Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

Researchers introduced a training-free method that quantizes the recurrent state once every 16 tokens and uses Compensator Tokens to hold large outliers in high precision, reducing memory and speeding up decoding while keeping 32-bit accuracy.

Why it matters

This could lower hardware needs and speed up running large language models, but results depend on the model and setup.

The method lowers how often the state is turned into a compact form by reusing a boundary state for a window and buffering high-precision token updates. A small set of high-precision Compensator Tokens captures major outliers, letting the rest be quantized more tightly.

Tests across several models show that using 8-bit state can match 32-bit accuracy on many tasks and run faster, though results depend on model type and hardware.

What remains uncertain

Results depend on model type and hardware; gains may not occur in every deployment or with all configurations.

Read the paper PDF ↗

Original sources · 1
  1. LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization ↗arXiv · 2026-09-29

Check the original paper for its authors, methods, version and access terms.