LeapQuant: Efficient Linear Attention Meets Near-Lossless State Quantization Claims, with Limits
A study on quantizing recurrent state in linear attention models shows big memory and speed gains, but results depend on model type and context length.
English translation prepared automatically for the Brief. Original sources and creator credits remain the reference.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the paper
- Authors
- Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
What they did and found
LeapQuant uses per-window quantization and Compensator Tokens to limit error. It reports near-FP32 accuracy with 8-bit state across several model families, plus kernel and end-to-end speedups.
Why it matters
Practitioners can consider 8-bit state for long-context models to reduce memory and improve throughput, while watching for context length and model mix.
LeapQuant lowers the memory tap on recurrent state by quantizing only at window ends and buffering high-precision updates. The approach also keeps the most significant outliers in high precision as Compensator Tokens to reduce quantization error.
In tests across Qwen, Kimi, and GLM families, accuracy remained near FP32 baselines while achieving notable speedups and memory cuts, especially when batching and longer contexts are used.
What remains uncertain
Results depend on model family and context length; not a universal deployment guarantee.
Original sources · 1
- LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization ↗arXiv · 2026-09-29
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.