Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

LeapQuant uses per-window quantization and Compensator Tokens to limit error. It reports near-FP32 accuracy with 8-bit state across several model families, plus kernel and end-to-end speedups.

Why it matters

Practitioners can consider 8-bit state for long-context models to reduce memory and improve throughput, while watching for context length and model mix.

LeapQuant lowers the memory tap on recurrent state by quantizing only at window ends and buffering high-precision updates. The approach also keeps the most significant outliers in high precision as Compensator Tokens to reduce quantization error.

In tests across Qwen, Kimi, and GLM families, accuracy remained near FP32 baselines while achieving notable speedups and memory cuts, especially when batching and longer contexts are used.

What remains uncertain

Results depend on model family and context length; not a universal deployment guarantee.

Read the paper PDF ↗

Original sources · 1
  1. LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization ↗arXiv · 2026-09-29

Check the original paper for its authors, methods, version and access terms.