$npx skillfedfor your agent
RESEARCH

Delta-rule recurrent states can hit 6-bit with no accuracy loss if you quantize both axes

on: STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Linear attention models like Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct replace growing KV caches with fixed-size recurrent states, but those states create their own memory problem at serving time. Each concurrent request needs its own persistent state, and at 70 simultaneous requests the FP32 state pool of Qwen already exceeds the memory footprint of its BF16 weights. Naive quantization makes things worse: errors don't just affect one step, they propagate forward through every subsequent Delta-rule update.

STEPQuant's central insight is that quantization error has two independent dimensions that prior work conflated. Temporally, state units with longer gate half-lives accumulate larger errors because the retention gate keeps propagating mistakes instead of washing them out—the Spearman correlation between gate half-life and accumulated INT6 error is 0.80 across all 2,304 Qwen heads. Spatially, equal-magnitude errors in different key rows produce very different readout damage, and the state matrix has large-magnitude outliers along both key rows and value columns simultaneously, not just one axis.

The temporal fix is lifetime-aware bit allocation: a mixed-precision scheme that weights reconstruction distortion by how long errors persist, then assigns higher precision to units where mistakes linger. A small fraction of the highest-risk units—just 1.39% of Qwen heads—get FP16 protection as sparse pivots. Removing those pivots alone costs 6.70 points on a three-task average at 4 bits.

The spatial fix is key-row-aware dual-axis fitting: separate scale factors for key rows and value columns, where row scales incorporate both magnitude and a calibrated impact score measuring how much each row's errors damage the output. Column scales are then fitted with heavier penalties on high-impact rows.

The results are striking. At 6 bits, STEPQuant matches FP32-state accuracy on both models across thirteen benchmarks, while uniform INT6 collapses to around 45% mean accuracy on long-generation tasks. At 4 bits, STEPQuant beats uniform INT8. Integrated into SGLang with fused CUDA kernels, the 6-bit configuration compresses recurrent-state memory by roughly 5× and cuts total serving memory by 68.7% on Qwen at a batch size of 512. State-update time drops by 65.6%.

The paper is honest about where the approach strains. The lifetime weight approximates error persistence through gate decay without fully modeling the key-dependent state transition. At 4 bits, KDA loses meaningful accuracy on long-generation tasks. The evaluation covers two specific architectures on fixed hardware; generalization to other hybrid designs remains open.

A principled two-axis quantization framework that gets Delta-rule recurrent states to 6 bits with negligible accuracy loss and 68.7% total memory reduction at serving scale.

Sources & links