$npx skillfedfor your agent
RESEARCH

Prefill and decode deserve separate quantization, and the accuracy gains prove it

on: Disaggregated Quantization: Specializing LLM Prefill and Decode

Prefill and decode have always shared weights in decoder-only LLMs, but they have never shared the same hardware bottleneck. Prefill is compute-bound at long context; decode is memory-bandwidth-bound at batch one. Disaggregated quantization (DQ) treats this asymmetry as a design axis rather than a nuisance, assigning each phase its own quantization format, and optionally its own weights and storage location.

The simplest intervention — format disaggregation — keeps a single set of weights but disables activation quantization specifically during decode. On NVFP4, this recovers meaningful accuracy on decode-heavy reasoning tasks across all seven tested Qwen 3 and Gemma 3 models, while decode latency actually improves by 2–3% because the activation quantization step is skipped. Prefill speed is unchanged. The cost is zero additional storage. That alone justifies treating it as a drop-in replacement for uniform NVFP4 inference.

Full disaggregation goes further: separate prefill and decode weight checkpoints, each trained toward the same response objective in a single forward-backward pass via QADD. The training trick is elegant — the SFT label mask that normally just gates the loss is repurposed to route each token position through either the prefill or decode linear pathway. Gradients reach the prefill weights through the prompt keys and values that decode attends to, so restricting supervision to response tokens still trains both sides. At 2-bit decode, fully-disaggregated models substantially outperform weight-only baselines on both decode-heavy and prefill-heavy tasks, while also delivering faster prefill via native NVFP4 compute.

The obvious objection is storage: two checkpoints on one device. Offloaded disaggregated prefill (ODP) answers this by streaming prefill weights from SSD block by block, reusing device memory carved out of the idle decode weights. The prefill checkpoint never fully occupies device memory. On Qwen3.8-27B in llama.cpp, ODP delivers a time-to-first-token speedup over the weight-only baseline at 8K context, with the crossover point where compute overtakes SSD loading sitting around 8K for all tested Qwen 3 models. Short prompts stall on loading; the paper is candid about this and explicitly flags ODP as unsuitable when time-to-first-token on short sequences is critical, and as largely inapplicable to mixture-of-experts architectures.

The prefiller concept extends this to frozen third-party checkpoints. Given any pre-quantized GGUF decoder — including opaque or closed-source ones — you can train only a new NVFP4 prefill checkpoint against it, leaving the decode weights untouched. For 1-bit IQ1_S decoders of Qwen3.8-27B, this more than doubles accuracy on both MMLU-Pro and MMMU-Pro, despite the QADD corpus containing no multimodal examples at all. The visual reasoning gains transfer from text-only training, which is a genuinely surprising result.

The paper's own limitations section notes that evaluations cover only batch-one decode and single-turn interactions. Multi-turn use is explicitly flagged as untested: cached assistant tokens carry decode-produced KV entries, but rebuilding that cache through prefill can produce different representations for the same token history. For agentic workloads with long conversation histories, that cache-policy dependence is an open question the authors do not resolve.

Phase-aware quantization that treats prefill and decode as distinct hardware problems — and solves both without requiring a single unified compromise.

Sources & links