$npx skillfedfor your agent
RESEARCH

Instruction-following feedback drowns out math in multi-teacher distillation

on: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Routing prompts to the right specialist is only half the problem. When you train separate RL experts for mathematics, coding, and instruction-following and then distill them into one student, the standard approach assigns each prompt to its matching teacher and treats all feedback equally. This paper shows that equal treatment is not neutral: instruction-following log-ratios are 2.3 to 4.4 times as dispersed as the pooled signal, mathematics log-ratios about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient under equal weights. The mathematics teacher is effectively drowned out before it can teach anything.

The fix is a single additional operation. Before updating the student, measure the standard deviation of each domain's teacher-student log-ratios within the current batch, compare it to the pooled spread, and scale each domain's advantages by the bounded ratio. No extra teacher calls, no learned router, no architectural change. Standard MOPD is the special case where every multiplier equals one.

The results across Qwen3.5 at 9B, 4B, and 2B are consistent. Label-routed MOPD never beats the strongest single-teacher student at any size and retains only 14–32% of the mathematics expert's gain at the 8K evaluation budget. DN-MOPD improves the six-task average over MOPD at every size and both evaluation budgets, with paired intervals above zero across three student seeds. At 8K the margin runs from 2.47 to 3.08 percentage points; at 16K from 1.17 to 2.36. Mathematics carries the largest gains; instruction-following changes little.

The fixed-weight controls are the most clarifying part of the paper. Amplifying mathematics feedback alone while leaving instruction-following unchanged recovers roughly half the improvement at 4B and little at 2B. Reducing instruction-following alone recovers most of it and raises mathematics by about three points. The problem is not that the mathematics teacher is too quiet; it is that the instruction-following teacher is too loud. DN-MOPD's per-batch estimation adds value mainly at 2B; at 9B and 4B, fixed weights measured once before training perform comparably.

One honest caveat: the benefit is configuration-dependent. An earlier Qwen3-4B construction with different expert checkpoints showed no clear gain, and the instruction-following multiplier in the Qwen3.5 runs almost always sits at the lower clipping bound, meaning the rule limits rather than equalizes that domain's scale. The experiments also use one expert pool per size, so variation from retraining experts is unquantified. Still, the core finding is hard to dismiss: deciding which teacher supervises a prompt and deciding how much its feedback counts are two separate design choices, and conflating them costs measurable capability.

Instruction-following feedback dominates multi-teacher distillation gradients; rescaling by batch-wise spread recovers the mathematics gain that label routing loses.

Sources & links