Small reasoning models need to learn when to ask for help, not just think harder
Small reasoning models fail in two distinct ways, and conflating them wastes compute. This paper draws a clean line between execution bottlenecks — where the correct answer is already reachable through the model's own reasoning and self-refinement can recover it — and knowledge bottlenecks — where the missing ingredient is factual information that no amount of internal reflection can conjure. The distinction sounds obvious once stated, but the empirical work here makes it precise and actionable.
The authors intervene at intermediate reasoning states across eight models from two families, measuring how self-refinement and injected external information each shift the probability of reaching a correct answer. The finding is stark: epistemic verbalizations like "wait" and "hmm" are common even in the smallest models, but their causal effect on correctness scales sharply with model size. Small models express uncertainty frequently; they just can't convert that uncertainty into progress. Relevant oracle information, by contrast, produces substantially larger value gains than either epistemic cues or random text insertions — and the gap is especially wide in scientific reasoning, where knowledge-like bottlenecks dominate.
There's a second, less obvious finding: smaller models not only hit knowledge bottlenecks more often, they also exploit external information less effectively when it's handed to them. The help-seeking capability itself must be learned.
FlyBy operationalizes this. A 4B model is trained to reason first, diagnose what's blocking it, and then issue a targeted query to a stronger external model — choosing both what to ask and how much external compute to spend. The external model never sees the original problem, only the query, which limits direct answer leakage (confirmed empirically: adding a tool observation to a thinking-disabled Qwen3-4B raises accuracy by under 2 percentage points). Supervised fine-tuning on just 475 carefully filtered rescue trajectories bootstraps the query action, and 80 steps of cost-aware reinforcement learning calibrates when querying is worth its price.
The numbers are competitive. FlyBy-4B reaches 45.96% pass@8 on 1,158 hard problems spanning math, science, medicine, and general reasoning, beating Qwen3-14B's 41.64% at 2.7 times lower serving cost. It also beats Qwen3-8B at pass@1 (16.85% vs. 15.31%), meaning the improvement isn't just a sampling artifact. The 8B variant pushes pass@8 to 51.81%.
One detail worth noting: RL learns iterative multi-turn querying even though SFT demonstrated only single-query rescues. The cost penalty coefficient controls how aggressively the policy economizes, but iterative querying emerges across all tested penalty strengths. The two training stages have distinct roles — SFT teaches the mechanics of querying; RL teaches when and how many times to do it.
The practical constraint is real: this approach requires reliable API access to stronger external models, with the attendant privacy and availability concerns the authors acknowledge. But the core diagnostic insight — that thinking harder is the wrong operation at a knowledge bottleneck — is both well-supported and directly useful for anyone designing inference-time compute strategies.
Thinking longer only helps when the answer is already reachable; this paper proves the distinction empirically and trains small models to act on it.
Sources & links
Related on SkillFed
A soft-token compression framework cuts reusable agent-skill prompts to 30-60% of their length while general-purpose compressors collapse on procedural tasks.
SkillGuard extracts executable contracts from agent skill docs and validates only the values a skill actually depends on, cutting drift-monitoring false positives from 40% to zero…