Treat agent harnesses as disposable discovery tools, not deployment scaffolding
on: Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Training agents on trajectories collected under specialized harnesses creates a distribution mismatch problem: the model learns to rely on workflow scaffolding that won't exist at inference time. Recursive Self-Rewrite addresses this by treating harnesses not as deployment infrastructure but as discovery tools, then systematically stripping their fingerprints from the resulting trajectories.
The mechanism has three stages. First, a single base model — Qwen-3.8-27B — runs under three distinct harnesses: a general terminal loop, a state-machine harness with explicit phase transitions, and a reflection harness that uses verifier outcomes to trigger continued execution up to a 180-minute wall-clock budget. Each harness solves a different subset of roughly 3K curated terminal tasks. The three-harness union solves 759 tasks, which is 34.3% more than the strongest individual harness alone. Of those 759, 288 are solved exclusively by one harness — 129 only by the general harness, 91 only by the reflection harness, 68 only by the state-machine harness. The complementarity is real and domain-specific, not just noise from unequal sampling.
Second, successful trajectories from the specialized harnesses are rewritten. A planner role extracts a runbook — key steps, checks, recovery strategies, but not the answer itself. A critic role screens runbooks for solution leakage and harness-specific artifacts. An executor role then re-solves the task from scratch in a fresh sandbox under the general harness, guided only by the private runbook. The runbook never appears in the public trajectory used for training. This expands the training set from roughly 2,001 source trajectories to 11,094 verified rewrites.
Third, the model is fine-tuned on those rewrites. Compared to direct SFT on the raw source trajectories, RSR gains 20.8 percentage points in pass@3 on Terminal-Bench 2, reaching 74.2%. On Terminal-Bench Hard it reaches 63.0% pass@3, up from 39.0% for the base model. The reflection harness produces trajectories that can exceed 100 turns; direct SFT on those causes the model to develop looping behavior — issuing the same commands repeatedly without progress. Rewriting compresses those trajectories and removes the harness-specific continuation signals, which appears to be what prevents that failure mode.
The process reward result on Long-Horizon Terminal Bench is more modest: 0.21 for the base model, 0.25 for direct SFT, 0.29 for RSR. No configuration solves any of the 46 long-horizon tasks outright, so the gain is partial progress rather than full completion. That's an honest ceiling to acknowledge.
The core insight — that harnesses are discovery tools whose successful outputs can be distilled into a general policy — is a cleaner framing than most self-improvement work, which tends to conflate the harness and the model. The critic's leakage detection and the fresh-sandbox re-execution are what make the distillation trustworthy rather than circular.
Harnesses as discovery tools, not deployment infrastructure — a clean framing that turns scaffold-dependent successes into genuinely portable agent capabilities.
Sources & links
Related on SkillFed
SkillLearnBench pits four automatic skill-generation methods against 20 verified real-world agent tasks. The best one closes only about 45% of the gap between no skill and a…
A web agent that turns its own successful runs into verified Python-function skills, then calls them as actions, beats both a static baseline and a text-memory skill agent on…