$npx skillfedfor your agent
RESEARCH

A cheap agent with a distilled playbook beats a strong agent running blind

on: Recursive Harness Distillation across Agents for Robot Manipulation

A frozen vision-language-action model can be steered by an external agent without touching its weights — but only if that agent knows how to steer it. That knowledge is the hard part, and Recursive Harness Distillation (RHD) is a systematic way to build it.

The core idea: a capable agent (here, GPT-6 Astra) interacts with a frozen GR00T policy, records which interventions work and why, and distills that experience into a "playbook" — a structured document specifying when to intervene, what to do, and how to verify the outcome. A lighter, cheaper agent (GPT-5.6 Luna) then executes tasks using the playbook, and Astra reviews Luna's failures to revise the guidance. The loop closes recursively: recipient experience shapes what the teacher teaches next.

The numbers make the case. On SimplerEnv Bridge, GR00T alone hits 41.7% success. Luna without any playbook barely moves the needle, reaching 43.8%. With the refined playbook, Luna jumps to 66.7% — and that same playbook pushes Astra to 79.2%. In real-world Franka Panda trials, the playbook lifts Luna from 37.3% to 64.0% across three physical tasks. Without the playbook, Luna scored zero on the physical robot, because access to intervention tools alone does nothing if you don't know when or how to use them.

The ablation on teacher quality is particularly clarifying. When Luna acts as its own teacher — constructing and refining its own playbook recursively — the refined result achieves only 22.9% success, actually falling below Luna's unguided 43.8%. The hierarchy matters: a stronger teacher sees things the recipient cannot, and that asymmetry is what makes the distilled guidance useful.

The playbook itself is concrete and grounded. Excerpts show entries like: check that a corrected grasp geometry is preserved before closing the gripper; verify object-target geometry before release; when an OPEN instruction still produces closing actions, apply direct action correction rather than trusting the instruction alone. These aren't abstract heuristics — they're failure modes encountered during development, encoded as checkable conditions.

Refinement also matters more than initialization. The initial playbook actually hurts Luna relative to no playbook at all (31.3% vs. 43.8%), while the refined version more than doubles the baseline. Iterative feedback from the recipient is what makes the guidance usable rather than merely plausible.

Cost is addressed directly. Luna costs roughly $0.63 per episode versus $3.43 for Astra — about 5.5 times cheaper — making the deploy-a-light-agent-with-a-playbook strategy economically coherent, not just technically interesting.

The framework sidesteps fine-tuning entirely. No model parameters change at any point. The playbook is the artifact, and it transfers across agents with different capabilities. That portability is the real contribution: accumulated intervention experience becomes a reusable asset rather than implicit knowledge locked inside one expensive model's context.

A playbook distilled from a strong agent's failures can make a cheap agent outperform the strong agent running blind — and the numbers hold on a physical robot.

Sources & links