Treating VLAs as callable tools inside executable code fixes failures raw scaling cannot
on: Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
Current VLA and world-action models have a structural problem that more data won't fix: a QwenGR00T policy trained on LIBERO with its instruction masked succeeds 92.3% of the time, versus 96.2% with the instruction present. The gap is tiny. The policy is mostly learning scene-to-trajectory mappings, not task understanding. Viewpoint shifts or layout changes collapse success toward zero, and nothing in an action chunk can detect or repair that.
The response proposed here is to give physical robots the same medium software agents already work in: executable code. "Physical Coding" means representing world state and execution procedure as programs that can be inspected, verified, and revised. Code as World tracks objects, relations, constraints, and progress predicates. Code as Policy organizes planning, tool calls, verification, and recovery. A VLA becomes one callable tool inside that program, not the sole locus of task intelligence.
HexaAnything is the instantiation. Its Harness maintains typed execution contracts, runs an independent verifier that returns semantic verdicts rather than accepting the model's self-report, and keeps provenance on every observation. On RoboCasa365, wrapping XR-1 inside this Harness raises Composite-Unseen success from 34.3% to 38.3%. On three individual tasks run over 100 seeds each, the gains are 31, 15, and 17 percentage points over the same underlying VLA. The mechanism is concrete: in LoadKebabSandwich, the coding agent holds a predicate requiring both ingredients inside the oven before the door can close; the baseline sometimes closes the door with one ingredient still outside, making recovery impossible.
Traces collected by the Harness then train HexaModel v0.1, a fine-tuned 27B model. Placed back in the same Harness, it improves over its base on every split, with the largest gain on Composite-Unseen, the split most dependent on decomposition and recovery. Tool revision on RoboDojo raises success on three tasks from 0%, 40%, and 0% to 80%, 100%, and 100% over two rounds, with weights unchanged throughout.
PhyBench extends the argument beyond manipulation. The benchmark requires an agent to design physics experiments, operate apparatus, read instruments, and submit quantitative estimates. With strong frontier models, HexaAnything completes all three tasks—spring constant, gravitational acceleration, coupled oscillator frequencies—with mean relative errors below 5% across valid runs. The gravitational acceleration estimate in one traced run lands at 0.004% relative error.
The paper is honest about what it hasn't shown: autonomous architecture search, unrestricted self-rewriting, and real-robot results beyond three trials per task. The real-robot numbers compare against published results on different hardware, which the authors flag explicitly. The broader four-target evolution program—Harness, model, data, embodiment—is a roadmap, not a demonstrated system. What the experiments do establish is that making state and procedure explicit and executable produces measurable, attributable gains at each level the paper actually tests.
Wrapping a VLA inside an executable world-and-policy program beats the same VLA alone by attributable, reproducible margins—the mechanism is the key result.
Sources & links
Related on SkillFed
SkillJect automates poisoned agent-skill generation with a closed-loop Attack/Victim/Evaluate agent loop, reaching 80.7% average attack success on Claude Code versus 0% for naive…
XSkill separates a multimodal agent's memory into a stable skill library and a disposable experience bank, both grounded in screenshots — beating single-memory baselines by up to…