Learned Curation Halves the Skill Bank, Keeps the 9.69-Point Gain
Notes on CODESKILL: Learning Self-Evolving Skills for Coding Agents (arXiv:2605.25430) — Yanzhou Li, Yiran Zhang, Xiaoyu Zhang, Xiaoxia Liu, Yang Liu · May 2026
Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made
AI-assisted notes · reviewed by SkillFed Skill evolutionCODESKILL treats skill-bank curation as something a model learns, not a heuristic bolted onto a frozen agent. A management policy — trained through a three-stage curriculum with GRPO — watches coding-agent trajectories and decides whether to extract a new procedural skill (a high-level task strategy, or a local event-driven reaction to something like a recurring command failure), revise an existing one against new evidence, or merge and drop entries during maintenance. Three signals shape the reward: a rubric-based reward from an LLM judge scoring grounding, reusability, and actionability; sparse execution feedback on whether injecting the candidate skill actually raised the frozen downstream agent's pass rate on held-out tasks; and an alignment check on whether the agent's behavior matched what the skill claimed to trigger. Training uses a group size of six generations per prompt, a quality-reward weight of 0.25, and a KL coefficient of 0.02, optimized against a frozen Qwen3.5-35B-A3B coding agent.
Across EnvBench (Python and Java), SWE-Bench Verified, and Terminal-Bench 2, CODESKILL lifts average pass rate 9.69 points over a no-memory baseline and 4.01 points over the strongest prompt-based or memory baseline tested per benchmark (a fixed-prompt "Prompt Skill Mgmt." approach on EnvBench and SWE-Bench Verified, "Subtask Memory" on Terminal-Bench 2). EnvBench-Python shows the widest gap — pass rate nearly triples over no-skill, 18.60 versus 6.98. Maintenance is doing real work, not just bookkeeping: extraction and evolution alone would grow the bank to 1,252 skills; running the full lifecycle with merge/drop maintenance compacts that to 676 — roughly half — without giving back the gain. Swap the frozen downstream agent for GPT-5.4-mini and the improvement over no-skill still holds at 8.93 points, evidence the learned skills capture something more portable than one model's prompting quirks.
Key numbers
| Average pass-rate gain vs. no-skill baseline | +9.69 pts |
| Gain vs. strongest prior baseline | +4.01 pts |
| EnvBench-Python pass rate (CODESKILL vs. no-skill) | 18.60 vs. 6.98 |
| Skill bank size after maintenance (from 1,252 unpruned) | 676 skills |
| Gain when swapped to a different frozen agent (GPT-5.4-mini) | +8.93 pts |
Skills related to this research
Related notes
- A skill library trained on 5x5 robot grids beats genetic search at 10x10 — on all 7 EvoGym tasks →
- MUSE-Autoskill's Self-Written Skills Beat Human-Authored Ones, 85.24% vs. 81.17% →
- Text-based skill edits beat trained baselines by 6.7 points →
- Typed contracts + call templates: 82 vs. 47 ALFWorld wins, −23% tokens per game →
- Route skill treatment by task difficulty, not skill category: +13.2 points OOD success →
- The validation gate turns a 40% agent into an 89% one — skill-writing without it changes nothing →
References
- Li, Zhang, Zhang, Liu, Liu — CODESKILL: Learning Self-Evolving Skills for Coding Agents, arXiv:2605.25430 (2026)