Letting agents build their own tools from failures beats human-curated tool libraries
on: TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
Two failures motivate this work, and both are quantified precisely enough to be uncomfortable. Equipping a frozen time series agent with 21 expert-curated tools makes it worse on anomaly detection by 5.8 to 8.8 points across every backbone tested — despite the library containing a dedicated anomaly tool. A single pass of generic self-revision changes 147 answers across ten tasks, breaks 56 previously correct ones, and moves the final score by less than a point. The second failure is the more insidious: aggregate accuracy is a bad loss function for self-improvement because harm and repair cancel inside it.
TimeEvo's response to both problems is architectural. Rather than supplying tools before deployment, it clusters an agent's own errors into failure buckets, writes a measurement contract for each bucket, and synthesizes Python functions that return numbers — never verdicts, never dataset-specific constants. The candidate library then faces two validation levels: a cheap per-tool prefilter requiring at least one attributable repair on held-out errors from the tool's own bucket, and a whole-library paired gate that counts exactly how many previously correct answers the new library breaks versus fixes. A library passes only when it demonstrably fixes more than it harms; a rejected round leaves the agent untouched.
The transfer result is the finding most worth sitting with. A library grown on the cheapest backbone, installed into three stronger models with no test data involved in the selection, lifts each of them on the ten-task mean. Without the library, those stronger models score between 54.6 and 57.1 — indistinguishable from the cheap one. What they were missing was not reasoning capacity but the right measurements. The tools compute the same numbers regardless of which model reads them; what differs is how deeply a stronger model exploits the evidence.
The ablation on pre-installed roots is equally pointed. Starting evolution from the 21 expert-curated tools rather than an empty library produces worse results in all four ablation cells, and the gate rejects the evolved library outright on two of them. Human-written tools do not become the right tools when an agent evolves on top of them — the same misalignment the paper identifies in baselines reappears inside the method itself when the root is non-empty.
The case studies make the mechanism concrete. In both examined repairs, the base agent had read the question correctly and simply had no number to compute. One deterministic measurement — a normalized residual against Fourier candidates, a dispersion comparison across windows — settled the answer. The repairs are not better reasoning; they are missing arithmetic, supplied exactly once.
Failure-driven tool synthesis with a paired admission gate that counts harm directly beats every human-curated and self-refinement baseline across all thirty task-backbone combinations.
Sources & links
Related on SkillFed
Payloads embedded in a coding-agent skill's documentation examples — not its instructions — bypass defenses 11.6% to 33.5% of the time, versus 0% for explicit attacks, across…
SkillForge traces cloud-support agent failures back to specific skill-file sections and auto-rewrites them, gaining +9-12 points of consistency per skill and beating a mature…