Self-graded exploration closes a 32-point reasoning gap — no labels needed
Notes on Unsupervised Skill Discovery for Agentic Data Analysis (arXiv:2606.06416) — Zhisong Qiu, Kang Song, Shengwei Tang, Shuofei Qiao, Lei Liang, Huajun Chen, Shumin Deng · June 2026
Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made
AI-assisted notes · reviewed by SkillFed Skill evolutionDataCOPE builds data-analysis skills without ever seeing a labeled example. Instead of grading trajectories against ground-truth answers, it manufactures its own quality signal out of the agent's exploration: for open-ended report tasks, an Adaptive Checklist Verifier writes a task-specific checklist, scores each report by how much of the checklist it verifiably covers, and rewrites the checklist itself whenever the agent starts gaming it; for fixed-answer reasoning tasks, an Answer Agreement Verifier clusters trajectories by their final answer and uses self-consistency — the relative size of a trajectory's answer cluster — as a secondary confidence signal. A Data-Analytic Agent samples the trajectories, the verifier sorts them into contrastive high- and low-quality groups, and a Skill Manager rewrites a Markdown skill file from that contrast, looping through generation, verification, and distillation with no human ever touching the exploration set.
Averaged over four matched base models — Claude, GPT, DeepSeek, and Qwen variants — DataCOPE lifts report-style accuracy on Deep Data Research from 47.4% to 57.1%, and reasoning accuracy on DABStep from 29.1% to 61.4%, with the reasoning gain concentrated on the benchmark's hard split. Both beat Anthropic's own Skill Creator tool run through Claude Code over the same exploration data, which manages only 51.3% and 51.7% respectively — the gap traces to the unsupervised verifier signal, not merely to having an agent write skills. The two verifier halves aren't interchangeable: on DABStep, keeping only self-consistency and dropping answer clustering scores worse than using no verifier signal at all (47.9% vs. 53.9%), because trajectories can converge confidently on the same wrong answer. The discovered skills also compress behavior, not just improve it — one configuration cut average token use by 73.4% (241k to 64k tokens) while accuracy rose from 44% to 64% under a fixed 15-turn budget. The one place unsupervised still trails supervised: with every exploration trajectory hand-labeled, the comparison baseline reaches 72.2% on DABStep, about nine points past what DataCOPE gets for free.
Key numbers
| Report-style accuracy, mean of 4 models (Deep Data Research) | 47.4% → 57.1% (+9.71 pts) |
| Reasoning-style accuracy, mean of 4 models (DABStep) | 29.1% → 61.4% (+32.30 pts) |
| Token use vs. accuracy (Claude Sonnet 4.6 + Claude Code, 15-turn cap) | –73.4% tokens (241k→64k) while accuracy rises 44%→64% |
| Margin over Anthropic's Skill Creator baseline (DABStep, mean) | 61.4% (DataCOPE) vs. 51.7% (Skill Creator) |
| Gap to full-label supervision (DABStep) | 62.8% unsupervised vs. 72.2% with every trajectory labeled |
Skills related to this research
Related notes
- Rubric-filtered training lifts a 9B model to 32% accuracy — outcome-only filtering caps out at 18% →
- Progressive Disclosure Triples Resource Touches — Pass Rate Moves Just 4 Points →
- Skill synthesis that checks its own work: +3 to +10 accuracy points, only 6% of skills still backfire →
- 215 Skills, 165 Contributors, No Fidelity Test →
- A weak model with a distilled skill beats its unaided teacher — at 1,000x lower inference cost →
- One in Four Model-Generated Skills Backfires on the Agent Using It →
- OpenSkill's verifier never sees the answer key, yet agrees with it 61% of the time -- and the skills it certifies beat closed-world baselines by 8.9 points →
- Decomposing agent traces into workflow, semantics, and attachments beats prompted summaries by 10.5% →
- The best skill scanner hits 98% recall — and still flags 937 of 4,000 safe skills as malicious →
References
- Qiu, Song, Tang, Qiao, Liang, Chen & Deng, "Unsupervised Skill Discovery for Agentic Data Analysis" (arXiv:2606.06416, 2026)
- Egg, Goyanes, Kingma, Mora, von Werra & Wolf, "DABStep: Data Agent Benchmark for Multi-Step Reasoning" (arXiv:2506.23719, 2025)
- Liu, Yu, Orini, Du & He, "Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models" (arXiv:2602.02039, 2026)
- Yao, Zhao, Yu, Du, Shafran, Narasimhan & Cao, "ReAct: Synergizing Reasoning and Acting in Language Models" (ICLR 2023)
- Anthropic, "Skill Creator" (github.com/anthropics/skills, 2026)