$npx skillfedfor your agent

Skill evolution beats verifier-only refinement by 4 points, no refinement by 10

Notes on COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows (arXiv:2607.01709) — Zongxia Li, Dawei Liu, Fuxiao Liu, Yuhang Zhou, Xiyang Wu, Jingxi Chen, Jingjing Xie, Xiao-Ming Wu, Lichao Sun · July 2026

Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made

AI-assisted notes · reviewed by SkillFed Skill evolution

ComfyClaw treats ComfyUI workflow construction as typed graph editing, not prompt rewriting. An agent inserts and connects nodes, tunes samplers, attaches LoRAs, and applies regional conditioning; invalid edits get reverted automatically. A region-level VLM verifier decomposes each prompt into a checklist of binary requirements, scores the image against them plus a holistic 1-10 detail score, and turns any failures into localized repair instructions — regional prompting to isolate a specific limb, for instance — that drive the next edit pass. Successes and failures get clustered across batches of prompts and distilled through a skill evolution loop that creates, revises, reinforces, merges, or deletes entries in a library of versioned Agent Skills, with each candidate mutation tested on synthesized held-out prompts and committed only if it doesn't degrade held-out performance.

Across four benchmarks (GenEval2, DPG-Bench, OneIG-EN, OneIG-ZH), three agent backbones (Claude Sonnet 4.5, Qwen-3.6-35B-A3B, Gemma-4-E4B-it), and two diffusion backbones, the full system beat a single-pass no-refinement baseline by roughly 10 points of average score, and a verifier-only variant without skill evolution by roughly 4 points. Human raters scored 2,400 generated images and preferred ComfyClaw's output across benchmarks — 4.49 vs. 3.65 on a 5-point scale for one backbone. The gains don't come mainly from better wording: only 39% of the agent's edits touched prompt text. Evolved skills accounted for 70% of skill lookups on DPG-Bench and 56% on GenEval2, but dropped to 7.5-16% on the more style-driven OneIG splits, where the agent leaned on predefined skills instead.

Key numbers

vs. no-refinement baseline+10 pts avg score
vs. verifier-only (no skill evolution)+4 pts avg score
human preference, Likert 1-5 (LongCat)4.49 vs 3.65
workflow edits that are not prompt-text60.7%
evolved-skill reliance, best vs. weakest benchmark70% → 7.5%

Skills related to this research

seedance-recipes Seedance-recipes provides production-ready recipe patterns for video content across genres: product, lifestyle, drama, music video, landscape, commercial, animation, and more. Each recipe preserves core creative constraints while inviting customization of subject, camera, lighting, and sound. Use recipes as proven starting shapes, not rigid templates.★ 5,445 team-frontend-debug team-frontend-debug orchestrates a multi-role team for frontend quality assurance, routing feature lists to a testing pipeline or bug reports to a debugging pipeline. Both flows leverage Chrome DevTools MCP for browser inspection, DOM analysis, console monitoring, and performance tracing. The coordinator role parses your input, spawns specialized workers (tester, reproducer, analyzer, fixer, verifier), and manages progress across phases.★ 2,142 paper-writing Paper-writing chains five specialized sub-skills into a single end-to-end workflow: outline planning, figure generation, LaTeX authoring, PDF compilation, and iterative review. Feed it a research narrative or existing plan, specify your target venue (ICLR, NeurIPS, ICML, CVPR, ACL, AAAI, ACM, IEEE), and optionally reference a style guide—the skill handles the rest, producing a polished paper directory with source and compiled output.★ 13,939 skill-evolution Skill Evolution monitors how your skills perform across sessions by analyzing user edits and success metrics, then suggests targeted improvements with confidence scores. Apply changes safely with automatic version snapshots and rollback capability, or use the holdout-promotion gate to validate candidates against a sealed eval set before graduating them.★ 208

Related notes

References

  1. Z. Li, D. Liu, F. Liu, Y. Zhou, X. Wu, J. Chen, J. Xie, X. Wu, and L. Sun, "ComfyClaw: Self-Evolving Skill Harnesses for Image Generation Workflows," arXiv:2607.01709, 2026.
  2. Agent Skills (2026), "Agent Skills Specification," agentskills.io/specification.
  3. G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar, "Voyager: An Open-Ended Embodied Agent with Large Language Models," arXiv:2305.16291, 2023.
  4. Z. He, S. Huang, X. Qu, Y. Li, T. Zhu, Y. Cheng, and Y. Yang, "GEMS: Agent-Native Multimodal Generation with Memory and Skills," arXiv:2603.28088, 2026.
  5. P. Xia, J. Chen, H. Wang, J. Liu, K. Zeng, Y. Wang, S. Han, Y. Zhou, X. Zhao, and H. Chen, "SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning," arXiv:2602.08234, 2026.