71% of public healthcare skills carry no safety-boundary statement
Notes on An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance (arXiv:2605.02709) — Gelei Xu, Ningzhi Tang, Xueyang Li, Toby Jia-Jun Li, Zhi Zheng, Wei Jin, Yiyu Shi · May 2026
Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made
AI-assisted notes · reviewed by SkillFed Agentic benchmarks Bridge: benchmarks × securityResearchers built the first systematic census of healthcare-focused agent skills — self-contained instruction packages that an agent loads only once a task matches the skill's stated description, a pattern the spec calls progressive disclosure. Starting from 58,159 public skills on ClawHub (an April 2026 snapshot), the team used an LLM classifier to pull 557 healthcare-related skills, then had a second LLM read each skill's full SKILL.md file to annotate ten dimensions covering function, care-cycle stage, intended user, input modality, autonomy level, and clinical decision impact.
The resulting picture is lopsided. Diagnosis and treatment reasoning — the tasks healthcare-agent papers foreground, with 40.0% of papers tagging diagnosis — show up in only 22.4% of skills; workflow automation, research support, and patient monitoring dominate instead. Specialized clinical inputs are nearly absent: medical images and physiological signals each serve as the primary input in just 9 skills apiece, an order of magnitude below structured forms and text chat. Autonomy clusters at L3 ("delegated execution"), but 13 skills combine high autonomy (L4 or L5) with outputs the authors classify as actively driving a clinical decision. Across the whole corpus, only 163 of 557 skills (29%) state an explicit safety boundary.
Key numbers
| Healthcare skills identified from ClawHub | 557 of 58,159 screened |
| Diagnosis-support skills, vs. 40.0% of healthcare-agent papers | 22.4% of skills |
| Skills with an explicit safety-boundary statement | 163 of 557 (29%) |
| Medical-imaging / physiological-signal skills (primary input) | 9 skills each (<2%) |
| High-autonomy skills (L4/L5) that drive clinical decisions | 13 skills |
Skills related to this research
Related notes
- 40,285 Skills Later, Supply Still Doesn't Match Demand →
- Same skill, +22 points for Claude Sonnet, +5.5 for Nemotron Nano →
- Agent-skill catalogs already top 700,000 entries — curation hasn't caught up →
- Curated Skills Lift Success Rates 16.2 Points — Self-Generated Ones Cost You 1.3 →
- 26.1% of Community Skills Ship With a Vulnerability →
- 0.449 vs. 0.300: An Automated Skill Audit Out-Agreed Its Human Reviewers →
- A skill compiler lifts Claude Code pass rates from 21% to 33% — and catches a missing safety guard in 95% of real-world skills →
- 97.6% of Injection and Poisoning Caught, Only 90.2% When Skills Interact →
- A fine-tuned 8B retriever hits 83 NDCG@10 — a 12B off-the-shelf model manages 55 →
References
- Xu, G., Tang, N., Li, X., Li, T. J., Zheng, Z., Jin, W., & Shi, Y. (2026). An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance. arXiv:2605.02709.
- Xu, G., Li, X., Chen, Y., Duan, Y., Wu, S., Yu, H., Chiu, C., Ni, J., Tang, N., Li, T. J., et al. (2026). A Comprehensive Survey of AI Agents in Healthcare. Journal of Biomedical Informatics, 105045.
- Ling, G., Zhong, S., & Huang, R. (2026). Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality. arXiv:2602.08004.
- Tang, N., Chen, C., Fang, Z., Xu, G., Dhakal, M., Shi, Y., McMillan, C., Huang, Y., & Li, T. J. (2026). Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions. arXiv:2604.00436.
- Wu, Y., & Zhang, Y. (2026). Agent Skills from the Perspective of Procedural Memory: A Survey. Authorea Preprints.