0.449 vs. 0.300: An Automated Skill Audit Out-Agreed Its Human Reviewers
Notes on MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills (arXiv:2604.20441) — Yingyong Hou, Xinyuan Lao, Huimei Wang, Qi Yao, Wei Chen, Bo-Sheng Huang, Fei Sun, Yu Lv, Weiqi Lei, Xueqi Wen, Pengfei Xia, Zhujun Tan, and 1 more · April 2026
Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made
AI-assisted notes · reviewed by SkillFed Agentic benchmarks Bridge: benchmarks × securityMedical-research agent skills carry a failure mode that general-purpose skill checks don't catch: a skill can run cleanly, pass every schema check, and still fabricate a citation or wander into diagnostic territory it has no business entering. MedSkillAudit is a two-gate pre-deployment audit built to catch exactly that, run before a skill ever ships. A structural veto gate checks crash rate, schema compliance, result determinism, and code-injection surface; a domain-specific research veto gate checks for fabricated citations or data, practice-boundary violations, methodological fallacies, and code usability. A final score blends a 25-criterion static check (weighted 0.4) with a dynamic execution rubric (weighted 0.6), sorting each skill into one of four release tiers: Production Ready, Limited Release, Beta Only, Reject. The team ran it against 75 real medical-research skills — 15 each across five categories: Evidence Insight, Protocol Design, Data Analysis, Academic Writing, and a catch-all Other. Two experts then independently scored every skill on the same 0-100 scale and disposition ladder, which let the researchers measure two things at once: how closely the system's verdicts tracked human consensus, and how closely the two humans tracked each other.
Most of the audited skills weren't actually ready: mean consensus quality landed at 72.4 (SD 13.0), with 57.3% falling below the Limited Release threshold outright. The sharper number, though, is about the audit itself, not the skills it audited. System-to-expert agreement reached ICC(2,1) of 0.449, beating the 0.300 ICC the two human experts managed against each other — smaller score divergence (SD 9.5 vs. SD 12.4), and no systematic bias toward harsher or softer verdicts (Wilcoxon p = 0.613). That agreement wasn't even across categories: Protocol Design scored best (ICC 0.551), while Academic Writing inverted outright (-0.567) — not just noisier than the rest, moving in the opposite direction from the human raters entirely.
Key numbers
| skills fell below the Limited Release deployment threshold | 57.3% |
| mean consensus quality score, out of 100 | 72.4 (SD 13.0) |
| system-vs-expert ICC(2,1), against a 0.300 human inter-rater baseline | 0.449 |
| Academic Writing category agreement — inverted, not just weak | ICC -0.567 |
| medical-research skills audited, 5 categories x 15 each | 75 |
Skills related to this research
Related notes
- 71% of public healthcare skills carry no safety-boundary statement →
- An 8B Model Beats 4 Frontier LLMs by 25%+ — By Mining Its Own Skill Bank →
- 0.000 to 0.805: a 42-skill library rescues a model that can't solve a single hard RTL problem alone →
- Splitting SKILL.md into three layers lifts retrieval 12%, risk detection 24% →
References
- Yingyong Hou, Xinyuan Lao, Huimei Wang, et al., "MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills," arXiv:2604.20441 (2026)
- X. Li, W. Chen, Y. Liu, et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks," arXiv:2602.12670 (2026)
- Y. Liang, R. Zhong, H. Xu, et al., "SkillNet: Create, Evaluate, and Connect AI Skills," arXiv:2603.04448 (2026)
- ISO/IEC 25010:2011, Systems and Software Engineering — Systems and Software Quality Requirements and Evaluation (SQuaRE) — System and Software Quality Models, International Organization for Standardization (2011)
- T. K. Koo & M. Y. Li, "A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research," Journal of Chiropractic Medicine 15(2):155-163 (2016)