skillfed

0.449 vs. 0.300: An Automated Skill Audit Out-Agreed Its Human Reviewers

Notes on MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills (arXiv:2604.20441) — Yingyong Hou, Xinyuan Lao, Huimei Wang, Qi Yao, Wei Chen, Bo-Sheng Huang, Fei Sun, Yu Lv, Weiqi Lei, Xueqi Wen, Pengfei Xia, Zhujun Tan, and 1 more · April 2026

Note published · written by SkillFed’s research pipeline from the paper above · how these notes are made

AI-assisted notes · reviewed by SkillFed Agentic benchmarks Bridge: benchmarks × security

Medical-research agent skills carry a failure mode that general-purpose skill checks don't catch: a skill can run cleanly, pass every schema check, and still fabricate a citation or wander into diagnostic territory it has no business entering. MedSkillAudit is a two-gate pre-deployment audit built to catch exactly that, run before a skill ever ships. A structural veto gate checks crash rate, schema compliance, result determinism, and code-injection surface; a domain-specific research veto gate checks for fabricated citations or data, practice-boundary violations, methodological fallacies, and code usability. A final score blends a 25-criterion static check (weighted 0.4) with a dynamic execution rubric (weighted 0.6), sorting each skill into one of four release tiers: Production Ready, Limited Release, Beta Only, Reject. The team ran it against 75 real medical-research skills — 15 each across five categories: Evidence Insight, Protocol Design, Data Analysis, Academic Writing, and a catch-all Other. Two experts then independently scored every skill on the same 0-100 scale and disposition ladder, which let the researchers measure two things at once: how closely the system's verdicts tracked human consensus, and how closely the two humans tracked each other.

Most of the audited skills weren't actually ready: mean consensus quality landed at 72.4 (SD 13.0), with 57.3% falling below the Limited Release threshold outright. The sharper number, though, is about the audit itself, not the skills it audited. System-to-expert agreement reached ICC(2,1) of 0.449, beating the 0.300 ICC the two human experts managed against each other — smaller score divergence (SD 9.5 vs. SD 12.4), and no systematic bias toward harsher or softer verdicts (Wilcoxon p = 0.613). That agreement wasn't even across categories: Protocol Design scored best (ICC 0.551), while Academic Writing inverted outright (-0.567) — not just noisier than the rest, moving in the opposite direction from the human raters entirely.

Key numbers

skills fell below the Limited Release deployment threshold57.3%
mean consensus quality score, out of 10072.4 (SD 13.0)
system-vs-expert ICC(2,1), against a 0.300 human inter-rater baseline0.449
Academic Writing category agreement — inverted, not just weakICC -0.567
medical-research skills audited, 5 categories x 15 each75

Skills related to this research

Related notes

References

  1. Yingyong Hou, Xinyuan Lao, Huimei Wang, et al., "MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills," arXiv:2604.20441 (2026)
  2. X. Li, W. Chen, Y. Liu, et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks," arXiv:2602.12670 (2026)
  3. Y. Liang, R. Zhong, H. Xu, et al., "SkillNet: Create, Evaluate, and Connect AI Skills," arXiv:2603.04448 (2026)
  4. ISO/IEC 25010:2011, Systems and Software Engineering — Systems and Software Quality Requirements and Evaluation (SQuaRE) — System and Software Quality Models, International Organization for Standardization (2011)
  5. T. K. Koo & M. Y. Li, "A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research," Journal of Chiropractic Medicine 15(2):155-163 (2016)