skillfed

Research

The literature behind SkillFed’s skill search

191 agent-skill papers read, verified, and distilled into claim-first notes — mapped across four research directions, with the bridges between them and one conspicuous gap.

2026 Feb Mar Apr May Jun Jul pre-2026 · 116 MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills Beware of Agentic Botnets: Scalable Untargeted Promptware Attacks via Universal and Transferable Adversarial HalluSquatting Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills Sealing the Audit-Runtime Gap for LLM Skills SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces When Skills Lie: Hidden-Comment Injection in LLM Agents Semia: Auditing Agent Skills via Constraint-Guided Representation Synthesis Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks Behavioral Integrity Verification for AI Agent Skills Methods for Formal Verification of Agent Skills: Three Layers Toward a Mechanically Checkable Capability-Containment Proof "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild When Agents Talk: Discourse, Manipulation, and Risk in an Agentic Social Network Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections POISE: Position-Aware Undetectable Skill Injection on LLM Agents SkillClone: Multi-Modal Clone Detection and Clone Propagation Analysis in the Agent Skill Ecosystem Malicious Or Not: Adding Repository Context to Agent Skill Classification How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study SkillMutator: Benchmarking and Defending Language-and-Code Cross-modal Attacks on LLM Agent Skills SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation SkillGuard: A Permission-Centric Framework for Agent Skill Security FORTIS: Benchmarking Over-Privilege in Agent Skills Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree Formal Analysis and Supply Chain Security for Agentic AI Skills HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents? SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills SkillTester: Benchmarking Utility and Security of Agent Skills Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces Skill security born 2025-Q3 ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning Reinforcement Learning for Self-Improving Agent with Skill Library Skill-R1: Agent Skill Evolution via Reinforcement Learning Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents Co-Evolving Skill Generation and Policy Optimization SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Evolving Programmatic Skill Networks SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills SkillOS: Learning Skill Curation for Self-Evolving Agents SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System SkillX: Automatically Constructing Skill Knowledge Bases for Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks SkillMaster: Toward Autonomous Skill Mastery in LLM Agents MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization Skill Coverage: A Test Adequacy Metric for Agent Skills SkillGen: Verified Inference-Time Agent Skill Synthesis SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents SkillGrad: Optimizing Agent Skills Like Gradient Descent SkillClaw: Let Skills Evolve Collectively with Agentic Evolver SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills SkillNet: Create, Evaluate, and Connect AI Skills CUA-Skill: Develop Skills for Computer Using Agent A Framework for Evaluating Agentic Skills at Scale Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale CODESKILL: Learning Self-Evolving Skills for Coding Agents OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction OpenSkill: Open-World Self-Evolution for LLM Agents MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality SkillOpt: Executive Strategy for Self-Evolving Agent Skills CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification EvoSkill: Automated Skill Discovery for Multi-Agent Systems An Empirical Study of Downstream Adaptation for Agent Skills A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications Agent Skill Evaluation and Evolution: Frameworks and Benchmarks SkillEvolver: Skill Learning as a Meta-Skill SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills From Anatomy to Smells: An Empirical Study of SKILL.md in Agent Skills Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries SkillWiki: A Living Knowledge Infrastructure for Agent Skills Harnessing Agent Skills: Architectural Patterns and a Reference Architecture for Skill-Mediated LLM Agents Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts Skill evolution born 2025-Q2 Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents SkillReducer: Optimizing LLM Agent Skills for Token Efficiency SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents Graph of Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills Skill Is Not Document: A Query-Conditional Benchmark and Two-Stage Retriever for LLM Agent Skill Routing Skill Retrieval Augmentation for Agentic AI SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale SkillRouter: Skill Routing for LLM Agents at Scale Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose SkillFlow: Scalable and Efficient Agent Skill Retrieval System Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings Generative Skill Composition for LLM Agents SkillsInjector: Dynamic Skill Context Construction for LLM Agents Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents SkillAxe: Sharpening LLM-Authored Agent Skills Through Evaluation-Guided Self-Refinement Skill retrieval born 2025-Q2 SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization ColPackAgent: Agent-Skill-Guided Hard-Particle Monte Carlo Workflows for Colloidal Packing NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills Agentic Publication Protocol: An Attempt to Modernize Scientific Publication Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System From Agent-Only Social Networks to Autonomous Scientific Research: Lessons from OpenClaw and Moltbook, and the Architecture of ClawdLab and Beach.Science EpochX: Building the Infrastructure for an Emergent Agent Civilization STEM Agent: A Self-Adapting, Tool-Enabled, Extensible Architecture for Multi-Protocol AI Agent Systems KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill AgentClick: A Skill-Based Human-in-the-Loop Review Layer for Terminal AI Agents MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis LEGO: An LLM Skill-Based Front-End Design Generation Platform Pomona: Continuous Code Quality Improvement via Small, Agentic Pull Requests at Bloomberg EffiSkill: Agent Skill Based Automated Code Efficiency Optimization How Agent Skills Fail under Long Contexts: A White-Box Study in Code Auditing Agentic benchmarks born 2026-Q1 DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback OpenAssistant Conversations - Democratizing Large Language Model Alignment Analyzing Modular Approaches for Visual Question Decomposition On Data Engineering for Scaling LLM Terminal Capabilities Automated Educational Question Generation at Different Bloom's Skill Levels Using Large Language Models: Strategies and Evaluation R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training Agent Skill Acquisition for Large Language Models via CycleQD Cerbero-7B: A Leap Forward in Language-Specific LLMs Through Enhanced Chat Corpus Generation and Evaluation A Comparative Study of Code Generation using ChatGPT 3.5 across 10 Programming Languages Rescue: Ranking LLM Responses with Partial Ordering to Improve Response Generation Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models Compute Optimal Scaling of Skills: Knowledge vs Reasoning LLM skill training born 2023-Q1 Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft Odyssey : Empowering Minecraft Agents with Open-World Skills Parallelized Planning-Acting for Efficient LLM-based Multi-Agent Systems in Minecraft MindAgent: Emergent Gaming Interaction Voyager: An Open-Ended Embodied Agent with Large Language Models Hierarchical Cooperative Multi-Agent Reinforcement Learning with Skill Discovery Learning Communication Skills in Multi-task Multi-agent Deep Reinforcement Learning STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models Scalable Multi-agent Covering Option Discovery based on Kronecker Graphs MetaAgents: Large Language Model Based Agents for Decision-Making on Teaming Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent Cooperation On Multi-Agent Learning in Team Sports Games OSExpert: Computer-Use Agents Learning Professional Skills via Exploration Exploration Based Language Learning for Text-Based Games Curiosity-Driven Exploration via Latent Bayesian Surprise Offline Multi-agent Continual Cooperation via Skill Partition and Reuse Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Improving Agent Interactions in Virtual Environments with Language Models Unsupervised Skill-Discovery and Skill-Learning in Minecraft Self-Supervised Exploration via Disagreement Curiosity-Driven Exploration by Self-Supervised Prediction Playful Agentic Robot Learning ALAN: Autonomously Exploring Robotic Agents in the Real World See and Think: Embodied Agent in Virtual Environment A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals Closed-Loop Vision-Language Planning for Multi-Agent Coordination Augmenting Autotelic Agents with Large Language Models SIMA 2: A Generalist Embodied Agent for Virtual Worlds Social Structure Matters in 3D Human-Human Interaction Generation The Information Geometry of Unsupervised Reinforcement Learning Learning with AMIGo: Adversarially Motivated Intrinsic Goals Bisimulation Makes Analogies in Goal-Conditioned Reinforcement Learning Latent Skill Planning for Exploration and Transfer Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents Lipschitz-constrained Unsupervised Skill Discovery ELSIM: End-to-end learning of reusable skills through intrinsic motivation Visual Reinforcement Learning with Imagined Goals SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning DexHoldem: Playing Texas Hold'em with Dexterous Embodied System Learning agile soccer skills for a bipedal robot with deep reinforcement learning RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents Goal-Conditioned Reinforcement Learning with Imagined Subgoals Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning MoCapAct: A Multi-Task Dataset for Simulated Humanoid Control WildLMa: Long Horizon Loco-Manipulation in the Wild Choreographer: Learning and Adapting Skills in Imagination Unsupervised Perceptual Rewards for Imitation Learning Deep visual foresight for planning robot motion Accelerating Reinforcement Learning with Learned Skill Priors PI-QT-Opt: Predictive Information Improves Multi-Task Robotic Reinforcement Learning at Scale Continual Quadruped Robots Coordination via Semantic Skill Discovery Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills Learning Predictive Models From Observation and Interaction Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds Learning Latent Plans from Play MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making One-Shot High-Fidelity Imitation: Training Large-Scale Deep Nets with RL RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation Meta-learning Parameterized Skills Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Composing Task-Agnostic Policies with Deep Reinforcement Learning AgentVLN: Towards Agentic Vision-and-Language Navigation Skill-based Model-based Reinforcement Learning GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks Inner Monologue: Embodied Reasoning through Planning with Language Models Chat with the Environment: Interactive Multimodal Perception Using Large Language Models Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations GEMS: Agent-Native Multimodal Generation with Memory and Skills CLIPort: What and Where Pathways for Robotic Manipulation MolmoWeb: Open Visual Web Agent and Open Data for the Open Web Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning PANDO: Efficient Multimodal AI Agents via Online Skill Distillation Do As I Can, Not As I Say: Grounding Language in Robotic Affordances Developmental Scaffolding with Large Language Models SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling Language Conditioned Imitation Learning Over Unstructured Data Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills Agentic Skill Discovery Agent Skills Should Go Beyond Text: The Case for Visual Skills Language to Rewards for Robotic Skill Synthesis Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck MMSkills: Towards Multimodal Skills for General Visual Agents XSkill: Continual Learning from Experience and Skills in Multimodal Agents Robotic skill learning born 2016-Q4 The shape of agent-skills research radius = publication month (pre-2026 compressed to the core; 2026 by month to the rim) · wedge = direction · arrow length = drift speed · outer dots = undifferentiated core (105) llm-agent-skillsrobotics-skillsrl-skill-theoryother / reject
How to read it. The field map from our published analysis: all 364 harvested papers — the 191 core notes below plus the adjacent and rejected work around them — clustered into directions. Each dot is one paper and fades in at its publication date, so the field blooms outward on its own. Radius is time (everything before 2026 packed into the core, 2026 fanned out month by month), the wedge is the direction, each wedge’s arrow shows how its center of mass has drifted, and color is the paper’s domain: violet for agent-skills work, coral for robotics, blue for RL theory, grey for the undifferentiated core. The direction cards below use our separate per-paper tags over the 191 core notes — the names overlap, the counts differ.Read the full field analysis →

five directions

Every paper in the corpus carries one of five direction tags. Each card below leads to that direction’s own page — every note it holds, newest paper first.

Skill security 42 papers

Attacks on and defenses for skill files — malicious skills, injection, supply chains.

Where the directions meet — and where they don’t

45 of the 191 papers in this corpus genuinely work in two directions at once. Counting those bridge papers by pair, over the papers published since October 2025:

benchmarks × evolution 13
benchmarks × security 13
benchmarks × retrieval 7
evolution × retrieval 7
retrieval × security 4
evolution × security 0 expected 3–13

Five of the six pairs are bridged. The empty one belongs to the corpus’s two largest directions: among 184 recent papers, not one connects skill evolution — agents writing their own skills — with skill security. Null models over these direction sizes expect 3–13 such papers; observed is zero. Security research audits skills written by others; evolution research builds skills the agent writes for itself; nobody yet audits a self-authored skill at authoring time.Read the full field analysis →

Put the research to work — search skills on SkillFed →

Reports & insights

  1. Insight · Aug 2026 Any AI chat can now run skill search — and you approve every request

    No install, no account, no connector. Your chat writes an abstract wish, you paste the link back, and it reads five security-swept skills. The whole request is a URL in plain English — the privacy boundary is something you check, not something you're asked to trust.

  2. Field report · Jul 2026 61 findings on a site we built for SEO

    A site with build-blocking structured-data lints, machine-readable mirrors and an enforced internal-linking floor still failed 61 checks drawn from the SEO skills our own editorial recommends — including FAQPage markup that same post called retired. 19% of the skills' criteria were stale too.

  3. Insight · Jul 2026 60,611 skills in the wild — what a full census of the public SKILL.md corpus shows

    SkillFed walked all 6,177 repositories in its discovery queue end to end: 2.5× more unique skills than listings claimed, 13,122 per-agent variant files merged, and 86,956 vendored aggregator copies excluded — more copies than originals.

  4. Insight · Jul 2026 The largest direction in agent-skill research is spreading outward, not settling down

    Papers on agents that write their own skills land steadily farther from the direction's own semantic center month over month — the only trend in our analysis that survives multiple-comparison correction (BH p = 0.0016) — with no single axis carrying the drift.

  5. Insight · Jul 2026 Zero of 184 recent papers connect skill self-authoring with skill security

    Five of the six research-direction pairs in the recent agent-skill literature are bridged by dual-topic papers. The pair formed by its two largest directions — agents authoring their own skills, and securing skill files — is empty, and three null models say that is not chance.

  6. Field report · Jul 2026 Agent-skills research didn't exist before 2023 — and its fastest-growing direction today is security

    A SkillFed field map of 364 agent-skills papers, 2016–2026: none of this work existed before 2023, and skill security went from nothing to the second-fastest-growing direction in about three quarters.

Editorial deep-dives on the skills themselves live on the blog.

Skill evolution 88 papers

Agents that write, revise, and govern their own skill libraries.

  1. 23% of Agent Skills Already Bundle Executable Code, Not Just Prompts

    2026-07-21 — Skillware is Fan and Lan's name for what an agent skill actually is once you stop treating it as a prompt: a three-layer object. The Skill Artifact is just the natural-language task spec. Wrapped…

  2. A co-evolving skill library lifts tool-use accuracy from 27.7% to 32.0% — with fewer tool calls, not more

    2026-07-15 — SPyCE trains multimodal agents that think with images by distilling every successful multi-step trajectory into a two-tier hierarchical skill library , rather than collapsing it into a scalar…

  3. Ten anchored examples recover 88-110% of an oracle metric's gains

    2026-07-14 — Self-evolving agent loops assume a reliable evaluator already exists to grade each attempt. This paper drops that assumption and evolves the evaluator itself. The metric takes shape as an expression…

  4. Flat retrieval breaks once a skill library hits the tens-to-hundreds range Bridge: evolution × retrieval

    2026-07-11 — This survey audits 124 papers on agent skill systems published between 2023 and 2026 (2 from 2023, 19 from 2025, 103 from 2026, cutoff May 31, 2026) and builds three shared tools for comparing them.…

  5. Skill retirement flatlines past a ~45% false-pass rate — and no amount of data brings it back

    2026-07-08 — Self-evolving agents that keep accumulating skills need a curator — a mechanism that retires a skill once its observed pass rate drops to a set threshold, which is what keeps a growing library from…

  6. Evolving the improver — not just the skill — accounts for all of ALFWorld's gain and half of SealQA's

    2026-07-06 — MetaSkill-Evolve doesn't stop at letting an agent revise its own skills — it lets the agent revise the machinery that does the revising. Each search branch pairs a task skill with a meta-skill :…

Latest 6 of 88 — all 88 in Skill evolution →

Skill security 42 papers

Attacks on and defenses for skill files — malicious skills, injection, supply chains.

  1. 15 cloned listings hijack skill retrieval 93% of the time

    2026-07-15 — SkillSec-Eval breaks the agent skill lifecycle into six stages — authoring, storage, retrieval, planner selection, execution, evolution — and gives each one its own threat taxonomy. Badhe and…

  2. Comparing a skill's claims to its code lifts misalignment detection from 0.45 to 0.89 Macro-F1

    2026-07-12 — SkillsMP, the largest open-source Agent Skills marketplace, supplied a corpus of 264,937 normalized skill packages out of 273,657 catalog entries, each split into three layers: metadata (name,…

  3. 216,938 skills, and only 114,565 come with a paper trail

    2026-07-08 — SkillCenter builds its library through a five-stage pipeline. Source acquisition feeds an LLM-based pre-filter called SkillGate , which screens raw material for actionability before any generation…

  4. Nearly 1 in 5 Skill Forks Add Security-Sensitive Instructions Bridge: security × benchmarks

    2026-07-03 — Researchers screened GitHub for agent skill repositories with at least 20,000 stars and 2,000 forks and landed on six, including Anthropic's own anthropics/skills, obra/superpowers, and…

  5. Stack five skills, multiply hidden-intent risk 14x

    2026-07-02 — SkillFuzz treats skill composition — not the individual skill — as the unit worth testing. An LLM first compiles each skill's natural-language instructions into a structured skill contract :…

  6. Whole-Trace Checking Catches 95.8% of Skill Policy Violations

    2026-06-25 — VIGIL is a runtime reference monitor for agent skills. It abstracts raw tool calls into typed events, then grounds each skill's natural-language specification into a policy that names the actual…

Latest 6 of 42 — all 42 in Skill security →

Skill retrieval 30 papers

Finding the right skill in a library too big to load — the problem SkillFed's own search is built on.

  1. Flat skill packs lift 20-book QA accuracy from 0.26 to 0.46 — a second routing level erases the gain

    2026-07-20 — Agent Skills packs — the folder-based standard for handing an agent on-demand expertise — expose only a short description until a task matches it, then load an indexed body, then the specific…

  2. Task-decomposition reranking beats the best baseline 78.7 vs 73.1 on ALFWorld-unseen, using just 1.3 skills per task

    2026-07-07 — SkillReranker treats skill selection as a graph-matching problem, not a similarity search. It decomposes a task into an ordered sequence of subtasks and intermediate sub-states, then parses every…

  3. A 3.9M-parameter skill sequencer closes 80% of the gap to hand-picked "gold" skill sets

    2026-06-30 — LLM agents built on skill libraries hit a bottleneck once the library grows: choosing what to load stops being a lookup problem and becomes a joint decision over subset, count, and order — three…

  4. Compose agents from skills, not fixed roles: +2 points over the best topology-only baseline, only a 0.96-point dip when the skill library changes

    2026-06-18 — Existing graph-based multi-agent design treats agents as closed-set entities : fix a roster of agents, roles, or groups first, then optimize who talks to whom. SIGMA drops that assumption. Given a…

  5. One Feedback Pass Takes Skill-Chain Decomposition From 51% to 68% Accuracy

    2026-06-16 — Compositional skill routing formalizes what happens when a query needs more than one skill: decompose it into atomic sub-tasks, retrieve a skill for each, then compose the results into an executable…

  6. Verification-gated skills add up to 12 points on KernelBench — pull retrieval at inference and most of it vanishes

    2026-06-15 — daVinci-kernel splits CUDA/Triton kernel generation across three roles under one shared LLM backbone. A Selection Agent retrieves candidate optimization techniques through BM25 pre-filtering plus…

Latest 6 of 30 — all 30 in Skill retrieval →

Agentic benchmarks 26 papers

Evaluation harnesses that test whether skills actually help the agents using them.

  1. 8/10 → 3/10: a 300K-character context collapses a code-audit skill's pass rate — relevant or not

    2026-07-20 — A fixed 24-check code-audit agent skill got stress-tested inside a production-derived auditing task, run through Codex with gpt-5.4-mini. The task and its verification checks stayed constant; only…

  2. A 96,401-skill curated corpus lifts agent pass rates +7.5pp — until coverage runs out Bridge: benchmarks × retrieval

    2026-07-17 — SkillCorpus turns the sprawling public SKILL.md ecosystem into one deployable, licence-clean corpus. A six-stage pipeline parses, deduplicates, and quality-scores roughly 821,000 crawled skill…

  3. Code review, testing, and security auditing claim 35% of task assignments; requirements analysis gets 2%

    2026-07-10 — The first large-scale empirical study of software-engineering skills starts with a brutal filtering funnel: 775,790 skills pulled from four public marketplaces — ClawHub, SkillHub, SkillNet, and…

  4. Coding agents violate their own skill's embedded logic in up to 70% of test cases

    2026-07-10 — SkillLogic is a static-analysis framework that reads an agent skill file and extracts the logical relations binding its instructions together: preconditions that gate an action, postconditions…

  5. LLM-generated skills move data-science accuracy 1.2 points — same as filler text

    2026-07-08 — The team built one reusable skill file per stage of a data-science agent's workflow — data preparation, data extraction, statistical analysis, and reporting — generating each with Gemini 2.5 Pro…

  6. Rubric-filtered training lifts a 9B model to 32% accuracy — outcome-only filtering caps out at 18% Bridge: benchmarks × evolution

    2026-07-02 — SkillCoach splits agentic skill-use into four dimensions scored separately — skill selection, skill following, skill composition , and skill-grounded reflection — pulled from real agent rollouts…

Latest 6 of 26 — all 26 in Agentic benchmarks →

Frontier & other 5 papers

Work that doesn't sit in a single direction yet — the field's unclaimed edges.

  1. A skill-specific LoRA beats prompting the full SKILL.md by 5.2 points and cuts token cost 6.6%

    2026-06-15 — Skill-to-LoRA (S2L) treats a SKILL.md file as training data, not runtime cargo. Offline, a teacher model reads the full skill document and generates synthetic task-response pairs that demonstrate…

  2. SKIM cuts agent skills to 30-60% of their length for a 1-2 point accuracy hit

    2026-06-10 — SKIM (SKIll coMpression) replaces a reusable agent skill's full instructions with a small set of learned soft tokens , so the skill no longer has to be pasted into every prompt in full. A compressor…

  3. Typed contracts + call templates: 82 vs. 47 ALFWorld wins, −23% tokens per game

    2026-05-27 — Skill-as-Pseudocode (SaP) rewrites markdown skill libraries into typed pseudocode, so agents stop re-deriving schemas and call syntax from prose on every retrieval. The pipeline clusters similar…

  4. Conditioning the perception latent on the text skill card cuts cross-modal redundancy 9x — and gets 2.3x the step-consistency of 5-sample self-consistency at roughly the same latency as 1 sample

    2026-05-08 — Agents built on vision-language models rarely repeat themselves. Ask the same policy to complete the same web task twice and the click sequences drift, even though the underlying reasoning hasn't…

  5. Compiling a skill for its model drops regressions from 15% to 4.5%

    2026-04-03 — Scale first: two public catalogs hold 118,000 agent skills between them — 28,990 on clawhub.ai, 89,280 on skills.sh. Running that catalog against eight LLMs and three harnesses (BareAgent, OpenCode,…