61 findings on a site we built for SEO
Field report · Mike Arbuzov · SkillFed Research ·
AI-assisted notes · reviewed by SkillFedA site with build-blocking structured-data lints, machine-readable mirrors and an enforced internal-linking floor still failed 61 checks drawn from the SEO skills our own editorial recommends — including FAQPage markup that same post called retired. 19% of the skills' criteria were stale too.
We ship a short editorial with every batch of skills we index. At the end of July one of them covered SEO and GEO skills — which ones we would actually reach for, which ones we would not. Publishing it raised the obvious question, and we ran it more or less on a whim: we recommend these. Do we pass them?
We expected to pass. This is not a site that treats SEO as an afterthought. The build fails if a page emits structured data outside an allowlist, or claims a rating we cannot substantiate. Every skill has a plain-text mirror and a JSON record for machines that would rather not parse HTML. There is a sitemap discipline, a uniqueness gate, and an internal-linking floor, all enforced as blocking lints rather than intentions. We had designed for this deliberately, with frontier models in the loop the entire way.
The audit distilled 355 checks from the ~25 skills the post covered and ran them against the live site. It returned 61 findings. Sixty of them survived an adversarial pass whose job was to refute them; none were refuted.
What it found
The worst one was ours coming back at us. Our own post says Google retired FAQ rich results in May 2026. Our skill pages were emitting FAQPage structured data at the moment we published that sentence. Honest markup, correctly formed, zero payoff, and we had written the paragraph explaining why days earlier.
The rest was the same shape — not sloppiness, but things nobody had thought to check:
Pages sharing an exact <title> |
110 across 38 groups; one name collided 23 times |
/best/playwright |
titled for one topic, listing the entire catalog |
| Missing URLs | returned 403, not 404 — no 404 page existed at all |
| Plain-text mirrors | indexable duplicates of their own HTML pages |
| Pages dropped from the build | 16 left live, self-canonical, unreachable, with no removal path |
| Author identity | absent everywhere, including on first-person editorial |
All of it was fixable and all of it is now fixed, across two generator releases. What deserves attention is not the list. It is why a team that had deliberately engineered for this still had a list.
The advice is stale too, and that is the useful part
69 of the 355 checks — 19% — were themselves stale or contested. Retired rich-result types presented as current. Sitemap attributes that stopped mattering years ago. llms.txt treated as an established standard when no major vendor has confirmed reading it.
We only saw that because roughly two dozen skills covered overlapping ground at different vintages and disagreed with each other. One skill saying "add FAQ markup" reads as authority. Several skills, one of which knows the format was retired, turns it into something you go verify against a primary source. The disagreement was the signal — and it is not available from a single source, because no skill reports its own decay.
This is not a capability problem
The temptation is to read this as models being not quite good enough yet. It isn't that.
Two things cannot be fixed by a better model. You cannot compress the long tail of a specialist domain into weights — the specific, local, and recent lose to the general every time. And whatever does make it in decays. SEO is close to a worst case: a model trained on the public corpus has absorbed a confident average of many vintages at once, with no way to date any of it, in a field where guidance is retired on a schedule.
Skills are how that knowledge gets to be current and specific: written by someone who works in the domain, dated, revisable without retraining anything. They are also, as above, wrong 19% of the time — which is an argument for reading several, not for reading none.
The part that changed how we think about our own product
Here is what actually happened, in order. We named the skills. Their contents went into context. The plan was built out of them, and the 355 checks effectively became the plan.
SkillFed's own installed workflow does the opposite. Plan mode produces a plan, the plan is approved, and skill search fires afterward, to help execute it.
Those are not the same operation, and the difference is structural rather than a matter of degree. Once a plan is approved, skills can only fill in steps that already exist. They cannot rescope it. Nothing downstream of an approved plan can introduce a check the plan never contained — and a plan that omits a check does not fail. It reports success.
Which makes us suspect our hook is in the wrong place, at least for work where coverage is the deliverable — audits, reviews, compliance passes, anything whose output is a list. For work where the deliverable is a change, and the shape of the job is already known, consulting skills at execution time is cheaper and probably fine.
If that is right, it has a requirement attached: skills have to be readable without being installed, or consulting a dozen of them before committing to a plan is too expensive to ever be the default. That is the argument for the plain-text mirrors and the machine index we had been thinking of as crawler surface. They are also how a skill gets into context early enough to change what you decide to do.
Method and limits
One site, one domain, one run, in July 2026. The 355 checks were distilled from the ~25 skills covered in a single editorial and are not a canonical SEO checklist; the findings are judgments with evidence attached, not scores, and the currency of each criterion was assessed against primary sources rather than taken on the skill's word.
The audit covered on-page, technical, and machine-surface quality. Off-site discovery — backlinks, index registration — is a separate problem with a separate fix, and it is not what this report is about.
We claim no traffic or ranking outcome. We fixed defects and shipped them; whether that moves anything is unmeasured and would be unmeasurable in this window regardless. And the ordering observation is a hypothesis about our own product drawn from a single instance, not a controlled comparison. It is a question we now think is worth asking, not a result.
Related: what a full census of the public skill corpus shows, and the research directory where we read the literature with the same discipline.
More from SkillFed Research
- Insight ·
Any AI chat can now run skill search — and you approve every request
No install, no account, no connector. Your chat writes an abstract wish, you paste the link back, and it reads five security-swept skills. The whole request is a URL in plain English — the privacy boundary is something you check, not something you're asked to trust.
- Insight ·
60,611 skills in the wild — what a full census of the public SKILL.md corpus shows
SkillFed walked all 6,177 repositories in its discovery queue end to end: 2.5× more unique skills than listings claimed, 13,122 per-agent variant files merged, and 86,956 vendored aggregator copies excluded — more copies than originals.
- Insight ·
The largest direction in agent-skill research is spreading outward, not settling down
Papers on agents that write their own skills land steadily farther from the direction's own semantic center month over month — the only trend in our analysis that survives multiple-comparison correction (BH p = 0.0016) — with no single axis carrying the drift.
- Insight ·
Zero of 184 recent papers connect skill self-authoring with skill security
Five of the six research-direction pairs in the recent agent-skill literature are bridged by dual-topic papers. The pair formed by its two largest directions — agents authoring their own skills, and securing skill files — is empty, and three null models say that is not chance.
- Field report ·
Agent-skills research didn't exist before 2023 — and its fastest-growing direction today is security
A SkillFed field map of 364 agent-skills papers, 2016–2026: none of this work existed before 2023, and skill security went from nothing to the second-fastest-growing direction in about three quarters.