Agent-readiness: what AI agents need to use your website
AI agents already browse ordinary websites on people’s behalf. An agent that cannot parse your structure, prices or contact route recommends a competitor it can parse. These are the ten signals that decide it.
Last reviewed 2026-09-04 · AgentProof, Harmony Future Holdings Limited, Dublin
What this page is
A plain-English account of the ten signals we measure, with the weight we give each and why. It is free, and it is the same corpus our paid report is built from. Unlike our sibling sites this is not a legal obligation — no law requires any of it, and we are not going to pretend otherwise.
Why this matters now
AI agents already browse ordinary websites on users’ behalf — Claude in Chrome since December 2025, ChatGPT agent mode, Gemini in Chrome. An agent that cannot parse a site’s structure, prices or contact route recommends a competitor it can parse.
Agent-readiness is 2026’s version of mobile-readiness in 2010: early, measurable, and cheapest to fix before the traffic arrives.
The ten signals, by weight
Weights follow the evidence, not the fashion. We re-based them in August 2026 after a market scan, and the re-basing moved things in directions that were commercially inconvenient for us.
| Signal | Weight | What a pass means |
|---|---|---|
| Labelled form inputs | 14 | Every form input has a label or aria-label. |
| Valid JSON-LD | 15 | Valid schema.org JSON-LD is present and parseable. |
| Real buttons and links | 12 | Interactive elements are real <button> and <a> elements, not bare clickable <div>s. |
| Organization or Product schema | 10 | The JSON-LD declares an Organization, Product, Service or LocalBusiness with real fields. |
| An explicit AI-crawler policy | 8 | robots.txt takes a deliberate position on GPTBot, ClaudeBot and Google-Extended rather than staying silent. |
| Title and meta description | 8 | A real <title> and meta description exist. |
| A machine-readable contact route | 8 | A mailto:, tel: or schema contactPoint exists. |
| Exactly one h1 | 6 | One <h1> states what the page is. |
| An llms.txt file | 6 | An /llms.txt file exists and describes the site for language models. |
| Declared language | 4 | <html lang> is declared. |
The two that carry the most weight
Labelled form inputs (14)
The highest-evidence lever on the list. Measured agent task-completion runs at roughly 78% on accessible sites versus 42% on inaccessible ones. An unlabelled input is an abandoned checkout in the agent economy.
Valid JSON-LD (15)
Structured data is the difference between an agent knowing your price and guessing it. It scores highest because it is the only signal that conveys facts rather than structure.
The useful surprise here
The two heaviest interaction signals — labelled inputs and real buttons — are accessibility work. If you have been putting off an accessibility pass, agent-readiness is the same job with a commercial return attached, and the evidence for it is stronger than for anything fashionable on this page.
Why llms.txt scores only 6
We weight this low on purpose, and it costs us to
llms.txt is the artefact everyone wants to sell you. The measured picture is thin: Google states it has no effect on Search, and Ahrefs measured 97% of llms.txt files receiving zero bot requests (May 2026).
It stays in our pack as a cheap hedge — it costs nothing, Lighthouse 13.3 audits for it, and it is the one site summary you fully control. But it is sold as exactly that, and it scores 6 out of 91. A vendor leading with llms.txt is selling you the easiest thing to deliver, not the thing that works.
Deciding about AI crawlers
The check is not “allow the crawlers”. It is take a deliberate position on GPTBot, ClaudeBot and Google-Extended rather than staying silent. Silence is a policy by accident.
One combination we flag as a warning: blanket-blocking AI crawlers while wanting agent customers. It is self-defeating, and it is usually the result of a copy-pasted robots.txt rather than a decision anyone made.
WebMCP, and why absence is not a failure
WebMCP lets a page register tools an agent can call directly, rather than making the agent drive the interface. It is a W3C Web Machine Learning Community Group draft, in Chrome origin trial through roughly Q1 2027.
Scored as INFO, never as a FAIL
Absence is normal today and we do not score a site down for it. Presence is being early to the interface agents will probably prefer. Anyone marking you down in 2026 for not shipping an origin-trial API is inventing a problem.
How the grade is worked out
Each check passes, warns or fails. A pass earns its full weight, a warn earns half, a fail earns nothing. INFO checks — WebMCP — are excluded from the denominator entirely, so they can never drag a score down. The percentage maps to a band:
| Score | Grade |
|---|---|
| 85 and above | A |
| 70–84 | B |
| 55–69 | C |
| 40–54 | D |
| below 40 | F |
The honest limits
This measures observable signals on the pages supplied, against conventions current agents and crawlers are documented to read, as at the corpus version shown on the report.
- It is not a guarantee of agent traffic, ranking or revenue.
- No law requires any of this. Every other product we make is about a legal deadline; this one is not, and we are not going to dress it up as one.
- The landscape changes monthly. That is why the report carries a corpus version and a regeneration entitlement — and why a copy of it goes stale.
The audit on your real pages, and the fix pack
The self-check below scores what you can see by eye. The audit runs the same checks against your actual pages, and the fix pack gives you paste-ready markup plus a retest.
See the pricing — from €49