We have been running an AI readiness scan on any site that asks for one, and publishing what we find. Fifty-eight of them since the end of July.
The thing that surprised me is not the failures. It is which failures.
The basics are fine
Across those 58 sites:
| Check | Sites passing |
|---|---|
| Title and meta description | 55 of 58 |
| OpenGraph markup | 55 of 58 |
| AI crawlers permitted in robots.txt | 52 of 58 |
| Homepage readable to a named AI crawler | 57 of 58 |
| Canonical URL present | 45 of 58 |
That is a market that did the work. Titles, descriptions, share cards, crawler access — the things an SEO checklist has told you to fix for the last decade are done, on nearly every site we look at.
I did not expect that. Going in, my assumption was that we would find a long tail of sites blocking crawlers by accident and shipping pages with no title. They barely exist. Whatever else is wrong, people have been paying attention.
The new list has barely been started
Same 58 sites. Same month.
| Check | Sites failing |
|---|---|
| Content schema | 57 of 58 |
| llms-full.txt content feed | 48 of 58 |
| Entity consistency and sameAs | 47 of 58 |
| Direct-answer patterns | 43 of 58 |
| Readable-content density | 42 of 58 |
| llms.txt manifest | 35 of 58 |
Fifty-seven out of fifty-eight sites have no content schema. One site had it. That number was strange enough that I re-ran the whole set to make sure the check was not broken.
The average site in the sample fails 8.6 of the 19 checks we run. Best score recorded is 89, worst is 21, average 58.8.
Why the split falls where it does
Every item on the first list is about presentation. Does this page have a title. Does it produce a decent card when shared. Can a crawler reach it. Those were the constraints when the job was ranking a page in a list of blue links, and an industry spent fifteen years getting good at them.
Every item on the second list is about description. Not "can a machine reach this page" but "can a machine say what this business is, confidently enough to recommend it to someone."
Those are different problems, and the second one is newer than most people's current checklist.
When someone asks an AI engine to recommend a company like yours, the engine is
not ranking pages. It is assembling a description, and it prefers sources it can
describe without guessing. Content schema is how a page says what it is.
Entity consistency is how a company says it is the same company across
different places. llms.txt is how a site says which parts of it matter.
A site can pass every item on the old list and give an engine almost nothing to work with.
The honest caveat
These 58 are not a sample of the web.
They are sites belonging to people who went looking for an AI readiness scan — a group already pre-selected for suspecting they have a problem. The true numbers across the web are probably better than this, and I would not quote these figures as though they described everyone.
What the sample is good for is the shape: the split between a completed old list and an untouched new one held across almost every site, whatever its score. The best-performing site we scanned and the worst-performing one both had that shape.
What I would do about it
Not a fix list, because most of these are not one-afternoon fixes and I am not going to pretend otherwise. Three things that are genuinely useful:
Find out which list you are on. Open your homepage, view source, and search
for application/ld+json. If there is nothing, you are in the 57. If there is
something, read what @type it declares — plenty of sites have a block that
describes their website and never says what they sell.
Then check /llms.txt on your own domain. A 404 puts you in the 35. That
one genuinely is quick: a good file takes an afternoon, a rough one takes ten
minutes, and a rough one beats nothing.
Stop expecting one fix to move the number. The average site here fails 8.6 checks. People fix one, re-run, see almost no movement, and conclude the score is arbitrary. It is not — they fixed one of nine. Anyone selling you a single lever is selling you the lever.
The part that should worry you slightly
Six of 58 sites score 80 or above. Thirteen score an F.
That distribution is what an early market looks like — a wide middle, a small group who have worked it out, and almost nobody deliberately excellent in between. It is also, briefly, an advantage available to anyone who moves.
The same distribution existed for mobile page speed around 2016. Being fast was a differentiator for about two years, and then it was table stakes that earned you nothing. The people who benefited moved while it was still unusual.
I do not know how long this window stays open. I know it is open now, and that the gap between the middle of this distribution and the top of it is eight or nine specific things rather than a rebuild.
If you want to know which nine are yours, the scan takes about a minute and tells you without a conversation.
For what happened when we pointed the same audit at our own domain — one citation across seven AI engines — the write-up is here.
Figures from 58 distinct external domains scanned between 30 July and 24 August 2026, most recent scan per domain, internal tests excluded. No client or scanned domain is identified.
