Faro Research · Study
We ran the same 47 agent-readiness checks on 1,103 company websites. Blocking is real, but it sits almost entirely in one industry. The bigger problem is the sites that let the crawlers in and then give them nothing worth reading.
Explicit robots.txt disallow for GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Industries with at least 15 sites in the sample.
Read 18.9% on its own and you would think a fifth of the web has decided to keep AI systems out. Split it by industry and it gets much narrower. Publishers made that decision. Almost nobody else did.
News and media sites block at 61.0%. Food and recipe sites, the other category whose whole business is text a model can answer with, block at 40.0%. At the other end, developer tools sit on 2.9% and marketing sites on 0.0%. Retail (6.9%), finance and banking (7.9%) and travel (6.7%) barely block at all.
It is a rational split, not a contradictory one. The industries blocking are the ones whose product is the content itself, and who have watched an answer engine stand in for the click. The ones not blocking need to be found so they can sell something that is not the page.
81.1% of these sites let every major AI crawler in. Most of them then give the crawler nothing it can use with any confidence.
58.6% publish no JSON-LD at all. 73.0% carry no sameAs links, so an agent has no way to match the site to the same company on LinkedIn, Wikidata or Crunchbase. 38.3% have a pricing page an agent can find, and 61.7% put a real price figure in HTML rather than behind a script or a contact form. 19.0% have no robots.txt, which means nothing points at a sitemap either.
The access argument is loud and it involves lawyers. The legibility problem is quiet, it involves nobody, and it affects four times as many sites.
Sort the sample by what a site is for rather than what industry it sits in and the gap gets sharper. Software vendors (137 sites) average 64.6 out of 100. Publishers (304) average 47.3. Merchants (291) average 46.9.
That puts vendors 17.3 points above publishers and 17.7 above merchants. The companies selling tools for the agent economy are ready for it. The publishers and retailers those agents have to read from and buy from are not. Readiness sits with the people who already think about this for a living.
The newer signals are rarer still. 25.9% of sites publish an llms.txt. 15.9% expose an MCP server. 9.2% publish an agents.json. 0.5% register a WebMCP tool, which is 6 sites out of 1,103. On this evidence, an agent arriving at a typical website can read it, cannot verify it, and cannot act on it.
The 18 industries with at least 15 sites, ordered by mean score. That covers 978 of the 1,103 sites. The other 125 sit in industries too small to report on their own. “Unclassified” is the 368 sites the classifier could not place with confidence, shown here rather than quietly dropped.
| Industry | Sites | Mean score | Blocks a crawler | llms.txt | JSON-LD | Visible price |
|---|---|---|---|---|---|---|
| Developer Tools | 34 | 67.3 | 2.9% | 76.5% | 76.5% | 88.2% |
| B2B SaaS | 33 | 63.9 | 3% | 60.6% | 69.7% | 90.9% |
| Marketing | 23 | 63.7 | 0% | 56.5% | 69.6% | 87% |
| Finance & Banking | 38 | 53.1 | 7.9% | 28.9% | 44.7% | 73.7% |
| Food & Recipes | 30 | 52.5 | 40% | 10% | 66.7% | 93.3% |
| News & Media | 59 | 48.5 | 61% | 15.3% | 64.4% | 78% |
| Consumer Internet | 20 | 47.9 | 35% | 25% | 25% | 65% |
| Unclassified | 368 | 47.7 | 19.8% | 21.5% | 39.9% | 61.4% |
| Gaming | 17 | 47.6 | 23.5% | 17.6% | 29.4% | 64.7% |
| Entertainment | 48 | 47 | 25% | 20.8% | 29.2% | 64.6% |
| Travel & Hospitality | 45 | 46.7 | 6.7% | 20% | 26.7% | 44.4% |
| Reference & Tools | 49 | 46.4 | 26.5% | 18.4% | 38.8% | 57.1% |
| Education | 35 | 46.3 | 34.3% | 17.1% | 25.7% | 45.7% |
| Health & Medical | 25 | 46.2 | 20% | 12% | 44% | 56% |
| Food & Restaurants | 33 | 46.1 | 3% | 27.3% | 18.2% | 39.4% |
| Retail | 87 | 46 | 6.9% | 25.3% | 26.4% | 36.8% |
| Sports | 16 | 44.1 | 31.3% | 18.8% | 31.3% | 50% |
| Automotive | 18 | 40.5 | 11.1% | 11.1% | 22.2% | 27.8% |
Faro (2026). Who blocks the AI crawlers. n=1,103, fieldwork August 25, 2026–September 9, 2026. https://byfaro.ai/research/ai-crawler-access-2026
The table is CC BY 4.0. Republish it, chart it, quote it. A link back to this page is the only condition. Journalists who need a different cut of the data can email hello@byfaro.ai.