Faro Research · Study

Who blocks the AI crawlers

We ran the same 47 agent-readiness checks on 1,103 company websites. Blocking is real, but it sits almost entirely in one industry. The bigger problem is the sites that let the crawlers in and then give them nothing worth reading.

Sample
1,103 websites
Fieldwork
August 25, 2026September 9, 2026
Published
September 9, 2026

Key findings

  • 18.9%of the 1,103 sites explicitly disallow at least one of GPTBot, ClaudeBot, PerplexityBot or Google-Extended in robots.txt.
  • 61.0%of news and media sites block at least one AI crawler. Developer-tools sites: 2.9%. That is a 21x gap between the two ends of the same sample.
  • 19.0%of sites serve no robots.txt at all. The file crawlers check for access rules simply is not there, and absence reads as permission rather than policy.
  • 51.0%of sites score below 50 out of 100 on agent readiness. Only 16 sites in the sample, 1.5% of it, score 80 or above.
  • 41.4%carry any JSON-LD structured data. 27.0% carry sameAs entity links. Most sites that let the crawlers in still cannot tell them who they are.
Share of sites blocking at least one major AI crawler, by industry

Explicit robots.txt disallow for GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Industries with at least 15 sites in the sample.

News & Media
61.0%
Food & Recipes
40.0%
Consumer Internet
35.0%
Education
34.3%
Sports
31.3%
Reference & Tools
26.5%
Entertainment
25.0%
Gaming
23.5%
Health & Medical
20.0%
Unclassified
19.8%
Automotive
11.1%
Finance & Banking
7.9%
Retail
6.9%
Travel & Hospitality
6.7%
Food & Restaurants
3.0%
B2B SaaS
3.0%
Developer Tools
2.9%
Marketing
0.0%

Blocking is an industry story, not a web-wide one

Read 18.9% on its own and you would think a fifth of the web has decided to keep AI systems out. Split it by industry and it gets much narrower. Publishers made that decision. Almost nobody else did.

News and media sites block at 61.0%. Food and recipe sites, the other category whose whole business is text a model can answer with, block at 40.0%. At the other end, developer tools sit on 2.9% and marketing sites on 0.0%. Retail (6.9%), finance and banking (7.9%) and travel (6.7%) barely block at all.

It is a rational split, not a contradictory one. The industries blocking are the ones whose product is the content itself, and who have watched an answer engine stand in for the click. The ones not blocking need to be found so they can sell something that is not the page.

The bigger number is the one nobody is arguing about

81.1% of these sites let every major AI crawler in. Most of them then give the crawler nothing it can use with any confidence.

58.6% publish no JSON-LD at all. 73.0% carry no sameAs links, so an agent has no way to match the site to the same company on LinkedIn, Wikidata or Crunchbase. 38.3% have a pricing page an agent can find, and 61.7% put a real price figure in HTML rather than behind a script or a contact form. 19.0% have no robots.txt, which means nothing points at a sitemap either.

The access argument is loud and it involves lawyers. The legibility problem is quiet, it involves nobody, and it affects four times as many sites.

The companies selling to the agent economy are ready for it. Their customers are not

Sort the sample by what a site is for rather than what industry it sits in and the gap gets sharper. Software vendors (137 sites) average 64.6 out of 100. Publishers (304) average 47.3. Merchants (291) average 46.9.

That puts vendors 17.3 points above publishers and 17.7 above merchants. The companies selling tools for the agent economy are ready for it. The publishers and retailers those agents have to read from and buy from are not. Readiness sits with the people who already think about this for a living.

The newer signals are rarer still. 25.9% of sites publish an llms.txt. 15.9% expose an MCP server. 9.2% publish an agents.json. 0.5% register a WebMCP tool, which is 6 sites out of 1,103. On this evidence, an agent arriving at a typical website can read it, cannot verify it, and cannot act on it.

Agent readiness by industry

Download CSV ↓

The 18 industries with at least 15 sites, ordered by mean score. That covers 978 of the 1,103 sites. The other 125 sit in industries too small to report on their own. “Unclassified” is the 368 sites the classifier could not place with confidence, shown here rather than quietly dropped.

IndustrySitesMean scoreBlocks a crawlerllms.txtJSON-LDVisible price
Developer Tools3467.32.9%76.5%76.5%88.2%
B2B SaaS3363.93%60.6%69.7%90.9%
Marketing2363.70%56.5%69.6%87%
Finance & Banking3853.17.9%28.9%44.7%73.7%
Food & Recipes3052.540%10%66.7%93.3%
News & Media5948.561%15.3%64.4%78%
Consumer Internet2047.935%25%25%65%
Unclassified36847.719.8%21.5%39.9%61.4%
Gaming1747.623.5%17.6%29.4%64.7%
Entertainment484725%20.8%29.2%64.6%
Travel & Hospitality4546.76.7%20%26.7%44.4%
Reference & Tools4946.426.5%18.4%38.8%57.1%
Education3546.334.3%17.1%25.7%45.7%
Health & Medical2546.220%12%44%56%
Food & Restaurants3346.13%27.3%18.2%39.4%
Retail87466.9%25.3%26.4%36.8%
Sports1644.131.3%18.8%31.3%50%
Automotive1840.511.1%11.1%22.2%27.8%

Method

  1. Population: 1,103 websites in the Faro score index carrying a complete result from the 47-check agent-readiness scan, run between 25 August and 9 September 2026. One row per registrable domain, newest scan only. A further 8 rows held partial results and were excluded.
  2. Each site was fetched once as an ordinary HTTP client, with no JavaScript execution. The checks read robots.txt, the homepage HTML and its headers, and a short list of well-known paths: /llms.txt, /agents.json, /sitemap.xml, /.well-known/mcp.json, /pricing.md and a few others. Nothing was crawled beyond that, and no site was scanned more than once a day.
  3. “Blocks at least one AI crawler” means robots.txt carries an explicit Disallow for GPTBot, ClaudeBot, PerplexityBot or Google-Extended. A site with no robots.txt counts as not blocking, because that is how a crawler treats it. That is why the 19.0% with no robots.txt sit outside the 18.9% blocking figure rather than inside it.
  4. Checks return true, false or uncertain. Uncertain counts as not passed throughout, so every pass rate here is a floor rather than an estimate. Only one check returns uncertain at any volume, the answer-first content check at 30.8%, and no figure above uses it.
  5. Industry labels come from Faro's own classifier and are shown for every industry with at least 15 sites. Segment labels (vendor, publisher, merchant) are assigned at the same time. 371 sites carry no segment, and those are excluded from the segment averages only.

What this cannot show

  • This is not a random sample of the web. It is the Faro score index: companies found through public lists and daily additions, weighted toward sites that are already commercially visible. Every figure describes this sample. None of them estimate a rate for the web as a whole.
  • robots.txt is a declaration, not an enforcement mechanism. This study measures what sites say about AI crawlers, not what any crawler does. Sites that block at the CDN or the firewall instead of in robots.txt count here as not blocking.
  • Homepage only, no JavaScript. A site that registers WebMCP tools from a bundled script, or renders its prices client-side, reads as not having them. Treat the 0.5% WebMCP figure as a floor on a signal that is hard to see in static HTML.
  • The fieldwork window is 16 days, so nothing here is a trend. It is one measurement. The method is published so that the same measurement can be taken again.

Cite this study

Faro (2026). Who blocks the AI crawlers. n=1,103, fieldwork August 25, 2026–September 9, 2026. https://byfaro.ai/research/ai-crawler-access-2026

The table is CC BY 4.0. Republish it, chart it, quote it. A link back to this page is the only condition. Journalists who need a different cut of the data can email hello@byfaro.ai.

Check your own site against this

Run the same 47 checks on your own siteCheck which AI crawlers your robots.txt allowsGenerate the JSON-LD 58.6% of these sites are missingWrite an llms.txtThe live Faro Score Index this study was cut from