Faro Research

AI readiness research

Faro scans websites for AI agent readiness all day, every day. These are the studies that come out of that index: how many sites block ChatGPT and Claude, how many publish structured data an AI system can actually use, and how far apart the industries are.

Every study states its sample, its method and what it cannot show. Every table comes with a CSV. Take the numbers and use them.

Published studies

StudySeptember 9, 2026 · n = 1,103 websites

Who blocks the AI crawlers

We ran the same 47 agent-readiness checks on 1,103 company websites. Blocking is real, but it sits almost entirely in one industry. The bigger problem is the sites that let the crawlers in and then give them nothing worth reading.

18.9% of the 1,103 sites explicitly disallow61.0% of news and media sites block19.0% of sites serve no robots.txt at

Where the data comes from

Faro runs a 47-check agent readiness scan across 7 categories and returns a score out of 100. It's the same scan behind the free AI Readiness Scan anyone can run on their own site, and the results feed the public Faro Score Index. The research is cut from that index, not from a separate crawl built to prove a point.

Sites enter the index through public company lists and a daily discovery job. Each one is fetched as an ordinary HTTP client with no JavaScript execution, no more than once a day, reading robots.txt, the homepage and its headers, and a short list of well-known paths. One row per registrable domain, newest scan wins.

That method has a real cost, and it's stated on every study: this is not a random sample of the web, and no figure here should be read as a web-wide rate. It's a large, consistently measured sample of commercially visible websites, measured the same way every time so the same measurement can be taken again.

The 7 categories behind every score

Agent permissions

Whether robots.txt exists, and whether GPTBot, ClaudeBot, PerplexityBot and Google-Extended are allowed through it.

Content discoverability

Sitemap, llms.txt, canonical tags, and the HTTP Link header that points an agent at them.

Structured data

JSON-LD, Organization and Product schema, sameAs entity links, OpenGraph, the OKF bundle.

Pricing transparency

Whether a price an agent can read exists in HTML, and whether /pricing.md is published.

Agent infrastructure

MCP server, WebMCP tools, agents.json, A2A card, OpenAPI, OAuth discovery, x402.

Agent operability

Whether buttons, forms and landmarks are labelled well enough for an agent to act on the page.

Trust and identity

Contact, about, privacy and terms pages, plus whether the homepage answers its own question up front.

Using and republishing the data

Every study ships the table behind it as a CSV on a stable URL, and the page carries Dataset markup pointing at that file. The point is that the table can be checked and reused as data rather than screenshotted.

The tables are licensed CC BY 4.0. Republish them, chart them, quote the figures in an article or a deck. A link back to the study page is the only condition, and you don't need to ask first. Each study also carries a ready-made citation line with the sample size and fieldwork window in it, so the number and its caveat travel together.

Working on a story and need a cut that isn't on the page, by industry, by country or by individual check? Email hello@byfaro.ai with the question you are trying to answer and we will run it.

What we are measuring next

One study a quarter. These are the questions in the queue, not published findings: agent readiness by US state, so the geography of the gap is visible; the share of the largest SaaS companies publishing llms.txt and agents.json; and Fortune 500 against Inc 5000, to test whether size predicts readiness or works against it.

If there is a question you want answered from this index, say so. It's cheaper for us to run a query than for you to build the sample.

Questions we get about the data

How many websites has Faro scanned?

The public score index held just over 1,100 domains on 9 September 2026 and grows every day. The first study uses the 1,103 of those that carried a complete 47-check result at the time it was cut.

Which AI crawlers does the research check for?

GPTBot (ChatGPT), ClaudeBot (Anthropic), PerplexityBot and Google-Extended (Gemini and AI Overviews). A site counts as blocking when its robots.txt carries an explicit Disallow for any one of them.

Can I republish these numbers?

Yes. Every table is CC BY 4.0. Republish it, chart it, quote it in an article. The only condition is a link back to the study page the figures came from. You don't need to ask first.

Is this a random sample of the web?

No, and no study here will claim otherwise. It's the Faro score index: companies found through public lists and daily additions, weighted toward sites that are already commercially visible. Each study says so in its own limitations section.

How often does a study get updated?

The numbers in a published study are frozen on purpose. Once someone cites a figure, that figure has to keep matching the page. Studies get re-cut with fresh numbers at six months, and the update is dated on the page when it happens.

Can I get a cut of the data that is not published?

Usually. Email hello@byfaro.ai with the question you're trying to answer. Splits by industry, by country, by company size and by individual check are all possible from the same index.

Measure your own site against this

Run the same 47 checks on your siteSee which AI crawlers your robots.txt allowsGenerate the JSON-LD most sites are missingBrowse the live Faro Score IndexGuide: robots.txt and AI crawlersGuide: what AI readiness actually means