Faro AI Signals·August 31, 2026

Forged crawler names hunt credential files as Google expands its answers

Crawler identity had a bad week. GreyNoise found 824 addresses wearing borrowed AI crawler names to hunt environment files and cloud keys, Cloudflare opened its bot directory so operators have to declare what their bots actually do, Google started expanding AI Overviews to full height on some queries, and two music publishers put scraping in front of a judge. If your AI crawler policy is a list of names, it is a list of names anyone can type.

ShareShare

This Week's Signals

1

Cloudflare opens its bot directory to operators who declare what they do

Source: Cloudflare

Why it matters for your score

Bot rules are shifting from a user agent string to a declared purpose, so decide now which purposes you allow on your site: search, reference, or model training.

2

824 addresses forged AI crawler names while hunting credential files

Source: GreyNoise Intelligence

Why it matters for your score

Any allowlist you built by pasting crawler names into robots.txt or a firewall rule is open to anyone who types the same name, and not one of the 824 addresses matched a published range.

3

Google expands AI Overviews to full height on some queries

Source: Search Engine Land

Why it matters for your score

On those queries, ranking and being read are different outcomes, and the citation inside the answer is the only placement that still reaches the reader.

4

Sony and Warner sue Anthropic, naming scraping alongside torrenting

Source: Axios

Why it matters for your score

How AI companies acquired content is now a live legal question, which turns your crawl policy into a record of what you actually permitted.

Emerging Audit Signals

What to do this week

  1. 1

    Test one crawler rule today. Send a request carrying the ClaudeBot or GPTBot name from an address outside the published range and see whether it gets through. If it does, your allowlist is a name filter, and the scan at /tools/ai-readiness-scan will flag the same gap across the rest of your crawl rules.

  2. 2

    Pull 30 days of robots.txt requests from your logs and list which named crawlers actually fetched it. Real crawlers ask for it, and in this data the forged ones never did once.

  3. 3

    Take your 20 highest value informational queries and record how far the first organic result now sits below the AI answer. Where the answer fills the screen, plan for the citation rather than the position.

Sources

← All editionsScan your site now →