Faro AI Signals·July 30, 2026

Cloudflare's September Crawler Block and the Citation Fingerprint Gap

Two forces are reshaping AI visibility at once: Cloudflare's three-way crawler split will block Training and Agent bots by default on ad-monetized pages from September 15, and new citation research across 25,337 data points shows that comparison pages and listicles — not homepages — are what AI engines actually pull from.

ShareShare

This Week's Signals

1

25,337-Citation Study Maps Exactly Which Page Types AI Engines Pull From by Industry

Source: DeltaV Digital

Why it matters for your score

Across 21,075 AI engine responses tracked between April and July 2026, comparison pages earned the highest citation rate of any page type at 1.87 citations per retrieval, while listicles captured 61% of citations in B2B technology services and homepages captured 55% in local services. If your content inventory does not include comparison and listicle formats, you are structurally absent from the page types AI engines favor in your category. Reddit appeared among the top-cited domains for 7 of 8 industries studied, which means unmanaged third-party mentions are likely outweighing your owned content in AI responses.

2

Cloudflare Splits AI Crawlers Into Three Classes, Blocks Training and Agent Bots by Default from September 15

Source: Cloudflare Blog (official)

Why it matters for your score

Cloudflare protects more than 20% of the global web, and its new three-way classification — Search, Agent, and Training — means a single block setting can now accidentally silence multi-purpose crawlers like Googlebot and BingBot, which Cloudflare evaluates under the most restrictive applicable rule. From September 15, Training and Agent crawlers are blocked by default on ad-monetized pages for all new domains and existing free-tier customers who do not opt out. Any business that wants AI agents to discover and act on its pages needs to review its Cloudflare settings before that date or those crawlers will be silently turned away.

3

Only 23.1% of SaaS llms.txt Files Match the Actual Spec — Most Published Files Add Noise, Not Signal

Source: Rankability / Arobis.ai

Why it matters for your score

A July 2026 SaaS audit found that while 57.7% of leading SaaS companies had something at /llms.txt, only 23.1% matched the actual spec with a correct H1, summary, and organized sections. An SE Ranking ML model found that removing the llms.txt variable from its citation-frequency model improved prediction accuracy, meaning a malformed file may hurt more than it helps. The priority is not publishing any file — it is publishing a correctly structured one, since among reachable top-1,000 sites the adoption rate is 15.8%, leaving real room for a well-formed file to stand out.

What to do this week

  1. 1

    Audit your Cloudflare crawler settings this week and confirm whether Agent and Training crawlers are permitted on your key landing pages — Cloudflare's September 15 default block applies to ad-monetized pages on free-tier accounts, and multi-purpose crawlers like Googlebot fall under the most restrictive rule you set.

  2. 2

    Check whether your site has a comparison page or listicle for each core service category — the DeltaV Digital data shows comparison pages earned 1.87 citations per retrieval, the highest of any page type, and listicles led citations in B2B technology services at 61%.

  3. 3

    Generate or validate your llms.txt file using Faro's /tools/llms-txt tool and confirm it includes a correct H1, a summary section, and organized content sections — the Rankability audit found only 23.1% of published SaaS files actually met spec, meaning most existing files are adding noise rather than signal.

Sources

← All editionsScan your site now →