AI Readiness

How to Get Your Business Cited by ChatGPT, Claude, and Perplexity

Faro Editorial

July 2, 2026 · 10 min read

8-signal checklist for getting cited by AI: ChatGPT, Claude, Perplexity

When an AI model recommends a brand, the downstream effect is substantial. A 2026 study of AI-influenced purchase behavior found that brands recommended first by AI are 389% more likely to be Googled afterward, and brands appearing in "best/top" framing are 425% more likely. When AI recommends a brand to a new customer, that customer is 182% more likely to search for it directly and 117% more likely to visit the site.

These are not marginal effects. They are category-defining advantages that compound over time, because AI citation builds brand awareness at the top of the funnel without requiring paid acquisition. The question most marketers are not asking yet: what actually determines whether AI recommends you?

Why AI Citations Are Not Random

AI systems do not recommend businesses by lottery. They form recommendations based on the evidence available to them: structured data, web content, user-generated signals, training data, and live retrieval results. Businesses that have strong, consistent signals across these layers are recommended more often, more confidently, and in more favorable positions. Businesses with thin or inconsistent signals are omitted or misrepresented.

The pattern mirrors how search worked in the early 2010s: businesses that understood the technical requirements earned disproportionate early advantage. The window where early movers gain durable ground over laggards is open now.

The Eight Signals That Drive AI Citations

1. llms.txt: your business context file

llms.txt is the most direct signal you can give an AI system about what your business does, who it serves, and what AI should say about you. It is a plain text file at the root of your domain, readable by AI crawlers, that functions like a brief for every AI model that visits your site. Without it, AI models either guess your context from web content or have no business context at all.

An effective llms.txt is specific: it names your primary use case, your customer type, your main differentiators, and links to your key pages. Vague descriptions ("we help companies grow") produce vague AI representations. Specific, verifiable descriptions ("we provide AI agent readiness audits for marketing agencies and B2B SaaS companies") give AI models exactly what they need to represent you accurately.

2. Open AI crawler access

AI citations depend on AI crawlers being able to read your site. GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, and Google-Extended each have their own user agent strings. If any of these is blocked in your robots.txt, either explicitly or via catch-all rules, that model cannot index your current content. The result is that it either omits you from recommendations or relies on stale training data.

Check your robots.txt. Add explicit Allow rules for each AI crawler. Use Faro's robots.txt Analyzer to verify your current status before making changes.

3. Structured data markup

JSON-LD structured data (Organization, Product, FAQPage, SoftwareApplication, and other Schema.org types) tells AI systems what type of entity you are, what you do, and how you relate to other entities. This is the layer that turns web content into machine-parseable facts. AI systems that have structured data to work with produce more accurate, more specific, more confident citations.

At minimum: Organization schema on your homepage, Product schema on product pages, and FAQPage schema on content pages where you answer common questions. At a more advanced level: SameAs connections to your authoritative profiles on LinkedIn, Crunchbase, and industry directories; and SoftwareApplication schema if you offer a software product. Faro's AI Schema Creator generates compliant JSON-LD for any URL.

4. Pricing transparency

AI models asked to compare tools or recommend solutions cannot recommend you if they do not know what you cost. Pricing locked behind a "contact us" wall, hidden in a PDF, or absent from your public site creates an information gap that prevents AI from confidently citing you in commercial queries. Publish clear pricing with plan names, price points, and the key features per tier. If you have custom enterprise pricing, publish a base tier with public pricing and note that larger plans are available.

5. Authentic user-generated content

AI systems weight authentic user testimony heavily because it is verifiable evidence rather than positioning. Reviews on G2, Capterra, and Trustpilot; Reddit discussions in relevant subreddits; forum posts in practitioner communities; and genuine case studies all function as evidence layers that AI can retrieve and use to form confident recommendations. The research finding that Reddit accounts for 54 to 71 percent of all UGC URLs retrieved by deep-research agents indicates that community presence is not optional for competitive AI citation rates.

6. agents.json: the capability manifest

agents.json is a machine-readable file at your root that tells AI agents what they can do with your business. It lists your API endpoints, your capabilities, your authentication methods, and your integration surfaces. AI agents evaluating software or services for a user can parse agents.json to understand whether you are actionable (they can take steps on the user's behalf) or just informational. As agentic workflows become standard, the absence of agents.json is increasingly a disqualifier for category-leader status in AI outputs.

7. Consistent entity information across the web

AI systems build entity understanding by cross-referencing your business name, domain, description, and founding information across multiple sources. If your LinkedIn company page, your Crunchbase profile, your website, and your structured data all describe you consistently, the AI's entity confidence is high. If there are inconsistencies (different descriptions, outdated information, name variations), the AI's confidence is lower, which typically results in more hedged recommendations or outright omission.

Audit your entity footprint: Google your business name and check the Knowledge Graph result if one exists, verify your LinkedIn and Crunchbase profiles are current, and ensure your structured data SameAs links point to your canonical profiles.

8. OKF knowledge bundle

Google's Open Knowledge Format (OKF), released in June 2026, provides a five-layer architecture for structuring all the knowledge an AI agent needs to understand and interact with your business. An OKF bundle at /okf/index.md gives AI agents a navigable, structured knowledge graph covering your product capabilities, use cases, pricing, integration options, and more. Fewer than 0.1% of businesses currently have an OKF bundle, which makes it one of the highest-leverage AI readiness signals available right now.

What the Measurement Looks Like

Running the eight signals above and measuring citation rate are two separate workstreams. You need to do both. The technical implementation (llms.txt, structured data, agents.json, OKF, robots.txt) is one-time work with ongoing maintenance. Measuring whether the implementation is working requires regularly querying AI platforms with your category prompts and tracking whether you appear, in what position, and with what sentiment.

Manual citation checking is time-consuming and non-reproducible. Faro's AEO Citation Monitor runs standardized prompts across Claude, ChatGPT, Perplexity, and Gemini and reports your citation rate, position, and sentiment across all four platforms, with 5-run confidence scoring on the primary category prompt to account for AI non-determinism. It is the fastest way to get a reliable baseline and track improvements over time.

How Long Does It Take to Earn AI Citations?

For live-retrieval systems (Perplexity primarily), well-indexed changes can affect citation rates within days to weeks. For training-data-dependent systems (Claude and ChatGPT when not using web search), the lag is longer and depends on when those models next update their training data. This is why a multi-platform citation monitoring strategy matters: Perplexity tells you quickly whether your current content is working; Claude and ChatGPT tell you whether your established entity authority is strong.

Businesses that have implemented the full eight-signal playbook typically see measurable citation improvements within 30 to 60 days in live-retrieval platforms. The compounding effect builds over 6 to 12 months as training data cycles and community content accumulates.

Common Mistakes That Suppress AI Citations

Several common configurations actively suppress citation rates regardless of how good the rest of the implementation is. Blocking AI crawlers in robots.txt is the most common: many sites have legacy Disallow rules that catch GPTBot and ClaudeBot. A paywall on key content pages prevents AI retrieval entirely. Missing or invalid structured data leaves AI systems guessing. Inconsistent pricing across different pages creates conflicting evidence that reduces AI confidence. No organic community presence means the UGC retrieval layer returns empty.

Identifying which of these applies to your site takes minutes with Faro's AI Readiness Scan, which checks all major citation-determining signals and prioritizes fixes by impact.

Run a citation baseline before making changes

Before implementing any of the signals above, run an AI citation check to establish where you start. Changes take weeks to affect citation rates, so you need the baseline to measure improvement.

In Short

AI citations are earned through eight measurable signals: llms.txt, open crawler access, structured data, pricing transparency, authentic UGC, agents.json, consistent entity information, and an OKF knowledge bundle. Most businesses are missing four or more of these. The compounding commercial effect of earning consistent AI citations is substantial: brands recommended first by AI are 389% more likely to be searched afterward. The implementation work is finite. The returns are ongoing.

Frequently Asked Questions

Does this work for small businesses, or is it only relevant at scale?

It is particularly effective for small businesses, because the competitive bar is low. Most small businesses have no structured data, no llms.txt, and no community presence. A small business that implements the full playbook will out-represent large competitors that have not in AI answers for their specific category and location. The technical implementation is not expensive or engineer-dependent; tools like Faro's AI Schema Creator and OKF Generator handle the heavy lifting.

Which AI platform should I prioritize first?

Perplexity is the fastest to respond to technical improvements because it relies primarily on live retrieval rather than training data. It is also a meaningful driver of discovery for B2B software buyers. Start with Perplexity as your leading indicator, then track Claude and ChatGPT over a longer horizon as their training cycles update.

Is there a risk of being penalized for optimizing for AI citations?

Not for the grounding-layer activities described here. Publishing accurate llms.txt, valid structured data, and authentic content is the same category of activity as well-implemented SEO. The activities that carry risk are those that involve deception: hidden instructions, fabricated reviews, or coordinated inauthentic behavior. None of the eight signals in this playbook are manipulative.

What is the most impactful single change I can make today?

If you have no llms.txt, create one. It is a 30-minute task that immediately gives every AI crawler a clear business context to work with. If you already have llms.txt, verify that GPTBot, ClaudeBot, and PerplexityBot are not blocked in your robots.txt. Either of these changes can produce citation improvements within weeks. Run a scan first to confirm which gap is largest for your site.

Use Faro's AI SEO tools to scan your site against the eight signals above, generate a compliant llms.txt, audit your schema, and benchmark against competitors — all in one place.

Related Reading

← Back to Blog

The Faro platform

Every tool you need to be found, understood, and chosen by AI.

Faro is building the complete infrastructure layer for AI discoverability. Scan first, then fix, monitor, and stay ahead. All from one platform.

AI Readiness ScanLive

Run 30+ checks across 6 categories. Get a score, a grade, and a prioritized fix list in 30 seconds.

Use tool →
llms.txt GeneratorLive

Give AI agents a structured map to your most important content. Download your file in under 60 seconds.

Use tool →
AI Schema CreatorLive

Paste your URL and get the exact JSON-LD markup your site is missing. No developer required.

Use tool →
robots.txt AnalyzerLive

See exactly which AI crawlers you're blocking and why. Get the precise fix lines in under 60 seconds.

Use tool →
Competitor IntelligenceLive

Side-by-side AI readiness scores across up to 3 competitors. See exactly where you lead and where you lag.

Use tool →
Pricing Clarity AuditorLive

Find out if AI agents can actually read and compare your pricing. 6-dimension check in seconds.

Use tool →
OKF GeneratorLive

Build the machine-readable knowledge bundle that tells AI agents exactly what your business does.

Use tool →
Revenue CalculatorLive

Calculate the monthly revenue gap between your current AI readiness and a fully optimised site.

Use tool →

One-Click Fix Engine

Connect your GitHub repo. Faro opens pull requests with every code fix automatically.

New tools ship continuously. Free tier always available.

Browse all tools →