How LLMs Prioritise Information When Answering About Brands
When someone asks ChatGPT, Perplexity, Claude or Google what your brand is, does or costs, the answer is not one decision. It is four, made in sequence, by different systems, with different owners. Most GEO advice collapses them into one and sells a tactic measured at the last stage as a result at the first. This piece separates them. For each stage: what the model does, what a publisher controls, one primary source, and what does not work — with the study that killed it.
How this was verified. Every claim below comes from a primary document an agent fetched and quoted. The evidence base behind it holds 140 claims: 100 confirmed, 34 weakened because the source said less than the claim, 5 misquoted and corrected. Zero citations were unreachable. Where two sources disagree, both are named.
The governing finding
A critical survey of 45 GEO studies (arXiv:2607.14035, July 2026) reached one sentence that reframes the category: "already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability." Nearly every published GEO effect size was measured on documents already inside the model's context window. Those are stage-four findings sold as stage-two outcomes. Keep that sentence in view through everything below.
Stage 1 — Parametric: what the model already believes
What the model does. With no web access, a model answers from weights frozen at a training cutoff. OpenAI documents the GPT-5.6 family at a "Feb 16, 2026 knowledge cutoff", and the Instant path at "Aug 31, 2025" — the same product, two knowledge states nearly a year apart, chosen by routing you cannot see (developers.openai.com/api/docs/models/gpt-5.6-luna).
What you control. Nothing, this quarter. A brand founded after the cutoff cannot appear in a retrieval-off answer regardless of what it publishes. The only channel into a future model is the training crawl — GPTBot, ClaudeBot, Google-Extended — and OpenAI's bot page describes GPTBot as used "to crawl content that may be used in training" and nothing more (developers.openai.com/api/docs/bots). The effect of allowing it is unmeasured by anyone, including the vendors.
What does not work. "Allow GPTBot so ChatGPT can find you." GPTBot governs training only; OAI-SearchBot governs search. OpenAI: "Each setting is independent of the others." The two are sold as one lever. They are not.
Stage 2 — Access: whether a crawler may fetch the page at all
What the model does. Before any answer-time search, an index has to hold your page. Three gates decide that, and every one of them is binary.
First, robots.txt, per agent. OpenAI: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links" (same page). Meta publishes the only positive first-party statement in the entire corpus: "Allowing Meta-WebIndexer in your robots.txt file helps us cite and link to your content in Meta AI's responses" (developers.facebook.com/docs/sharing/webmasters/web-crawlers).
Second, the network layer. OpenAI's help centre states eligibility as a conjunction: "allow OAI-Searchbot to crawl the site and confirm that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses" (help.openai.com/en/articles/9237897). Robots.txt says allow, the WAF says 403, and no dashboard tells you. From 2026-09-15 Cloudflare sets Training and Agent crawlers to blocked by default on ad-displaying pages for new domains, new sites on existing accounts and all existing free-tier customers.
Third, rendering. Vercel's crawler study: "while ChatGPT and Claude crawlers do fetch JavaScript files, they don't execute them. They can't read client-side rendered content" (vercel.com/blog/the-rise-of-the-ai-crawler). Re-confirmed in June 2026 for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Bytespider. Googlebot renders; the others do not. A price that appears only after a script runs does not exist to them.
What you control. All of it. This is the only stage where the publisher owns every variable, the mechanism is documented by the operators themselves, and the check is deterministic. For Google specifically, the requirement is stated once: "a page must be indexed and eligible to be shown in Google Search with a snippet ... There are no additional technical requirements" (developers.google.com/search/docs/appearance/ai-features). The corollary is that suppression controls — nosnippet, max-snippet:0, data-nosnippet — remove you. Bing documents the same: content so marked "will be excluded from snippets and AI summaries" (blogs.bing.com, October 2025).
What does not work. Additive markup. Google's AI guidance (updated 2026-07-10): "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add" and, on llms.txt and similar files, "Google Search ignores them" (developers.google.com/search/docs/fundamentals/ai-optimization-guide). Ahrefs' matched-control test of 1,885 pages found no lift on any platform and a significant 4.6% drop in AI Overview citations after schema was added (ahrefs.com/blog/schema-ai-citations). There is no additive counterpart to any of the subtractive controls above. You can only be ineligible.
Stage 3 — Retrieval: whether the page enters the context for this query
What the model does. It does not run your customer's prompt. OpenAI: ChatGPT search "typically rewrites your query into one or more targeted queries". Google may fan out across subtopics. The Gemini API generates its own queries. The retrieval input is unobservable from outside — except in Bing Webmaster Tools, which is the only surface anywhere that exposes the grounding queries an AI actually used.
Then a candidate list is built and ordered. Position in that list is the strongest effect in the experimental literature. C-SEO Bench (arXiv:2506.11097, NeurIPS 2025): "traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective" than any content edit. The 252,000-trial factorial (arXiv:2605.25517) agrees: "topical relevance and list position are the biggest drivers of being cited first."
What you control. Topical relevance — the one substantive factor three independent studies agree on. Not keyword coverage: Google says you need not "worry that you don't have enough 'long-tail' keywords", and the original GEO paper found keyword stuffing lowered its own visibility metric. Relevance means the page answers the question the query asks. You do not control position; nobody outside the retrieval operator does.
What does not work. "Rank in Google's top 10 and AI Overview citations follow." Ahrefs measured 863,000 SERPs (March 2026): 37.9% of URLs cited in AI Overviews also appeared in the first ten blocks; roughly 31% came from positions 11-100 and 31% from beyond 100 (ahrefs.com/blog/ai-overview-citations-top-10). BrightEdge reports 16.7% with a different method and no published sample size. The two disagree by twenty points, and neither should be read as a trend — the earlier Ahrefs figure of ~76% came from a different parser.
Stage 4 — Generation: which retrieved documents get used and cited
What the model does. Given a context of candidates, it writes an answer and emits citations. The two are not the same act. OpenAI exposes retrieved sources and cited URLs as separate fields, with the retrieved set larger. Mechanistic work (arXiv:2606.28358) repaired over 90% of missed citations by amplifying specific attention heads without changing the sources — citation emission is partly decoupled from what influenced the answer. Anthropic's search tool cites text spans of up to 150 characters; Perplexity cites whole sources by integer index. The units are not commensurable, which is why no single cross-platform "AI visibility score" has a denominator.
What you control. Whether the facts on the retrieved page are correct and agree with each other. This is the stage where "ChatGPT gets our pricing wrong" is actually caused, and the plausible mechanism is your own site stating a price two ways. Fixing that is fully publisher-controlled and maps directly onto what customers complain about.
What does not work. Rewriting for citation. C-SEO Bench, on the whole family of GEO content tactics: "most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking." The factorial: "formatting-only edits have little impact." The survey: "citation-oriented rewrites can impair retrieval." The widely quoted "up to 40% more visibility" came from an engine the authors built, with sources pre-loaded into context, measured by a metric that counts the words the intervention added. It never touched organic discoverability.
What is cross-cutting: brand prominence
Across all four stages, one association dominates every correlational study and nothing controls for it. In the largest citation dataset, "the top 20 most frequent news sources account for 67.3% of all citations for OpenAI models" (arXiv:2507.05301). Semrush's 36 brands visible on every surface every month are YouTube, Google, Reddit, Amazon, Facebook, Apple, Walmart, Disney, Nintendo. That is brand size wearing a GEO label. "Domain authority" is a third-party regression built from backlinks — a proxy for the confounder, not a control for it. Google publishes no site-level authority score and, on E-E-A-T, says it "isn't a specific ranking factor" (developers.google.com/search/docs/fundamentals/creating-helpful-content).
What this means for your brand: five checks
- Fetch your key pages with JavaScript disabled. If the prices, product names and claims are not in the raw HTML, no non-Google AI crawler has ever read them. Server-render them.
- Resolve robots.txt per agent, then test the CDN. Allow OAI-SearchBot, PerplexityBot, Claude-SearchBot and Meta-WebIndexer. Then request a page as each one and confirm a 200, not a 403 or a challenge page. If you are on Cloudflare's free tier, check AI Crawl Control before 2026-09-15.
- Grep for accidental suppression.
nosnippet,max-snippet:0,data-nosnippet,noindex— and the Search Console generative-AI inclusion control, which is opt-out and inherits from parent properties. - Find where your own site contradicts itself. Pricing table, FAQ, footer, schema, About page. One version of every fact. If you have JSON-LD, every fact in it must also be in the visible text.
- Know which answers you cannot reach. A retrieval-off answer reflects a cutoff of February 2026 or August 2025, and nothing you ship changes it until the next model. Measure with retrieval state disclosed, or you are averaging an unobserved mixture.
Everything on that list is subtractive, mechanical and checkable. None of it is what the category sells. It is also the only part of the category the evidence supports. Run a free audit to see which of the five your site fails today, or read why schema markup is not the lever it was sold as.
Ready to see what AI says about your brand?
Run your first audit free. Get visibility scores, detect hallucinations, and get specific fixes.
Start Free AuditRelated Articles
Schema Markup for AI: What the Evidence Actually Shows
Google says no special markup is needed for AI search. The one controlled test found a 4.6% drop. Here is why the industry sells schema anyway, what it is still good for, and the one AI use that holds up: parity with your visible text.
What Is GEO? The Complete Guide for 2026
Generative Engine Optimization (GEO) is how you get your brand into AI-generated answers. This guide covers everything: what GEO is, why it matters, the 7 ranking factors, and how to start.
Google AI Overviews: What Marketers Need to Know
Google AI Overviews appear in 70-80% of search queries. Here is how they work, what content they pull from, and how to optimize for them.
