Scout7 logo

Scout7

framework

How ChatGPT Decides Who to Cite: The 8 Steps Between a Question and a Citation

September 2, 2026 · 10 min read · Scout7

A plain-English walkthrough of the 8 steps between a user question and an AI citation, plus how to rewrite content into citation-ready passages.

How ChatGPT Decides Who to Cite: The 8 Steps Between a Question and a Citation

Eight stages sit between someone's question and a link appearing in the answer. Here is what happens at each one — and where your page quietly drops out.

Introduction

You rank well, publish often, and still vanish inside ChatGPT. That happens because AI search optimization does not follow the same path as classic rankings: answer engines often choose citations by extracting, comparing, and selecting passages, not by rewarding the highest-ranked page.

That shift is already material. According to Salesforce’s 2026 marketing report, nearly 9 in 10 marketers (88%) are already optimizing for AI-driven answers, while McKinsey found half of U.S. consumers intentionally use AI-powered search.

Key takeaways:

  • AI citation and search ranking are now parallel jobs
  • Pages often fail after extraction into passages
  • Passage-level optimization for AI search is structural, not mystical
  • The fix is self-contained, answer-first paragraphs

Steps 1-3: Fan-Out, Index, and Fetch Budget

Steps 1-3: Fan-Out, Index, and Fetch Budget

So what happens first? Before any citation appears, the system expands one question into several hidden sub-questions, searches its own candidate set, and opens only a few pages.

That means your page can lose before it is properly read.

  • Fan-out rewrites one prompt into related intents and angles
  • Index searches a candidate corpus, not necessarily a human-visible SERP
  • Fetch budget limits how many pages get opened for reading
  • Selection pressure starts before your page text is even parsed

Google’s filed patents describe parts of this likely pattern. WO2024064249A1 discusses query fan-out, and US20240362093A1 describes citation selection as independent of document ranking, but these are filed documents, not proof of production behavior.

Third-party reverse-engineering suggests some systems may open only roughly five to ten pages, which is useful as a working model but not platform-confirmed. That narrow reading set matters because Similarweb reported more than 1.1 billion AI referral visits in June 2025, up 357% year over year.

How sure are we? Three grades of evidence

How sure are we? Three grades of evidence

And that leads to the next problem: too much writing about AEO treats guesswork and evidence as the same thing. A better answer engine optimization strategy starts by labeling confidence.

Use three evidence grades, and keep them separate.

  • Filed patents show likely retrieval patterns, not shipped behavior
  • Measured studies show observed outcomes in the market
  • Reverse-engineering offers practical models, not confirmations
  • Confidence rises when different evidence types point the same way

Measured studies are the strongest evidence for outcomes. Ahrefs found that when an AI Overview appears, the top result’s clickthrough rate is 58% lower on average. In separate analysis, Ahrefs also observed that a #1 ranking page has roughly a one-in-three chance of being cited, and that 47% of citations come from pages below position five. Those are two different measurements of two different things — a click effect and a citation effect — and they should be read separately rather than stacked into one claim.

Behavioral data supports the same shift. Research reported by Search Engine Land found about 68% of U.S. Google searches (68.01%) in early 2026 ended without a click. That is why AI Search matters: AI search is a method of information retrieval that utilizes large language models and natural language processing to synthesize answers from multiple sources rather than providing a list of links.

Steps 4-6: Parse, Chunking, and Rerank

Steps 4-6: Parse, Chunking, and Rerank

Once a page gets opened, it stops being a page. The system often strips rendering down to text, cuts that text into chunks, and ranks those passages against passages from other sites.

This is where most structurally sound SEO pages fail as citation candidates.

  • Parse removes design, layout, and sometimes JS-dependent meaning
  • Chunking breaks one article into smaller retrieval units
  • Rerank compares your paragraph against another site’s paragraph
  • Winning unit is often the passage, not the page

Most content teams write continuous arguments. One paragraph depends on the previous one, so extraction breaks the logic chain.

The pattern is structural, not mystical: pages built like essays collapse when a system strips them to text and judges each paragraph on its own.

A passage that cannot stand alone usually cannot be cited.

That matters for scaling B2B organic growth because unprepared brands may lose 20% to 50% of traditional search traffic, according to McKinsey, while Adobe found more than three-quarters of organizations (76%) report faster, higher-volume content production from generative AI. Speed is rising; structure now decides whether that output survives rerank.

Steps 7-8: Word Budget and Attribution

Even then, winning passages still get cut. The model has limited room, so only a few supporting passages survive into the final answer.

That is where word budget and attribution decide who gets cited.

  • Word budget limits how many claims and sources fit
  • Compression favors concise but complete passages
  • Attribution follows supporting evidence, not page prestige
  • Parallel discipline means SEO success does not guarantee citation success

This is why Scout7 takes an anti-hype position. AI citation is not replacing search; it is creating a second surface with different rules.

Organic SEO still matters because it drives discovery, authority, and organic traffic. But citation works differently: Ahrefs found the top-ranked page has only roughly a one-in-three chance of being cited, and 47% of citations come from pages ranking below fifth place.

That is a clean definition of parallel work. Search ranking and AI citation now run side by side, and B2B teams need both.

Trust also continues after the citation. Forrester found, from a survey of nearly 18,000 buyers, that generative AI is reshaping discovery, but buyers still verify through human networks.

The passage-first audit

So what should you actually do on Monday? Stop auditing pages like essays and start auditing paragraphs like standalone answers.

That is the practical core of passage-level optimization for AI search.

  • Check standalone meaning: can this paragraph survive alone?
  • Lead with the answer: put the direct claim first
  • Anchor context: replace “this” and “it” with named subjects
  • Increase entity density: include product, concept, audience, and action
  • Keep support nearby: define, explain, then prove in one passage

Before: “This is why the approach above matters so much for modern teams.”

After: “Passage-level optimization helps B2B teams earn AI citations because each paragraph answers one sub-question clearly enough to be quoted without surrounding context.”

This is where AEO matters. AEO, or answer engine optimization, is the practice of improving content so answer engines can retrieve, trust, and cite it inside generated responses.

Faster drafting does not help here. Writing more paragraphs that each depend on the one above them just produces more text that cannot be quoted.

When the framework fails

When the framework fails

Still, not every page should be flattened into citation bait. Some content is meant to be navigated, not extracted.

That is where good judgment matters more than dogma.

  • Technical documentation often depends on sequence and prerequisites
  • Interactive tools lose meaning when parse removes the interface
  • Comparators and calculators work better with human navigation
  • Step-dependent tutorials may need page context to stay accurate
  • Selective optimization beats forcing every asset into one mold

If a page only works when readers move through it step by step, extraction may damage it. In those cases, optimize for navigation first and citation second.

Forcing every asset to do every job is how you break the pages that were already working.

A second discipline to earn

A second discipline to earn

The opening mistake was assuming a citation works like a ranking, just faster. It does not. AI systems fan out the question, search a private candidate set, open only a handful of pages, strip them to text, break them into passages, rerank those passages, compress the final answer into a tight word budget, and attribute the supporting claims that survive.

Key takeaways:

  • AI citations are earned passage by passage
  • Organic SEO and AEO now run in parallel
  • Self-contained paragraphs survive extraction best
  • Start with one page and five rewritten passages

That should change how you think about organic traffic. Organic SEO still earns visibility, branded trust, and discovery, especially as AI interfaces and classic search coexist. But answer engine optimization strategy adds a second job: making sure your best ideas still make sense after extraction.

The practical path is narrow. Keep your existing SEO program running, then layer passage-first rewrites into the pages that answer discrete questions best.

If you want a concrete next step, pick one high-intent article, identify five weak paragraphs, and rewrite each one as a complete answer to a likely sub-question. If those passages start earning citations, you have found your pattern — repeat it across the rest of the library.

Frequently asked questions about AI search optimization

How does ChatGPT decide which sources to cite?

It runs a retrieval pipeline, not a popularity contest. The question is rewritten into sub-questions, a search returns candidates, only a handful of those pages are actually opened, and the opened pages are cut into passages that compete against passages from other sites. The few passages that survive into the final answer are the ones that get the link.

Does AI cite pages or paragraphs?

Paragraphs. Once a page is opened and parsed, it is split into passages and each passage is judged on its own merits. That is why a page can be excellent as a whole and still never get quoted: if no single paragraph on it stands up without the paragraphs around it, there is nothing clean to lift.

Why doesn't ranking #1 get me cited?

Because ranking and citation are scored by different systems. Ahrefs found that a #1 ranking page has roughly a one-in-three chance of being cited in an AI Overview, and that 47% of citations come from pages ranking below position five. A strong ranking helps you get found; it does not buy you a quote.

How many pages does an AI answer engine actually read?

Far fewer than it finds. Third-party reverse-engineering of ChatGPT's retrieval stack suggests only roughly five to ten pages get opened for a given answer. That figure is inferred from observed behaviour rather than confirmed by the platform, but the direction is consistent: most of what the search returns is never read at all.

What is query fan-out?

Query fan-out is the step where one question is expanded into several machine-written sub-questions before any search runs. Google's filed patent WO2024064249A1 describes generating these synthetic queries across different intents, including comparative, exploratory and entity-based variants. You never see them, but they are what your page is really being matched against.

Why do I get different sources every time I ask the same question?

Because these systems are not deterministic. The same prompt can return a different set of cited sources from one run to the next, so checking once is taking a single sample rather than making a measurement. To know whether you are actually cited, run the same question several times and across more than one engine.

References