Guide · Fundamentals

How ChatGPT, Perplexity and Google AI Overviews Choose Sources to Cite

Retrieval-augmented generation in plain language, plus what each answer engine publicly documents about its crawlers and index, and where the documentation stops.

By InTheAnswer Editorial · Updated · 8 min read

Answer engines don't pick citations from everything their model has ever read. When ChatGPT search, Perplexity, Microsoft Copilot or Google's AI features answer with sources, they run a search first: they turn the prompt into one or more queries, pull candidate pages from a search index, pick the passages that answer each part, and write a response that attributes those passages to their pages.

That process is called retrieval-augmented generation, or RAG, and it explains most of what marketers notice in AI answers: why the same question cites different sources on different days, why pages that rank well in search show up so often, and why a brand missing from the retrieved pages is missing from the answer.

This guide explains each step in plain language, then sets out what each major engine publicly documents about its crawlers and index. Where the documentation stops, it says so. None of these companies publishes its ranking or citation algorithm, so anything beyond crawler names and stated policies is inference.

Retrieval-augmented generation, without the jargon

A language model on its own answers from what it learned in training. That knowledge is broad but frozen at a cutoff date, and the model can't point to where a particular fact came from. Retrieval fixes both problems by handing the model fresh documents to read at the moment the question is asked.

Think of it as an open-book exam. The model still writes the answer, but it writes from a short stack of pages it just looked up, and it can footnote those pages. The citations you see in an AI answer come from that stack. If your brand isn't on any page in it, the model can only mention you from training memory, with no source to link.

Not every prompt triggers a search. Engines tend to search when a question involves recent events, specific products, prices, comparisons or anything the model is unsure of, and answer from memory for stable general knowledge. Commercial questions such as "which CRM suits a 10-person sales team" usually fall into the first group, which is why retrieval matters so much for brands.

From prompt to cited answer, step by step

The exact pipeline differs between engines and changes often, but the publicly described pieces fit a common shape.

  1. Decide whether to search. The system judges whether the answer needs fresh or specific information from the web.
  2. Rewrite and fan out the query. The prompt becomes one or more short search queries, and a single conversational question can turn into several searches covering its sub-topics.
  3. Retrieve candidates. Those queries go to a search index, which returns ranked pages, much like a normal results page.
  4. Select passages. The system reads the candidates and picks the chunks of text that best answer each part of the question.
  5. Generate with attribution. The model writes the answer from the selected passages and attaches citations to the pages they came from.

Query rewriting and fan-out

This step is better documented than most. Google describes AI Overviews and AI Mode as using a "query fan-out" technique, running multiple related searches across subtopics to build a response (Google Search Central). Microsoft's documentation for its workplace Copilot says Copilot generates a short search query from the user's prompt, sends it to the Bing search service and composes the response from what comes back. In both cases, what gets searched is not the user's exact wording.

For brands, fan-out means a question like "best invoicing software for freelancers in the UK" might set off searches about UK invoicing rules, freelancer software comparisons and individual product reviews. A page that ranks for any one of those can land in the stack.

Retrieval and passage selection

Retrieval is where search ranking matters, because a candidate list drawn from an index reflects that index's judgment of relevance and authority. Passage selection is where writing matters: a page can be retrieved and still go unquoted if no single passage answers the question cleanly. Both points are inferences from how these systems are described and how they behave rather than published rules, but they hold across engines.

Key takeaway: Citations are the visible end of a search process. To be cited, a page that mentions your brand has to rank for at least one of the searches the engine runs and contain a passage worth quoting.

Engine by engine: what's publicly documented

Each company names its crawlers and explains, at least briefly, what they do. That's the most reliable public information available, and it matters because a site's robots.txt decides which of these crawlers may read it.

OpenAI documents its user agents in its crawler documentation. OAI-SearchBot is used to surface websites in ChatGPT's search features, and OpenAI says sites that opt out of it won't appear in ChatGPT search answers. GPTBot collects content that may be used to train OpenAI's models, so blocking it is a training opt-out. ChatGPT-User handles page visits a user triggers inside ChatGPT, and OpenAI notes that robots.txt rules may not apply to those user-initiated requests. OpenAI has also said ChatGPT search draws on third-party search providers alongside its own crawling and partner content.

Perplexity

Perplexity documents PerplexityBot as the crawler that surfaces and links websites in its search results, and says it doesn't crawl for AI model training. Perplexity-User fetches pages when a user's question calls for it, and Perplexity says it generally ignores robots.txt because a person initiated the request.

Google (AI Overviews, AI Mode and Gemini)

Google's AI features in Search use Google's existing index, crawled by Googlebot. Pages must be indexed and eligible for a snippet, and the usual controls apply: noindex keeps a page out, while nosnippet, data-nosnippet and max-snippet limit what can be shown. Google-Extended is a separate robots.txt token that controls whether content may be used to train Gemini models and to ground Gemini's answers, and Google states that it doesn't affect inclusion or ranking in Google Search. AI Overviews get their own treatment in Google AI Overviews and backlinks.

Microsoft (Copilot)

Copilot's web answers are grounded through the Bing search service, so being crawlable by Bingbot and ranking in Bing are what count. Microsoft documents the step where Copilot converts a prompt into a short Bing query, and says its workplace Copilot Chat shows those generated queries alongside its citations.

Anthropic (Claude)

Anthropic documents three user agents in its crawler help article. ClaudeBot collects content that may contribute to model training. Claude-SearchBot analyzes content to improve search quality, and blocking it reduces a site's visibility in Claude's search results. Claude-User retrieves pages when Claude users ask questions that need current information. Beyond the crawler documentation, Anthropic hasn't published how Claude ranks the results it retrieves.

Documented versus inferred

Keep a clear line between what these companies say and what practitioners infer from watching their output.

Publicly documented:

  • Crawler and user-agent names, what each one is for, and how robots.txt applies to it
  • That Google's AI features draw on its search index, require snippet eligibility and respect snippet controls
  • That Google uses query fan-out, and that Copilot sends generated queries to Bing
  • That Google-Extended doesn't affect Google Search

Reasonable inference, not documented:

  • Pages that rank well in the underlying index are more likely to be retrieved
  • Established publications are favored when sources disagree
  • Clear, self-contained passages are quoted more often than vague ones
  • Freshness weighs more heavily for time-sensitive questions
  • Engines prefer several independent sources over a single one

Anyone claiming to know the exact weighting of these factors is extrapolating. That includes every scoring system, InTheAnswer's own AEO score method among them: it estimates citation likelihood from observable signals such as authority, traffic and topic, rather than reading it from the engines.

Key takeaway: Plan around the documented facts first (crawler access, indexation, snippet controls), then use the inferences to break ties between otherwise similar options.

The pipeline points to a short list of practical consequences.

  • Be on pages that rank for the sub-queries. That means your own pages, but also the roundups, comparisons and news articles engines retrieve for your category. Placements on publications that already rank there put your brand into the candidate set. How backlinks help your brand get cited covers that chain end to end.
  • Check crawler access on both ends. Your site and any publication you appear on need to admit the search crawlers: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot. Blocking a training crawler like GPTBot or Google-Extended is a separate decision. The free AEO audit reports which of these crawlers a page's robots.txt allows.
  • Write passages that survive selection. A placement that states what your brand is, who it's for and what it does, in one or two plain sentences, gives the selection step something to quote.
  • Match publications to the engines you care about. Copilot leans on Bing and AI Overviews on Google, so a publication that ranks well in both is worth more than one that ranks in neither. What makes a publication citable shows how to judge that before you buy.

If you'd rather start from a ranked shortlist than a blank page, the niche link finder suggests publications by niche, budget and audience, ordered by how likely answer engines are to cite them.

Frequently asked questions

Why does the same question cite different sources each time?+

Because the search step runs fresh each time. Query rewriting can produce slightly different searches, the index changes as pages are published and updated, and the selection step can pick different passages from similar candidates. Judge citations across many runs over several weeks, not from a single answer.

If a site blocks GPTBot, will it disappear from ChatGPT?+

Not according to OpenAI's documentation. GPTBot is the training crawler; appearing in ChatGPT search answers is governed by OAI-SearchBot. A site can block one and allow the other: blocking GPTBot signals that its content shouldn't be used for training, while blocking OAI-SearchBot keeps it out of search answers.

Do AI answer engines use their own index or Google's?+

It depends on the engine. Google's AI features use Google's index and Copilot uses Bing. OpenAI and Perplexity run their own search crawlers, and OpenAI has said it also uses third-party search providers. Anthropic documents its own search crawler. Because the indexes differ, a page can be cited by one engine and absent from another.

Can I see which searches an engine ran for my prompt?+

Sometimes. Some interfaces list the searches they ran or the sources they consulted, and Microsoft documents that its workplace Copilot Chat shows the generated web queries with its citations. Most engines show only the final sources, so you usually have to infer the sub-queries from the pages that were cited.

Keep reading