Skip to content
Getting cited by AI · Source selection
Summarise with

How do AI search engines choose sources to cite?

Each platform's ranking system is proprietary, but the public documentation agrees on the shape: retrieve candidate pages, expand the question into related searches, select relevant supporting sources, generate an answer with links. Here is what that means for your pages.

Updated Jul 2026First published Jul 2026

This is the source-selection stage of the getting cited by AI pillar. It assumes crawler access is already working.

What happens between a buyer question and an AI citation?

The documented sequence has four steps, whatever the platform. First, the system needs candidate content: an index built by its search crawler, or a live retrieval at answer time. Second, it interprets the user's question, often expanding it into multiple related searches. Third, it selects the pages that best support the specific claims the answer needs. Fourth, it generates the answer and attaches links to the supporting pages.

Google describes its AI features this way: pages must be indexed and snippet-eligible through normal Search, per its AI features documentation, and its generative AI guidance frames optimization as standard content quality work. OpenAI's publisher guidance describes OAI-SearchBot surfacing and linking sites in ChatGPT search results. The internals differ per engine; the shape is shared.

What is query fan-out?

Query fan-out is the expansion of one user question into several searches. Google's AI Mode documentation describes issuing multiple related queries on the user's behalf and assembling the response from what they return. A question like "best accounting software for a small construction business" may fan out into searches about accounting software categories, construction-specific features, pricing and reviews.

This matters because your page competes at the sub-question level, not the headline-question level. A page that thoroughly answers one of the fanned-out searches can be cited in an answer whose headline question you never targeted. It also means citation opportunities are wider than your keyword list suggests.

Why are different sources cited for the same question?

Because answers are generated fresh each time, from retrieval that varies. Ask the same question twice and the fan-out queries, the retrieved candidates and the generated text can all differ. Region, language, account context and time all move the result.

The practical consequence: a single AI answer is an anecdote, not a measurement. Judging visibility requires the same questions asked repeatedly across engines and regions, which is why the measurement article in this cluster treats one-off checks as unreliable.

What makes a page relevant to a grounding query?

The page has to contain the answer to the specific retrieved question, stated clearly enough to support a sentence in the generated response. In practice that favors pages where the direct answer is present in text (not implied, not in an image), where the page covers one intent rather than five, and where claims are concrete enough to be quoted or paraphrased safely.

None of this is exotic: it is the same relevance judgment search has always made, applied at the level of passages that can support an answer. The citable content article covers how to structure for it.

Does a citation mean the page influenced the answer?

Not necessarily, and the distinction matters for measurement. A link in an AI answer means the page was selected as a supporting source for some part of the response. It does not tell you how much of the answer's substance came from that page, and a brand can be mentioned in an answer without any of its pages being linked.

That is why mention, recommendation and citation should be tracked as separate outcomes. An answer can mention you without citing you, cite you without recommending you, or recommend you based entirely on third-party sources.

Why are competitors cited instead of your business?

Usually for one of four reasons, each with a different fix:

  1. They are retrievable for the sub-questions and you are not. Their pages answer the fanned-out searches; yours target different phrasing or do not exist.
  2. Their pages state the answer more extractably. Same information, but theirs is stated directly where a model can lift it.
  3. The independent sources talk about them. Reviews, directories and community threads retrieved alongside the brand pages describe the competitor, so the answer does too.
  4. Their business entity is clearer. The engine can confidently say who they are and what category they belong to; yours is ambiguous.

Diagnosing which reason applies to each lost question is most of the work. Guessing wrong sends you into a quarter of content production when the actual gap was a review profile, or vice versa.

How should marketers analyse cited sources?

Work backwards from real answers, question by question. For each buyer question you care about, record which sources every engine cited, classify each source (your site, a competitor's site, a review platform, a directory, a community thread, an editorial page), and note whether your brand appears on any of the cited pages.

Patterns emerge quickly: certain domains recur across many questions in your category, and those recurring cited sources are where presence is worth earning. This beats generic authority-building because it names the exact pages doing the deciding. The third-party proof article covers how to act on those source lists legitimately.

How CiteAgentic maps prompts to cited URLs and competitors

Product example. CiteAgentic tracks your buyer questions across engines and regions and keeps the evidence at question level: for each tracked question it records which brands were mentioned or recommended and which URLs were cited, classifies each cited source, and shows where competitors appear in sources you are absent from. Recommendations are then tied to the specific questions and cited sources behind them, so "earn presence on this source" always names the real source and the questions it decides, rather than generic advice.

See which sources decide your questions

Track your real buyer questions and see exactly which pages each engine cites, and where competitors appear instead of you. Explore AI Visibility →

Frequently asked questions

Is there a ranking factor list for AI citations?
No credible one. The platforms do not publish citation algorithms, and their public guidance points to ordinary indexing, relevance and content quality. Anyone selling a definitive factor list is extrapolating beyond what the vendors have said.
Do AI engines prefer big, high-authority sites?
Recurring cited sources in a category often include large review platforms and editorial sites, but engines also cite small, specific pages that directly answer a fanned-out query. Relevance to the specific sub-question is the lever you control.
Can I see the fan-out queries an engine used?
Not directly. You can infer them by studying which sources were cited for a question and what those pages answer. Repeated observation across many questions reveals the sub-topics an engine associates with your category.
Keep reading
Published by CiteAgentic, the AI visibility platform. We run these audits for a living; everything above is measured on real scans, not opinion.
Reviewed by the CiteAgentic research team · citeagentic.com