TL;DR

Ranking helps Google find your page, but it doesn't mean AI Mode will use or cite it. Google's patents show that it may prefer content that adds something new, avoid relying too heavily on one website, search related questions to find more sources, and use separate sources to check claims before showing a citation. For marketers, focus on original information, clear sections that answer specific questions, and claims backed by evidence.

AI search advice assumes that the pages Google ranks highest are the pages AI Mode is most likely to use. But current data and my findings don't support that.

For example, seoClarity analyzed 12,011 Google AI Mode citations and found that in 81% of the queries, AI Mode used at least one source from the top 20 organic results. But only 19% of all citations came from those top 20 results. Even the URL ranking first was cited only 25% of the time (source).

Semrush had a similar finding in a 5,000 query analysis where the most common AI Mode response format had a 51% domain overlap with Google's top 10 results, but only 32% URL overlap (source).

Ranking influences which pages Google can retrieve. It doesn't impact what's included in the final source list.

I wanted to know what could be happening between Google finding the page and then citing or using the source in the answer.

So I went through Google's patents, and found several extra steps that can sit between those two points.

1. The Most Relevant Source Doesn't Always Get Cited

One common assumption about AI search is that the most relevant source is the one Google uses. If one page matches the question better than another, it seems logical that it should win the citation.

When I went through Google's March 2026 generative-search patent, I found that the patent says: "some portions are included in the diversity set that are less relevant to the query but represent diversity of content from other more relevant portions." (source)

Basically, the system described in the patent scores individual passages (sections, paragraphs or sentences from a page) for relevance. After selecting one strong passage, it can compare later passages with the information already chosen.

If another passage is highly relevant but says substantially the same thing, the system can skip it. A less relevant passage can be selected instead if it adds different information.

A page can answer the query perfectly and add nothing Google doesn't already have. Another page can be a weaker match overall but contain one fact, example, comparison or piece of evidence that fills a gap in the answer.

Google's newer Search guidance on generative AI points in the same direction. It recommends unique, non-commodity information rather than pages that recycle what is already available elsewhere (source).

This doesn't prove AI Mode runs this exact redundancy filter on every query, and it doesn't mean Google has a published 'uniqueness ranking factor'.

The patent shows that Google has designed a generative-search system where a highly relevant passage can be dropped because it repeats information already selected.

So when you're using ChatGPT to generate the same definitions and text as everyone else - think about what that means for citation probability in Google AI Mode.

What This Means for Marketers

  • Don't use organic position as a proxy for citation potential. A page can rank and add nothing new to the source set.
  • When you refresh content, compare the claims, examples and evidence against other pages covering the topic. Don't just copy their structure and rewrite the same points.
  • Add information that changes the answer: first-party data, original examples, product detail, expert input, specific comparisons, evidence or experience that isn't already repeated across the SERP.
  • Review content at passage level. One strong section with distinct evidence may be more relevant to AI source selection than making the entire page longer.

2. Google May Limit How Much It Uses From One Domain

Ranking several URLs for the same topic looks like an obvious advantage in traditional Search. In a generated answer, Google may have a reason to stop one site taking too much of the source list.

One March 2026 patent describes limits at both domain and page level. It states: "the domain constraint may limit the number of portions added to the diversity set 235 for resources from the same domain." (source)

One example allows only one selected passage from the same domain. Another allows two. The patent also describes limits on how many passages can come from the same individual resource.

Those numbers are examples, not published AI Mode rules. The patent explicitly includes domain diversity in generative-search retrieval.

A site can have multiple pages that are relevant enough to be retrieved without all of them being carried into the material used for the answer. If Google already has enough from one domain, another source may add more insight.

Google's documentation supports this. It says query fan-out can identify additional supporting pages and produce a wider, more diverse set of helpful links than classic web search (source).

Semrush saw that in its AI Mode study. In 92% of the responses it analyzed, the most common sidebar format contained about seven unique domains. That format had only 51% domain overlap and 32% URL overlap with the traditional top 10 (source).

That doesn't prove a one or two passage cap is active in AI Mode. It shows why multiple organic rankings from one domain shouldn't be assumed as multiple guaranteed slots in an AI answer.

What This Means for Marketers

  • Don't assume ranking multiple similar URLs gives you more chances to appear in the same AI answer. Google has patented mechanisms that can restrict repeated evidence from one domain.
  • Prioritize the strongest page for each topic or use case instead of producing multiple pages that make the same points in different ways.
  • Look at source diversity when you analyze AI results. If seven domains appear, the question is 'what did each source contribute that the others didn't?'
  • Treat multiple URLs from your own site as competing for a limited role in the answer. Make the purpose and evidence on each page different.

3. Query Fan-Out Can Pull Different Sources for Different Parts of the Answer

Query fan-out is where Google, and other LLMs, take one search and turn it into several related searches.

Google's current documentation confirms that basic behaviour. AI Mode can search across subtopics and data sources to find more supporting information.

The March 2026 patent describes what can happen to those results after the searches run.

The main query can build its own set of relevant, varied passages. Each related query can build a separate set. Google can then select across those groups to create what the patent calls a 'completeness set' for the answer.

The main query can carry more weight. Related queries can add information that the original search didn't find. Repeated information can be removed again before the final material is sent to the model.

User query

  1. Related searches and subquestions
  2. Each search finds its own sources
  3. Repeated information is removed
  4. The remaining material is combined
  5. The answer is generated

So fan-out isn't just keyword expansion.

A source doesn't need to rank for the exact wording the user typed if it answers one of the related questions Google searches in the background.

  • A pricing page could enter through a cost-related search.
  • A comparison page could enter through an alternatives query.
  • A forum or video could enter because one part of the question is better served by a different source type.

Google's current AI Mode documentation confirms that query fan-out runs multiple related searches across subtopics and data sources. The March 2026 patent describes building a separate set of relevant, diverse passages for the main query and each related query, then selecting across those sets for the final answer (source).

The related queries don't have to be synonyms. The patent says they can be chosen for both relevance to the original query and diversity from it. That means one user question can create several different routes into the answer, with each route bringing back different evidence before duplicate information is filtered out (source).

What This Means for Marketers

  • Stop researching only the exact prompt. Map the subquestions Google could fan out into: cost, comparisons, alternatives, use cases, objections, definitions, examples and supporting facts.
  • Build pages and sections that answer those subquestions directly. A page can enter the answer through one part of the fan-out even if it doesn't rank for the original wording.
  • Track citations against clusters of related prompts, not one head query. The original SERP won't show the full set of routes Google can use to find sources.
  • Consider the source type each subquestion needs. Some parts of an answer may pull from forums, video, places, product data or other specialist sources rather than standard web pages.

4. More Complex Queries May Pull From More Sources

AI search tracking tools often put different prompts into one report: short questions, product comparisons, research queries and multi-part prompts.

Those queries might not be pulling from the same amount of source material.

Google's March 2026 patent describes assigning a complexity score to the query and changing retrieval based on that complexity. More complex questions can use more related queries and larger sets of passages (source).

One example in the patent uses three related queries for a less complex question and five for a more complex one. Those are examples, not live AI Mode thresholds.

A simple factual query may need a small amount of supporting information. A multi-part comparison can require more subquestions, searches, and sources before the model has enough material to answer.

The patent doesn't tell us the live thresholds or how many searches AI Mode runs for a given prompt. It shows that Google has designed query complexity as something that can change retrieval depth.

Ahrefs' analysis of query fan-out found AI Mode typically runs 5 to 11 searches for a single query. That doesn't prove complexity controls the number of searches, but it supports the broader point that AI Mode retrieval isn't a fixed one-search process (source).

What This Means for Marketers

  • Segment AI-search reporting by query type. Don't put short factual prompts, comparisons and multi-part research questions into one citation-rate number and assume they're comparable.
  • Expect detailed questions to expose your content to a larger source pool. A citation rate on a small question doesn't tell you how the same page will perform on a complex one.
  • Create content that can answer a specific part of a larger question, not only the whole query. Detailed comparisons, pricing, use cases, and supporting facts can each match different subquestions.
  • When citation performance changes, check whether the prompt set changed in complexity before assuming the content itself changed in performance.

5. The Citation You See May Not Be the Source That Generated the Answer

The common misconception at the moment is 'Google cited this URL, so this must be where the answer came from.'

Google has patented a process where that isn't necessarily true.

Its Factuality of generated responses patent describes retrieving source material to generate an answer, then checking factual statements after that answer exists. The system can identify a factual span in the generated response and search for a document that supports it (source).

To clarify, 'document' just means a source Google can retrieve, such as a web page or another indexed resource.

That supporting document can come from the original source set, or the system can submit the generated statement itself as a new Search query and find another document.

The flow can look like this:

Original query → sources A, B and C → answer generated

Then:

Generated claim → new search → source D verifies the claim → source D is attached as the supporting link

Source D can therefore appear as the citation without being the source that originally supplied the information used during generation.

Google's Generative summaries for search results patent family describes the same separation in a different way (source).

So, Google can attach a different source to verify part of the answer, even if that source wasn't used to generate it.

Ahrefs found an interesting example of how different visible source lists can support similar outputs. Across 730,000 paired AI Mode and AI Overview responses, the answers had 86% semantic similarity but only 13.7% citation overlap (source).

That doesn't prove Google selected those citations after generation, but it shows that similar answers can be paired with very different visible sources.

What This Means for Marketers

  • Use citation tracking to measure which URLs Google displays, not as a complete record of every source used to build the answer.
  • Separate three questions in your reporting: Was the brand cited? Did the cited page support the information? Was that page necessarily the source used during generation? Those aren't the same thing.
  • Audit citations. A URL appearing beside an answer tells you less than whether the page actually supports the sentence Google attached it to.
  • Don't judge content only by whether it appears as a visible citation. A page could have been retrieved or used earlier in the process without being the link shown to the user.

6. More Trusted Sources May Increase Confidence

AI citation analysis often focuses on which single domain got the citation.

Google has also patented a system where several sources can support the same generated claim. In Generative summaries for search results, Google describes calculating confidence for individual parts of a generated answer based partly on the documents that verify them (source).

The patent gives an example; confidence can be higher when four search-result sources verify part of the answer than when only one does. It can rise again when those sources are considered more trustworthy.

Four isn't a published threshold. It's an example used to explain the mechanism.

A May 2026 academic study ran 55,393 trending queries and broke Google AI Overview responses into 98,020 individual claims. It found that about 11% of those claims weren't supported by the cited pages (source).

Oumi found a similar gap in a different test: 91% of its sampled AI Overviews contained the correct answer, while 67% of individual claims were supported by the cited sources (source).

Those studies are about AI Overviews, not AI Mode, and they show why 'Google has a verification mechanism' and 'every citation perfectly supports the answer' aren't the same.

What This Means for Marketers

  • Back important numbers, comparisons, and product info with evidence that can be checked independently. If Google may verify a claim against multiple sources, unsupported statements give it less to work with.
  • Build original evidence where you can: first-party data, experiments, benchmarks, customer data, expert input and documented examples give Google something concrete to verify or cite.
  • Don't read one visible citation as the only source behind a claim. Several documents may be supporting the same answer even when only one is shown.
  • When reviewing AI citations, check whether the cited page actually supports the specific claim. Citation presence on its own isn't enough.

What This Changes for Marketers

Google says its generative Search features are rooted in the same core Search ranking and quality systems used to retrieve web pages (source).

The patents don't contradict that. They describe more decisions that can happen after retrieval.

  • A relevant passage can be dropped because it repeats information already selected.
  • Evidence from one domain can be limited.
  • Related searches can bring in sources that never ranked for the original query.
  • More complex queries can gather more material.
  • A displayed citation can be found after the answer is generated.
  • Several trusted sources can support confidence in the same claim.
  • Treating AI citations as another set of rankings misses those extra stages.

Organic rankings can tell you whether Google is likely to find and retrieve a page. Citation tracking tells you which sources Google chose to show. Neither reveals every source or decision between those two stages.

For marketers working on AI search, the better question isn't only "where do we rank?" It is also "what information do we add that Google may need when it builds and verifies the answer?"

FAQ: How Google AI Mode Selects Sources

Does Google AI Mode use the same ranking systems as normal Google Search?

Google says its generative Search features are rooted in its core Search ranking and quality systems, and normal SEO best practices remain relevant. Google hasn't said the final AI citation list follows the organic rankings in order. Its patents describe additional possible stages involving passage selection, source diversity, related searches and verification (source).

Does ranking number one guarantee a Google AI Mode citation?

No. In seoClarity's study of 1,000 US transactional queries, the position-one organic URL appeared in AI Mode 25% of the time. Higher rankings increased the chance of being cited in that dataset, but they didn't guarantee inclusion (source).

Does Google AI Mode use query fan-out?

Yes. Google's own documentation says AI Mode and AI Overviews can use query fan-out, issuing multiple related searches across subtopics and data sources. Google patents describe possible implementations where those searches build separate source sets that are later combined (source).

Can a Google AI citation come from a source found after generation?

Google has patented that capability. Its Factuality of generated responses patent describes identifying a factual span in an already-generated answer, searching for a supporting document and attaching that result as a hyperlink or footnote. Google hasn't confirmed that every AI Mode citation is selected this way (source).