ContextBolt SEO Free for 7 days. Keyword data, Google SERPs and backlinks, inside Claude. SEO data inside Claude. Start free trial
Guide · How AI Assistants Choose Citations

How AI Assistants Choose Citations (2026)

An AI assistant picks citations in three steps, and they are easy to keep separate once you have seen them. It rewrites your question into several different searches. It retrieves candidate pages for each of those searches. Then it writes an answer and attaches a citation to the specific sentences it lifted from. Being findable decides the first step, being quotable decides the second, and being checkable decides the third. Most pages fail at a different step than their owner assumes.

That matters because the advice going around treats citation like a ranking. Publish good content, rank well, get cited. It is a tidy story and the data stopped supporting it more than a year ago.

This is the mechanical version. What actually happens between someone typing a question and your domain name appearing under an answer, sourced from the documentation each company publishes rather than from what the process feels like from outside. Every figure and user-agent string below was checked against the vendor’s own docs on August 5, 2026.

Quick answer
  • Three steps decide it: retrieval, extraction, attribution. Ranking only helps with the first.
  • You are never competing for the question that was typed. Google’s AI Mode splits it into a “multitude of queries” before anything is retrieved.
  • Only 37.9% of AI Overview citations rank in the top 10 for the query, per Ahrefs in March 2026. 31% rank nowhere in the top 100.
  • The unit that gets cited is a passage, not a page. Claude’s web search returns up to 150 characters of cited text per citation.
  • Blocking the wrong crawler removes you from a citation pool without touching your Google rankings.

How do AI assistants choose which sources to cite?

Start with the word itself. A citation in an AI answer is a link attached to a specific claim in generated text, marking the source the model drew that claim from. It is not a ranking position and it is not a link the model chose to reward you with. It is a receipt.

That distinction explains most of the confusing behavior. A ranking answers “which page is best for this query”. A citation answers “where did this sentence come from”. Those two questions have different winners more often than you would expect.

The pipeline behind the receipt is the same shape everywhere, whether you are in ChatGPT, Claude, Perplexity or an AI Overview.

Step 1, retrieval: The assistant turns your question into one or more searches and pulls back a candidate set of pages. This is the step traditional SEO touches. If you are not in the candidate set, nothing else can save you.

Step 2, extraction: The model reads the candidates and pulls the fragments that answer the parts of the question. Length, structure and clarity decide this step. Authority barely enters it.

Step 3, attribution: The model writes the answer and links the fragments it used back to where they came from. A page can be retrieved and read and still never appear, because nothing in it made it into the final text.

Three steps, three different jobs. Most “why am I not cited” questions are actually a question about step 2 being answered with step 1 advice.

What happens to your question before anything is retrieved?

This is the part people skip, and it changes everything downstream.

Google is explicit about it. Its May 20, 2025 announcement states that AI Mode uses a query fan-out technique, “breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf.” Google’s own name for the mechanism, in Google’s own words.

Anthropic documents the same behavior differently. Its web search tool documentation says the search process “can repeat multiple times throughout a single request”, and puts numbers on it: “Simple factual queries typically use 1-3 searches; comparative or multientity research can use 10 or more.”

So when a reader asks an assistant “what is the best bookmark manager for researchers”, nobody searches that string. The assistant searches something closer to six different things. What tools exist. What researchers need. What each one costs. What people complain about. Which are still maintained. Each of those runs its own retrieval, and each pulls its own candidate set.

The practical consequence is blunt. You are not competing for the query. You are competing for one sub-question inside it. A page that is the tenth-best overview of a topic but the single clearest answer to “what does it cost” can win the citation while the definitive guide gets nothing, because the definitive guide never separated that answer out cleanly enough to be lifted.

This is also why keyword-level thinking underperforms here. A fan-out does not have a keyword. It has a question shape, and the pages that get quoted are the ones that already answered a specific version of it.

Does ranking first still get you cited?

Less than it used to, and the drop is measurable.

Ahrefs analyzed 863,000 keyword SERPs and 4 million AI Overview URLs and published the breakdown on March 2, 2026. Only 37.9% of URLs cited in AI Overviews also appeared in the first 10 results for the same query. Another 31.2% ranked somewhere between position 11 and 100. And 31.0% ranked nowhere in the top 100 at all.

The comparison that makes it land is Ahrefs’ own earlier number. In July 2025, the same team measured 76.10% overlap with the top 10. So in roughly seven months the relationship between ranking first and being cited went from “mostly the same thing” to “a coin flip with worse odds”.

Where the cited page ranksShare of AI Overview citationsWhat it means for you
Top 1037.9%Ranking still helps. It is no longer the whole game
Position 11 to 10031.2%Page two content gets quoted constantly
Outside the top 10031.0%Retrieved by a sub-query you never targeted

Ahrefs attributes part of the shift to better parsing on their side and part of it to the fan-out itself, which is the honest reading. Some of the change is measurement. Plenty of it is real.

Here is the take I will defend. That bottom row is the most useful number in AI search right now, and almost nobody optimizes for it. Roughly a third of citations go to pages that do not rank for the query at all, which means they were retrieved by a sub-query nobody targeted, using a phrasing nobody researched. You cannot chase that with a keyword tool. You get it by answering narrow questions completely, on their own pages, in their own words. That is a different production habit, not a different tactic.

SEO tool ContextBolt SEO· Get found on Google and in ChatGPT· $35/mo See it

Which crawler actually has to reach your page?

Retrieval has a prerequisite. Somebody’s crawler has to have your page, and it is usually not the crawler you were thinking about.

The important thing here is that the AI companies run separate bots for separate jobs, and the docs are unambiguous about which does what. OpenAI’s bot documentation states plainly that “OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features”, while GPTBot “is used to crawl content that may be used in training our generative AI foundation models”. Two different bots. Two completely different consequences if you block one.

Perplexity splits it the same way. Its crawler documentation describes PerplexityBot as “designed to surface and link websites in search results on Perplexity”, adding that “it is not used to crawl content for AI foundation models”, and recommends allowing it in robots.txt. Perplexity-User is separate again, fetching a page live because a person just asked something, and the docs note that “since a user requested the fetch, this fetcher generally ignores robots.txt rules”.

AssistantCrawler that feeds citationsSeparate training crawler
ChatGPT searchOAI-SearchBotGPTBot
ChatGPT, live fetchChatGPT-Usern/a, user-triggered
PerplexityPerplexityBotNone. Says it does not train on crawls
Perplexity, live fetchPerplexity-Usern/a, ignores robots.txt by design

The failure mode this creates is quiet and common. A site adds a blanket AI block in robots.txt to keep its writing out of training data, catches the search bots in the same rule, and drops out of a citation pool without a single ranking moving. Nothing in Search Console reports it. The full user-agent list is in our AI crawlers list, and the file-level detail is in llms.txt vs robots.txt.

Does an AI assistant cite a page or a passage?

A passage. This is the single most useful thing to understand about the whole system, and it is sitting in public documentation that almost nobody in SEO reads.

Anthropic’s web search documentation describes exactly what a citation object contains. Citations are always on for web search, and each one carries a url, a title, and a cited_text field holding “up to 150 characters of the cited content”. Not the page. Not the section. Up to 150 characters.

Sit with that number. A citation is roughly one sentence long. The model is not deciding your page is good. It is deciding that one specific sentence on your page is the cleanest available way to state one specific fact, out of every sentence it retrieved from every candidate page.

That reframes what a well-optimized page even is. A 4,000-word guide is not one candidate. It is a few hundred candidate sentences, most of them useless for citation because they only make sense with the paragraph above them. A model reaching for a quotable fragment cannot use “as we saw in the previous section, this drops to about half”. It can use “Claude’s web search is priced at $10 per 1,000 searches”.

So the writing habits that get you cited are not the ones that get you shared.

  • Self-contained sentences beat flowing prose. Each claim should survive being lifted out of the page with no context around it.
  • Named numbers, dates and versions beat adjectives. “38% of citations” is quotable. “A significant share” is not.
  • One question per section. If a heading covers three questions, the answer to each is diluted into the other two.
  • Front-load the answer. The sentence that answers the heading should be the first sentence under it, not the conclusion of the argument.

None of that is a trick. It is just writing for a reader who is going to quote you and needs the quote to stand up alone. It is also, almost word for word, what both sides of the content chunking for SEO argument end up recommending, which is worth knowing before you pay anyone to optimize your chunks.

Why do assistants cite pages that are obviously worse than yours?

Because extraction and quality are different tests, and only one of them is being run at that moment.

Picture the model at step 2. It has eight candidate pages open and a sub-question to answer. It is not judging which site knows the most. It is scanning for the passage that answers this narrow thing most directly, then moving on to the next sub-question. A thin page with a clear one-line answer wins that scan against a thorough page that reaches the same answer in paragraph nine, every time.

There is a second reason, and it is less flattering to all of us. Assistants favor sources whose claims are checkable, because a checkable claim is a safer thing to repeat. A page with a specific figure, a date, and a named source is low-risk to quote. A page with confident prose and nothing to verify is a liability the model routes around, however good the prose is.

This is where the old authority work still pays, just indirectly. Being named on other sites, having a real author, and citing your own sources all raise the odds a model treats you as safe to repeat. Answer engine optimization covers that layer in full, and the assistant-specific tactics live in how to rank in ChatGPT.

What actually changes your citation odds?

Five things, ordered by how much they move and how fast.

1. Be in the candidate set at all: Allow the search crawlers, keep pages fetchable without JavaScript, and rank somewhere respectable for the sub-questions. Position 15 is genuinely in the game now. Position 90 mostly is not.

2. Split questions onto their own pages: One narrow question answered completely beats the same question buried inside a broader guide. This is the fan-out lesson applied to your sitemap.

3. Write liftable sentences: Answer in the first line under each heading. Keep the claim inside one sentence. Assume every paragraph will be read alone, because it will be.

4. Put checkable specifics in: Numbers, dates, versions, prices, user-agent strings, named sources. Every one of them is a reason for a model to pick your sentence over an equally true but vaguer one.

5. Keep it current: Retrieval systems carry freshness signals. Claude’s search results include a page_age field showing when a page was last updated, which tells you the age of a page is data the model sees, not a detail buried in your CMS.

What is not on that list is as important. Publishing an llms.txt file will not do it, and neither will schema markup on its own. Google’s own AI features guidance says you do not need to create new machine-readable files or AI text files to appear in AI Overviews or AI Mode. The work is in the pages.

How do you know whether any of it worked?

This is where most AEO advice stops, right at the point it gets useful.

You cannot check citations by asking an assistant about yourself once. Answers vary between runs, they are personalized, and one flattering response tells you nothing. What you need is the same set of prompts run repeatedly across assistants, with your brand’s mention rate tracked against named competitors over time. That measurement is the whole job, and doing it by hand collapses within a week.

ContextBolt SEO does it from inside the agent you already work in. You connect one MCP URL to Claude Code, Claude Desktop or Cursor, then ask in plain language whether AI tools are naming you, and it reports your AI visibility and share of voice against competitors alongside normal keyword, rank, backlink and SERP research. No dashboard to learn, because the agent is the interface. It is $35 a month with a 7-day free trial, and every lookup saves to your project as markdown so the research is still there next week. If you just want a single reading with no signup, the free AI visibility checker will give you one.

The short version

An assistant splits your reader’s question into pieces you never targeted, retrieves candidates for each piece, and quotes roughly a sentence from whichever source states that piece most cleanly. Ranking gets you into the running for one of those pieces. Clear, self-contained, checkable writing wins it.

The uncomfortable part is that this rewards a kind of writing the industry spent a decade training itself out of. Long guides that cover everything and refer back to themselves were the winning format for years. They are a bad shape for a system that quotes one sentence at a time. The pages getting cited now are shorter, narrower, and far more willing to just answer the question in the first line and stop.

How AI Assistants Choose Citations: FAQs

How do AI assistants choose which sources to cite?
In three steps. The assistant rewrites your question into several searches, retrieves candidate pages for each one, then generates an answer and attaches a citation to the specific sentences it drew from. Being retrievable, quotable and checkable matters at a different step each time.
Does ranking number one get you cited by AI?
It helps, but far less than it used to. Ahrefs analyzed 863,000 keyword SERPs and 4 million AI Overview URLs in March 2026 and found only 37.9% of cited pages also appeared in the top 10. Another 31% ranked nowhere in the top 100.
Do AI assistants cite whole pages or single passages?
Passages. Anthropic's web search documentation shows each citation carrying up to 150 characters of cited text from the source. The unit that gets quoted is a fragment, so a page that answers one question cleanly in one paragraph beats a page that answers it across four.
Which crawler has to reach my site for AI citations?
It depends on the assistant. OpenAI uses OAI-SearchBot to surface sites in ChatGPT search, separately from GPTBot for training. Perplexity uses PerplexityBot for its index. Blocking the search crawler removes you from that assistant's citation pool entirely. See our AI crawlers list.
Why does an AI assistant cite a worse page than mine?
Because the model is picking the source it can quote most cleanly for one sub-question, not ranking the best page on the topic. A thin page with a direct, self-contained answer beats a thorough page that buries the same answer in paragraph nine.