How each answer engine cites its sources

Inline footnotes, a list at the end, or nothing at all. What a team can and cannot see about citations in ChatGPT, Perplexity, Gemini, Google AI Overviews, AI Mode and Claude, and which crawler each one sends.

Sample and methodEngine note. Descriptions are qualitative and based on what is visible in stored answers. We report no rates here and have no internal knowledge of how any engine retrieves, ranks or decides to cite.

A citation is the only visible link between an answer and the web. It is where a brand mention can be traced to a page, where a PR placement can be shown to have mattered, and where most measurement of “AI visibility” either becomes evidence or stays a guess. The six engines we track expose citations very differently, and what they expose sets the limit on what you can measure.

This note describes what an answer shows. It does not describe how the engines work inside, because we do not know, and we are wary of anyone who says they do.

Two ways an answer gets made

From the outside, an answer arrives one of two ways. Either the engine ran a live web retrieval and shows you what it used, or it answered from what the model already held and shows nothing. The same engine can do either for the same question on different days. Some show a visible hint (a “searching the web” status, a sources panel, a footnote); some do not.

The distinction matters for measurement. A brand mention with displayed sources can be traced: you can open the pages and see whether they name you. A mention with no sources is still a real observation, but its origin cannot be known from the answer. We store both, label which is which, and never guess at a source that was not displayed.

Engine by engine

ChatGPT, web answers

When ChatGPT searches, citations appear as inline markers in the text and as a sources list that can be expanded. When it does not search, the answer carries no sources at all, and nothing in the text tells you what it drew on. Which mode you get depends on the question, the account settings and, from what we observe, the day. Three crawlers are relevant: GPTBot for training data, OAI-SearchBot for the search index behind web answers, and ChatGPT-User for pages fetched on a user’s behalf during a conversation.

Perplexity

Perplexity is retrieval-first. Nearly every answer carries numbered inline citations that map to an ordered source list, often several per paragraph. This makes it the most legible engine to measure: the link between sentence and source is explicit. Its crawlers are PerplexityBot for the index and Perplexity-User for on-demand fetches.

Gemini

Gemini mixes modes. Some answers show a sources or related-links section, sometimes attached to specific passages, sometimes as a block after the text. Others, particularly for well-known facts, show nothing. The robots.txt token Google documents for Gemini is Google-Extended. It controls whether content Google crawls may be used for Gemini, and blocking it does not remove you from Google Search.

Google AI Overviews and AI Mode

Both surfaces are built on Google Search and show their sources as link cards beside or beneath the generated text, sometimes with the citing passage highlighted. AI Overviews appear only for some queries, and not consistently for the same query. AI Mode is a conversational surface and shows links throughout. There is no separate crawler: both use Googlebot, so you cannot opt out of one without leaving Search. The documented levers are the usual snippet and indexing controls.

Claude, with web search

With web search on, Claude shows citations to the pages it used, attached to the relevant passages. With it off, it answers from training with no sources. As with ChatGPT, the same question can go either way. Its crawlers are ClaudeBot for training, Claude-SearchBot for search, and Claude-User for pages fetched at a user’s request.

What a team can see per engine, as observed in stored answers in mid-2026. Qualitative; behaviour changes without notice.
EngineVisible citationsWhere they appearAnswers with no sourcesCrawlers to allow
ChatGPTWhen it searchesInline markers and a sources listCommon when it answers from the modelGPTBot, OAI-SearchBot, ChatGPT-User
PerplexityAlmost alwaysNumbered inline, with an ordered listRarePerplexityBot, Perplexity-User
GeminiSometimesPassage-level, or a block after the textHappens, especially for well-known factsGoogle-Extended
Google AI OverviewsWhen an overview is shownLink cards beside the textThe overview may not appear at allGooglebot
Google AI ModeYesLinks throughout the conversationVariesGooglebot
ClaudeWhen web search is onAttached to passagesCommon when search is offClaudeBot, Claude-SearchBot, Claude-User

What you can and cannot measure

Given the above, here is what an honest citation metric can contain.

  • Can: whether a domain or page was displayed as a source for a question, on which engine, in which market, on which date.
  • Can: how many stored answers displayed it, and how that count compares with other domains for the same questions.
  • Can: whether the brand mention in an answer sits on a cited page, or whether the page is cited without naming you.
  • Cannot: why a source was chosen, or how it ranked inside the engine’s retrieval.
  • Cannot: whether a mention without displayed sources came from training data or from a retrieval the engine did not show.
  • Cannot: what a citation is worth in traffic, beyond what your own analytics record by referrer, which is partial and differs by engine.
How a citation becomes a count
  1. 1
    Store

    The full answer and every source URL the engine displayed, exactly as shown.

  2. 2
    Normalise

    Sources are counted at page level and at site level, so both questions can be answered.

  3. 3
    Link

    Each displayed source is tied to the mentions in the same answer, without inferring which sentence it supports unless the engine shows that.

  4. 4
    Count

    Cited domains are ranked over stored answers, per question, engine and market, with the count shown next to every rank.

The same pipeline for every engine. Where an engine shows less, the count is smaller and says so; it is never filled in.

What to do with it

Three practical consequences. First, allow the crawlers for the engines you want to appear in, and know which ones you are blocking; the free crawler access check reads your robots.txt bot by bot. Second, when you report visibility, report citations and mentions as separate counts, because they are separate observations. Third, weight your source strategy toward the engines whose citations you can see: a placement that shows up in Perplexity’s source list is evidence, while the same placement’s effect on a source-less ChatGPT answer is a hypothesis.

Our research page keeps the current per-engine notes and the date each was last checked.

Find your next AI visibility opportunity

Choose a market, compare your brand, and see the evidence behind your next move.

Start free trial