A Citation Is Not a Recommendation
Four things can happen to your brand inside an AI answer. Every tool on the market reports them as one number, and the number is quietly wrong.
- Two 2026 studies measured the same ratio, the share of AI citations that also name the brand, and published 17.6% and 69.9%. Neither made an arithmetic error.
- Five of the pages currently ranking for this topic define citation in mutually incompatible ways, and three of them disagree on whether a linked brand name is a mention at all.
- Google does not use the word citation in its AI features documentation. It says supporting link.
- Citation, mention, inclusion and recommendation are separately countable events with different commercial value. Being cited in the tray while a competitor is named as the pick is a loss that reports as a win.
- The biggest driver of any citation and mention overlap number is not the engine. It is which prompts you chose. One overlap measure moved from 7.2% to 39% on prompt type alone.
Two studies, one metric, a four-fold gap
In 2026 two teams published research on the same narrow question: when an AI engine cites your page, how often does the answer also say your brand's name out loud?
Semrush, working with Kevin Indig, logged 3,981 domain appearances across 115 prompts, 14 countries and four engines. BuzzStream logged 12,000 AI responses covering roughly 200 brands across 10 industries, resolving 221,946 cited URLs of which 47,550 belonged to tracked brands.
Semrush published three buckets that sum to every appearance in its dataset: 61.7% cited with no brand mention, 13.2% cited and mentioned, 25.1% mentioned with no citation. Citations therefore account for 74.9% of appearances. Divide the 13.2 by the 74.9 and the share of citations that also name the brand is 17.6%.
BuzzStream published that exact ratio directly. It got 69.9%.
Same metric. Four times apart. No arithmetic error on either side.
A citation, a mention, an inclusion and a recommendation are four separate events inside an AI answer. Each can happen without the others, each carries different commercial value, and effectively every AI visibility number in circulation merges at least two of them. Count them apart or the number tells you nothing.
The temptation is to call this a methodology dispute and move on. It is not. Both teams measured carefully and documented their samples. What they did not share was a definition, and the field has never forced them to. That is the actual problem, and it is a vocabulary problem, not a data problem.
Nobody agrees what a citation is, and Google does not use the word
I fetched the pages currently ranking for this topic and pulled their definitions verbatim. They do not agree, and the disagreement is not cosmetic.
| Source | A citation is | A brand name shown with a link is |
|---|---|---|
| Semrush and Growth Memo | the domain appearing as a source link in the response | both a citation and a mention |
| Similarweb | explicit credit, typically a clickable link, footnote or inline attribution | a citation, and explicitly not a mention |
| Writesonic | the model referencing your content as the source, often with a link | a citation, and explicitly not a mention |
| BuzzStream | a reference generated alongside an answer to support or verify it | a mention too, since mentions count with or without a link |
| ZipTie | in-text source attribution, as in "According to Brand X's research" | decided by the wording, not by the link |
Read the third column again. Similarweb defines an AI mention as your brand appearing "without linking to your content," which means a linked brand name is not a mention. BuzzStream counts mentions "with or without a link," which means the identical event is a mention. Two pages ranking for the same query give opposite answers about the same pixel on the same screen.
ZipTie is stranger still. It classifies by evaluative framing rather than by linking, so a bare source link in the tray is not a citation at all under its rule, while a sentence in the answer body can be. That is a defensible taxonomy. It is simply not the one anyone else is using, and nobody says so.
This is what a young field looks like before it standardizes. Ahrefs' Louise Linehan describes the last twelve months as a shift from counting mentions to interrogating their value, one that has spawned an entire measurement vocabulary. She is right about the direction. The problem is that the vocabulary arrived faster than the definitions did, and vendors shipped dashboards on top of it anyway.
The cleanest evidence that the industry invented this vocabulary is that the largest engine does not share it. Google's AI features documentation never uses the word citation. It says pages are eligible "to be shown as a supporting link in AI Overviews or AI Mode." The noun the entire measurement category is built on does not appear in the primary vendor's own spec.
The four events, with a counting rule for each
Here is the taxonomy I use and the one the rest of the measurement posts on this site run on. Four events, four counting rules, four different commercial meanings.
1. Citation. Your URL appears in the answer's source apparatus: tray, sidebar, footnote or inline link. Count one per answer per domain, no matter how many links point at you. It does not require your name to appear anywhere a human reads. Commercial value is the lowest of the four, because Pew's passive panel of 900 US adults found users clicked a link inside an AI summary in just 1% of visits to pages that had one.
2. Mention. Your brand name appears in the answer text. Count one per answer, not one per repetition. This matters more than it sounds, and I will come back to it.
3. Inclusion. Your brand appears inside an enumerated set of options offered for the same job. Record two values, never one: the size of the set and your ordinal position in it. "In the set at four of six" is a materially different fact from "in the set."
4. Recommendation. The answer names your brand as the choice, with endorsing language directed at the user's stated need. Count it only when the endorsement is actually directed at that need, not when your brand is the example in a definition.
The ordering is by commercial value, not by nesting. You can be recommended without being cited, which happens constantly on Gemini, and cited without being named, which is the majority case on ChatGPT. Semrush's per-engine split makes that concrete: ChatGPT ran an 87% citation rate against a 20.7% mention rate, while Gemini ran 21.4% citations against 83.7% mentions. Those are dataset-wide rates rather than per-brand ones, so do not read them as your own odds. The direction is still unmistakable. The two engines make opposite default choices about attribution and naming, which means a dashboard that averages them is reporting on a system that does not exist.
The AI knows the information about the brand came from somewhere, but doesn't feel the need to explicitly say so to users. The brand name carries on its own.
That sentence is the reason event one and event two have to be separated. The engine treats attribution and naming as unrelated decisions. Any metric that fuses them is fusing two independent choices the model made for different reasons.
A mention with no citation indicates memory from training data. A citation indicates that the model needed more evidence and is typically retrieving that evidence live.
French, quoted in BuzzStream's study, gives the most useful mechanistic framing published so far, and it upgrades the taxonomy from bookkeeping into diagnosis. An unlinked mention is the model recalling you. A citation with no mention is the model checking something and not caring who you are. Those two failures need completely different fixes, and a blended visibility score hides which one you have.
Your prompt set decides your number, not the engine
This is the likeliest explanation for the 17.6 against 69.9 gap, and it is uncomfortable for anyone selling a dashboard.
BuzzStream segmented overlap by prompt type, running the ratio the other way round: the share of mentions that arrived with a citation. For single-brand prompts, where the user already named the company, 39% of mentions came with a citation. For open list and category prompts, the kind a real buyer asks, only 7.2% did. Head-to-head comparisons landed between them at 35.8%. Same engines, same brands, same window. That overlap moved more than five-fold on prompt selection alone.
Neither team published its prompt mix in a form that lets you recompute the other's number, so I cannot close the four-times gap arithmetically and will not pretend to. But a variable that swings one overlap measure five-fold is the first place to look, well before you blame the engines. Two agencies reporting a "citation to mention rate" to two clients are not reporting the same quantity, and neither client can tell.
SparkToro's volatility work compounds it. Across 2,961 prompt runs from 600 volunteers, the engines almost never returned the same brand list twice.
AIs do not give consistent lists of brand or product recommendations. If you don't like an answer, or your brand doesn't show up where you want it to, just ask a few more times.
The supply of citations is not stable over time either, which quietly breaks month-over-month reporting. Seer Interactive's John Lovett analyzed 206,412 ChatGPT responses across 43 brands and found citations per response jumped from 5.7 to 10.4, an 81% increase, starting on a single date in December 2025. Any brand tracking raw citation counts booked a large win that month for doing nothing. That is why a defensible prompt set and an honest sample size matter more than the engine you pick, and why tool comparisons are meaningless until you know which of the four events each tool counts.
MEASURE is stage two of the Cited Method, and these four counting rules are exactly what it counts.
See how the Cited Method measures this→The loss that gets reported as a win
Now the commercial argument, which is the whole reason the taxonomy is worth the trouble.
Picture the answer to "best commercial roofer in Salt Lake City." Your page is in the source tray. The answer body names a competitor as the pick and lists two others as alternatives. Your brand name never appears in the text a human reads.
Every tool on the market records that as a citation. Most blend it into a visibility score. The score goes up. You lost the sale.
The cited-and-unnamed half of that is not an edge case. It is the largest bucket Semrush measured: 61.7% of appearances were a source link where the brand name never appeared in the answer. Note the denominator carefully, because the number is widely misquoted. 61.7% is the share of all domain appearances, not the share of citations. Run it against citations only and 82.4% of cited appearances named nobody. The correction makes the finding worse, not better.
Position inside the answer is its own separate event, and Peec AI is the only team I found measuring it at scale: roughly 200,000 responses across eight engines, September 2025 to March 2026. Ranking first inside a frequently cited third-party listicle moved a brand's position in the answer earlier by 1.80 places in US Finance, 1.17 in B2B SaaS and 0.82 in Emerging MarTech.
A brand can appear exactly once at the very top and be the clear winner. Another brand might be repeated five times simply because the AI is comparing caveats, alternatives, or tradeoffs.
That is the strongest argument against frequency-based scoring I have seen, and it is why my counting rule for a mention is one per answer. Any tool that counts repetitions rewards being the cautionary example. Ehrlinspiel's own instruction is to track visibility and position as two separate metrics, which is the same conclusion arrived at from the other direction.
There is also a supply-side reason recommendations show up where you did not ask for them. Profound ran 50,000 prompts across seven industries and found that nearly half of every AI response contained unsolicited editorial content: comparisons, opinions and recommendations the user never requested. Roughly 47% of the content in verbose responses was material nobody asked for. Your brand is being ranked in answers to questions that were not ranking questions.
What to report instead
Replace the single visibility score with four lines. This is the format the rest of the measurement work on this site uses.
- Citation rate: share of sampled answers where your domain appears in the source apparatus, one count per answer.
- Mention rate: share of sampled answers where your brand name appears in the answer text, one count per answer.
- Inclusion rate and mean position: share of answers where you appear in an option set, plus the average of your ordinal position and the set size.
- Recommendation rate: share of answers where you are named as the pick for the user's stated need.
Report all four per engine and never averaged across engines. Averaging ChatGPT's 87% citation rate with Gemini's 21.4% produces a number describing no system that exists. Semrush's index work separately found the overlap between mentioned brands and cited domains on Gemini can run as low as 30%, with ChatGPT averaging about 15 sources per response against Gemini's three. These are different instruments pointed at different things.
The practical order of operations does not change much. Confirm crawler access first, because a blocked fetcher zeroes all four events at once. Understand that retrieval happens at the subquery level, which is why citations arrive from pages you were not targeting. Then work on the events themselves: earning citations is a different job from earning brand mentions, and both are different from being described correctly when you do get named, which is its own accuracy problem.
What changes is what you promise. A retainer that promises "AI visibility" has promised nothing measurable. A retainer that promises to move mention rate on a named prompt set, sampled at a stated cadence, has promised something a client can check and something you can fail. That is a better business to be in, and it is the only version of this work I am willing to sell.
These four terms are load bearing across everything else I publish. The MEASURE stage of the Cited Method is built on them, and every post in the insights archive uses them in exactly this sense. If you see me write "citation" anywhere on this site, I mean event one, and nothing more than event one.
Frequently asked questions
What counts as an AI citation?
Your URL appearing in the answer's source apparatus: the tray, sidebar, footnote or an inline link. Count one per answer per domain regardless of how many links point at you. It does not require your brand name to appear anywhere in the text a human actually reads.
Is a brand mention the same as a citation?
No, and the two happen independently. Semrush found 25.1% of appearances were mentions with no citation and 61.7% were citations with no mention. Only 13.2% were both. Treating them as one metric fuses two separate decisions the model made for different reasons.
What is a ghost citation?
A ghost citation is when an AI engine uses your page as a source link but never names your brand in the answer text. Semrush measured this at 61.7% of all domain appearances. Expressed against citations only rather than all appearances, it is 82.4%.
Why do two AI visibility tools report different numbers for my brand?
Usually because they count different events under the same label, and because their prompt sets differ. BuzzStream's mention-to-citation overlap moved from 39% on single-brand prompts to 7.2% on open category prompts. Prompt selection changes the number more than the engine does.
What is the difference between being included and being recommended?
Inclusion means your brand appears inside a set of options offered for the same job. Recommendation means the answer names you as the pick for the user's stated need. Record inclusion with two values, your ordinal position and the size of the set, never just presence.
Does an unlinked brand mention have any value?
Yes, and arguably more than a ghost citation. An unlinked mention puts your name in front of the reader, which a source link usually does not. Pew found users clicked a link inside an AI summary in only 1% of visits to pages that had one, so the tray is a weak traffic channel.
Should I count a brand mentioned five times in one answer as five mentions?
No. Count one per answer. Peec AI's research notes a brand can appear once at the top and be the clear winner while another is repeated five times because the engine is comparing caveats and tradeoffs. Frequency scoring rewards being the cautionary example.
Can recommendations be measured reliably?
Less reliably than citations or mentions. No large public study measures recommendation as a distinct class, and classifying an endorsement requires a judgement call two analysts can make differently. Measure it anyway, because it carries the commercial value, but disclose the softness.
Sources
- Semrush with Growth Memo (Kevin Indig). The Ghost Citations Study (2026-06)
- BuzzStream (Vince Nero). How AI Mentions and Cites Your Brand (New Study) (2026-06)
- BuzzStream (Vince Nero). AI Citations vs Mentions: What's the Difference and What's More Important? (2026-06)
- Similarweb (Shai Belinsky). AI Mentions vs AI Citations: Key Differences for GEO (2026-04)
- Writesonic (Pragati Gupta). AI Brand Mentions vs AI Citations: What's The Difference? (2025-08)
- ZipTie (Ishtiaque Ahmed). Mentions vs Citations vs Recommendations in AI (2026-03)
- Google Search Central. AI Features and Your Website (2025-12)
- Peec AI (Jan Ehrlinspiel). The Listicle Rank Effect: What Nearly 200,000 AI Responses Across 8 AI Engines Reveal About Brand Visibility (2026-05)
- Profound (Jasman Singh, Matthew Huo, Ryan Van). The Parrot Problem (2026-06)
- Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results (2025-07)
- SparkToro with Gumshoe.ai (Rand Fishkin). New Research: AIs Are Highly Inconsistent When Recommending Brands or Products (2026-01)
- Seer Interactive (John Lovett, VP Analytics). ChatGPT Changed Their Algorithm 46 Days Before Announcing Ads (2026-01)
- Semrush. Semrush Releases Expanded 2026 AI Visibility Index Analyzing 126 Million AI Search Prompts (2026-06)
- Ahrefs (Louise Linehan). AI Search Trends 2026 (2026-07)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →