What AI Visibility Tools Actually Measure
Not one of these platforms has a first party feed from an AI engine. Every score you are about to put in a client report is a reconstruction from a sample the vendor chose, and the sampling rules are published in places nobody reads.
- No AI engine sells a brand-level share of voice feed, so every tracking tool is estimating a population it cannot observe.
- Runs per prompt is the largest driver of divergence between two tools looking at the same brand, and most vendors do not publish theirs.
- Ahrefs Brand Radar runs prompts through the free public web interfaces and refreshes the assistants monthly. Evertune samples each prompt 100 times per model through direct API access. Those two numbers cannot be compared.
- A mention and a citation are different events, and 61.7 percent of AI citations never name the brand at all.
- The two free first party reports, Google Search Console and Bing Webmaster Tools, give you less than the paid tools promise and more than the paid tools can verify.
No engine sells you the truth, so every tool is estimating
Point one tool at a brand this week and it returns a visibility score. Point a second tool at the same brand in the same week and it returns a different one. Neither vendor is lying to you.
They sampled different things.
No AI engine publishes a brand-level share of voice feed. Google reports impressions only, Microsoft reports an explicitly sampled slice of citations, and OpenAI, Anthropic and Perplexity report nothing to brands at all. Every AI visibility tracking tool is therefore reconstructing a population it cannot observe, from a sample it chose. The sampling method, not the dashboard, decides whether the number belongs in a client report.
This is not a secret the vendors are keeping. Semrush states it in its own help centre: AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility. Microsoft says the same thing about its own first party data, writing that the data shown represents a sample of overall citation activity. Ahrefs publishes a full methodology page and volunteers that it does not filter out hallucinated or malformed links, because they reflect real model output.
The vendors are more candid in their documentation than the comparison posts are on their behalf. That is the actual gap in this category. Frase's ten-tool roundup does not state runs per prompt, interface versus API, or raw response export for any tool on the list. Search Influence's analysis of over ten platforms prints a Typical Refresh Rate column without defining whether that means automated monitoring, scheduled batch runs, or manual spot checks. Brainlabs, to its credit, breaks rank and says outright that the question of how many people actually saw an AI response mentioning your brand is unknowable.
So the useful review is not a feature grid. It is a methodology teardown.
The seven facts that decide whether a number is usable
Everything a buyer needs to know about an AI visibility tool sits in seven facts, and I would rank them in this order of impact on the reported number.
- Runs per prompt. The single largest source of divergence. One capture of a probabilistic system is a coin flip with more sides.
- The counting rule. Whether a brand string in the answer text, a linked URL, or both, count as one event.
- Locale and language. Which market the prompt was issued from, and whether you can set it per prompt.
- Interface or API. Whether the tool queried the consumer product with live retrieval, or the raw model.
- Refresh cadence and reporting window. Whether today's figure is a same-day reading or a trailing average.
- Prompt set stability. Whether the vendor rotates the prompt library under your trend line.
- Raw response export. Whether you can audit the score or only accept it.
Notice what is absent from that list: number of engines covered, which is the column every affiliate grid leads with. Engine coverage changes what you can report on. It does not change whether the number is trustworthy. A tool covering eight engines at one run per prompt is producing eight noisy numbers instead of one, and I would take a single engine measured properly over that every time. If you have not yet frozen the prompt set you are measuring against, none of the seven matter, because you do not have a measurement yet, you have a search box.
What each tool documents about its own sampling
Below is what each vendor publishes about itself, drawn from vendor documentation and help centres rather than from reviews. Where a cell reads "not published," it means the fact did not appear on the pages I fetched, not that the tool lacks the capability. That distinction matters, and I have kept it honest rather than filling gaps with inference.
| Tool | What it queries | Cadence | Locale control | Counting rule | Raw export |
|---|---|---|---|---|---|
| Ahrefs Brand Radar | Free public web interfaces of ChatGPT, Gemini, Perplexity and Copilot | Assistants refreshed monthly, AI Overviews and AI Mode continuously, 90 day reporting window | Locale mirrors the ratio of queries by country and language in its keyword database | Citations are linked URLs, mentions are string matches, hallucinated links not filtered out | Stored responses searchable in product |
| Semrush AI Visibility Toolkit | Real requests, explicitly not LLM APIs | Daily on a rolling basis | 40+ regional databases | Not defined in the help centre | Not published |
| Profound | Consumer AI search engines rather than static API calls | Not published | Not published | Not published | Not published |
| Peec AI | A fixed set of models per prompt | Daily, a regular 24 hour cycle | Per prompt country, set with ISO 3166-1 alpha-2 codes | Not defined in the prompt setup docs | Not published |
| Otterly.AI | All AI search engines in your account | Daily | Not stated on the monitoring help page | Brand Mentions and Domain Citations reported as separate metrics | Response by response breakdown in product |
| Evertune | Direct model APIs, plus live consumer apps for the retrieval layer | Not published as a cadence, but 100 runs per prompt per model | Not published | Not published | Not published |
| Scrunch AI | Browser automation and official platform APIs across eight engines | Not published | Any language and any country | Not published | Responses API with full text, citations and CSV export, 90 day history |
| SE Ranking AI Results Tracker | ChatGPT, Perplexity, Gemini, AI Mode, AI Overviews | Not stated on the product page | 7 supported markets | Linked and unlinked brand references counted separately | Not stated on the product page |
Read down the cadence column and the problem announces itself. Peec AI runs every active prompt on a 24 hour cycle. Semrush updates daily on a rolling basis. Ahrefs refreshes ChatGPT, Perplexity, Gemini and Copilot monthly against a 90 day reporting window, while its Google surfaces update continuously. Those are not slower and faster versions of one metric. A monthly capture inside a 90 day window is a smoothed estimate of a trend. A daily capture is a time series with all the noise still in it. Plotting them on the same chart is a category error.
The interface-versus-API split is the other fault line, and the vendors argue about it in public.
At Profound, our unique advantage lies in actively monitoring AI Search Engines like ChatGPT, Perplexity, and Copilot, rather than relying solely on static API model calls.
Ahrefs takes the same side and says so mechanically: all prompts run through the free, publicly available web interfaces of ChatGPT, Gemini, Perplexity, Copilot and other supported platforms, to reflect typical user experiences. Semrush is blunter still, stating that prompt responses are captured from real requests and not via any APIs of LLMs.
Evertune argues the opposite case, and it is not a weak one. Through direct API access it isolates what it calls the model's foundational knowledge, the view of your brand baked in during training, before any search happens. Scrunch splits the difference, using browser automation and official platform APIs together.
Here is the ranking nobody in the category will state plainly. If you are reporting on what a buyer sees, interface capture is the only defensible source, because API calls do not reproduce the retrieval layer, the system prompt, or the product's own re-ranking. If you are diagnosing why a model dislikes your brand before it searches, API capture is the only source that isolates the prior. Those are two different jobs, and a tool that does one well is not a worse version of a tool that does the other.
Runs per prompt is the biggest single driver of divergence
If you only interrogate one column, make it this one.
SparkToro and Gumshoe.ai put 600 volunteers and 12 prompts through 2,961 runs across ChatGPT, Claude and Google's AI, and found less than a 1 in 100 chance of the same brand list appearing in any two responses to the same prompt.
There's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses.
Now apply that to the table above. A tool that captures each prompt once per day is reporting a single draw from that distribution and calling it a reading. A tool that captures 100 draws per model is reporting a mean with a knowable spread. Both can be labelled "visibility score" in a dashboard header. Only one of them survives a client asking why the number moved.
Evertune's 100 runs per prompt per model is, as far as I can find, the only sampling depth any vendor in this category states as a number. Kevin Indig, writing in SE Ranking's guide to choosing prompts, suggests around five consecutive runs of a prompt once a week and roughly 15 prompts per persona, which is a practitioner's floor rather than a statistical one. SE Ranking's own research in the same piece found only 35 percent of domains repeat across AI answers, with two thirds vanishing between runs.
The rank order is therefore: runs per prompt beats engine count, beats refresh cadence, beats everything else on the spec sheet. I have written separately about how many runs a prompt set actually needs before a delta means anything, and the short version is that most published GEO reporting is running an order of magnitude short.
Mentions and citations are different events, and half the dashboards blur them
A brand name appearing in answer text and a domain appearing in the source list are not the same thing happening. Semrush and Growth Memo logged 3,981 domain appearances across 115 prompts in 14 countries and found 61.7 percent of citations never mentioned the brand in the answer at all. Count citations and you overstate brand impact by roughly two and a half times.
Brands need to map not just which topics they want to appear in, but which phrasing patterns produce mentions versus ghost citations.
The tools handle this unevenly. Otterly.AI reports Brand Mentions and Domain Citations as two separate metrics, which is the correct default. Ahrefs defines citations as linked URLs and mentions as string matches, which is precise and also means a competitor writing about you registers as your mention. SE Ranking counts linked and unlinked references separately. Several others do not publish a definition at all, which means their headline score is an undocumented blend of two events with different commercial value.
There is a second trap underneath. Semrush's expanded index of 126 million prompts found ChatGPT cites around 15 sources per response while Gemini cites around 3. Citation slots are roughly five times scarcer on one engine than the other. Any cross-engine citation count that is not normalised for slot scarcity is measuring the engine, not the brand. I have not seen a single tool disclose whether it normalises for this, and it is the first thing I would ask.
If you want the underlying taxonomy, the four distinct events an AI answer can produce are worth internalising before you pick a counting rule, and the case for brand mentions as the stronger correlate is the reason the distinction is commercially live rather than pedantic. Accuracy is a third axis again, because a mention that gets your service area wrong is a liability, and what the answer says about you is not captured by any presence metric.
Locale and personalization are the columns nobody fills in
Geography is where the reported numbers quietly stop describing your market.
Ahrefs parameterises locale to mirror the ratio of queries by country and language in its keyword database. Read that carefully. It is a globally weighted sample, tuned to the shape of the keyword corpus, not a control you set for a client in Salt Lake County. It is a defensible design for a market-level product and the wrong instrument for a single-location business. Peec sets location per prompt with ISO country codes. Scrunch supports any language and any country. SE Ranking lists seven supported markets. Semrush covers 40 or more regional databases.
Those are not equivalent capabilities and no roundup I read distinguished them. For local service businesses this is the decisive column, and it is the one I check first when a tool's number disagrees with what I see in the product myself. The problem compounds because AI local surfaces compress the winner set hard, which I cover in the local AI visibility methodology piece.
Personalization is worse, because almost nobody documents it. Ahrefs at least tells you its prompts run through free public web interfaces, which implies a logged-out, memory-free session. Most vendors say nothing about account state, memory, or custom instructions. Semrush's admission that responses are highly personalized is the closest thing to an industry position, and it is an admission, not a control.
If you are about to sign an annual contract for a visibility platform, an hour spent on the sampling questions is cheaper than a year of reporting a number you cannot defend.
Book a working session→The free first party data is better than the category admits and worse than you hope
Two engines now report to site owners directly. Both are free, both are narrower than the paid tools imply, and both are more verifiable than anything a third party can offer.
| Google Search Console | Bing Webmaster Tools | |
|---|---|---|
| Surfaces | AI Overviews and AI Mode, combined into one view | Copilot, AI summaries in Bing, and select partner integrations |
| Core metric | Impressions only | Total citations and average cited pages per day |
| Query data | None | Grounding queries the AI used to retrieve content |
| Clicks or CTR | Not included | Not included |
| Completeness | Rolling out by property, standard 1,000 row limit, newest data preliminary | Explicitly described as a sample of overall citation activity |
| Competitor view | None | None |
Google's generative AI performance report counts an impression when a link to your site is shown in a generative AI feature, and stops there. No clicks, no CTR, no queries, and no way to separate AI Overviews from AI Mode. Microsoft's AI Performance report goes further on one axis by exposing grounding queries, which is Microsoft's own name for the fan-out subqueries behind a single user prompt, and it explicitly warns that its citation counts do not reflect ranking, authority, or the role of any page within an answer.
Here is the uncomfortable synthesis for the paid category. On the two engines where first party data exists, the free report is the ground truth and the paid tool is the model of it. On every other engine, the paid tool is all you have and there is nothing to validate it against. That asymmetry should shape what you buy, and almost no vendor comparison acknowledges it exists.
Server logs are the third free instrument and the most underused. They will not tell you whether you were cited, but they will tell you whether the retrieval bots reached the page in the first place, which is why AI crawler log analysis belongs upstream of any visibility subscription.
By the time a user sees an answer, the model has already decided which sources matter.
What I actually put in a client report
I report three things from tooling and one thing from the site.
First, a directional presence trend from a frozen prompt set, captured at a fixed cadence, labelled as an estimate with the run count printed next to it. Second, the source domains appearing in those answers, because the competitive read is more actionable than my own score. Third, the first party impression and citation counts from Search Console and Bing Webmaster Tools, which are the only figures in the deck a client could independently verify.
The fourth is not a visibility metric at all. It is whether the retrieval crawlers can reach the pages, which is the one mechanical lever with near-certain causality and the one most sellers skip. If OAI-SearchBot is getting a challenge page, every score above it is measuring the consequence of a fixable server-side decision. Google states the same prerequisite in its own documentation: a page must be indexed and eligible to be shown with a snippet before it can appear in a generative AI feature at all, which means a stray nosnippet directive can zero a visibility score no tool will ever explain. That is why a real audit starts with access and ends with measurement, not the other way around.
What I do not report is a share of voice percentage presented as fact, a month-over-month delta from single-run data, or any claim that a specific piece of work caused a specific citation. AI search attribution does not currently support that claim, and saying so out loud has never once cost me a client. Presenting a tool's estimate as a measurement is the thing that eventually does. The broader discipline of reporting GEO work honestly is mostly this: label the estimate, print the sample size, and never let a dashboard header do your thinking.
Buy the methodology, not the dashboard
The feature grids are converging. Every platform in this category will track the same six or seven engines by the end of the year, and the interface differences will be cosmetic. What will still differ, and what will still determine whether your number is defensible, is how each one samples.
- Do you query the consumer interface or the model API, and which specific models?
- How many runs per prompt, and is the reported figure a mean across runs or a single capture?
- What locale and language does each prompt run in, and can I set it per prompt?
- Is the session logged out and memory-free, or personalized in any way?
- Does a mention mean a linked citation, a string match on the brand name, or both counted once?
- How often does the prompt library itself refresh, and does that break my trend line?
- Can I export the raw response text, not just the score?
- What is the reporting window, and is today's figure a same-day reading or a trailing average?
A vendor that answers all eight without hedging is selling you an instrument. A vendor that cannot answer four of them is selling you a chart.
And hold the line on what the tooling is for. It is a directional signal that tells you where to look. It is not evidence, it is not attribution, and it is not a substitute for the mechanical work. The industry spent two years selling llms.txt files and schema markup as citation levers on exactly the same reasoning: a plausible mechanism, a dashboard that moved, and no controlled evidence underneath. The measurement layer deserves more scepticism than the tactics did, not less, because it is the layer that decides which tactics you believe. More of my working notes on this sit in the insights archive, and if you want the sampling questions applied to your own stack, that is a conversation worth booking.
Frequently asked questions
Why do two AI visibility tools show different numbers for the same brand?
Because they sampled differently. Runs per prompt, locale, refresh cadence, and whether the tool queried the consumer interface or the model API all change the result. Divergence is usually a methodology difference, not an error by either vendor. Compare their published sampling rules before assuming one is wrong.
Does any AI engine give brands a first party share of voice number?
No. Google Search Console reports impressions in AI Overviews and AI Mode combined, with no clicks, CTR, or query data. Bing Webmaster Tools reports citations and grounding queries, and calls its own figures a sample. Neither shows competitors. Every share of voice metric on the market is a third party estimate.
How many runs per prompt does an AI visibility tool need?
More than one, and most tools do not publish theirs. SparkToro measured under a 1 in 100 chance of two runs of the same prompt returning the same brand list. Evertune publishes 100 runs per prompt per model. Kevin Indig suggests around five consecutive runs weekly as a practical floor.
Is API-based tracking or interface-based tracking better?
They answer different questions. Interface capture reproduces what a buyer actually sees, including live retrieval and product re-ranking, which is what belongs in a client report. API capture isolates the model's trained view of your brand before any search happens, which is better for diagnosis. Neither is universally superior.
What is the difference between a mention and a citation in these tools?
A mention is usually a string match on your brand name in the answer text. A citation is a linked URL in the source list. Semrush and Growth Memo found 61.7 percent of citations never name the brand, so the two counts diverge sharply. Several tools do not publish which one their headline score uses.
Can I use a free tool instead of paying for AI visibility tracking?
Partly. Google Search Console and Bing Webmaster Tools give verifiable first party impression and citation data at no cost, and server log analysis tells you whether retrieval crawlers reached the page. What free tools cannot give you is competitor comparison or coverage of ChatGPT, Claude, and Perplexity.
How often should AI visibility be reported to a client?
Match the reporting interval to the sampling interval, not the meeting calendar. If the tool refreshes monthly against a 90 day window, weekly reporting invents movement that the data cannot support. Print the run count and the reporting window next to every figure so the client can judge it.
Does geolocation matter for AI visibility tracking?
Enormously, and it is the least documented column in the category. Some tools let you set a country per prompt with ISO codes. At least one weights locale to mirror its own keyword database rather than your market. For a single-location business, a globally weighted sample describes someone else's customers.
Should I trust an AI visibility score in a sales pitch?
Only if it comes with a run count, a locale, a date range, and a counting rule. A single screenshot of an AI answer is not evidence, because the same prompt run again will likely return a different brand list. Ask for the raw responses, not the score.
Sources
- Ahrefs. Ahrefs Brand Radar Methodology: How we collect and model AI visibility data (2026)
- Semrush. Where does the data in Semrush's AI Visibility Toolkit come from? (2026)
- Profound. Seeing what customers see: direct AI search engine monitoring vs. API limitations (2026)
- Peec AI documentation. Setting up your prompts (2026)
- Otterly.AI help centre. How does Prompt Monitoring with OtterlyAI work? (2026)
- Evertune. How Evertune Measures AI Visibility, Methodology (2026)
- Scrunch AI. Scrunch FAQs (2026)
- SE Ranking. AI Search Visibility Tool (2026)
- SE Ranking (Yevheniia Khromova). How to Choose Prompts to Track for AI Visibility (2026-04)
- Google Search Console Help. Generative AI performance report (Search) (2026)
- Microsoft Bing Webmaster Blog. Introducing AI Performance in Bing Webmaster Tools (Public Preview) (2026-02)
- SparkToro with Gumshoe.ai. New research: AIs are highly inconsistent when recommending brands or products (2026-01)
- Semrush with Growth Memo. The Ghost Citations Study (2026-06)
- Semrush. Semrush releases expanded 2026 AI Visibility Index analyzing 126 million AI search prompts (2026-06)
- Brainlabs. Your AI Visibility Data Is Wrong (And That's Okay) (2026)
- Search Influence. AI SEO Tracking Tools 2026: Comparative Analysis of Over 10 Platforms (2026)
- Frase. The 10 Best AI Visibility Tools in 2026 (2026)
- Clearscope (Bernard Huang). The search engine wars are back with AI (2026-03)
- Google Search Central. AI Features and Your Website (2025-12)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →