Auditing What AI Gets Factually Wrong About a Brand
A field set of eight, a named public adjudicator for each, and two fields deliberately thrown out. The exclusions are the reason the score is worth anything.
- Google's AI Overviews returned the correct answer about 91% of the time on a 4,000-plus query factual benchmark, but only 39% of those overviews were both correct and fully supported by the sources they cited.
- Most brand accuracy audits cannot be re-run because they never fixed a field set, never named an adjudicator, and sampled each prompt once.
- Within-prompt resampling alone accounts for 34.8% of the variance in LLM brand answers, so a single-shot audit is not a measurement.
- The eight fields worth scoring are legal name, founding year, service area, service list, hours, phone, ownership and licensing status, because each has a public source of truth outside your control.
- Price band and stated differentiator are excluded on purpose. No public authority adjudicates them, so including them would let the audited party score itself.
Ninety-one percent accurate is a number you cannot act on
In late 2025 and early 2026 the AI startup Oumi ran a fixed benchmark against Google's AI Overviews for the New York Times. Oumi reran the same experiment in both periods, once on Gemini 2 and again after the switch to Gemini 3. Search Engine Land reported the scale and the result: 4,326 searches, with accuracy rising from 85 percent to 91 percent.
Then Oumi checked whether the cited pages actually supported the answers. Only 39% of the total overviews were both correct and fully supported by their citations. Only 67 percent of individual claims were supported at all.
Brand accuracy in AI answers is only measurable if you fix the field set before you run the audit. Score eight objectively checkable facts (legal name, founding year, service area, service list, hours, phone, ownership, licensing status), name the public source that settles each one, and re-run the identical prompts on the identical engines. An audit that scores different things every time is a folder of screenshots, not a metric.
Two different failures hide inside the word accuracy. The engine can state something false. The engine can also state something true that nothing retrievable supports. Those have opposite fixes, and a single headline percentage tells you which one you have exactly never.
That is the flaw I want to fix here, and it is not a Google flaw. It is an audit design flaw, and it is ours.
Three reasons your last accuracy audit cannot be re-run
I have looked at a lot of these. They fail in the same three places.
No fixed field set. The auditor scores whatever the engine happened to say. One run catches a wrong phone number, the next catches a wrong founding year, and the two runs are not comparable. You cannot show improvement between two audits that measured different things.
No named adjudicator. The audit calls a claim wrong because the client said so. That is a complaint, not a finding. If the party being measured is also the referee, the score is decoration.
One sample per prompt. This is the one that quietly destroys everything. A 2026 variance-components study of 12,933 LLM brand responses across three models and eight languages found that within-prompt resampling alone accounts for 34.8% of the variance in brand answers. Ask the same question twice and you get materially different output, at temperature 0.3, with nothing else changed.
of variance in LLM brand answers comes from asking the same prompt again, nothing else changed
The same paper gives the practical stopping rule. A repeat past the fifth reduces relative-error variance by roughly 0.0003. Five runs per prompt is not a hedge, it is the point where more runs stop buying you anything. That is a better answer than any vendor has published on how many samples an AI visibility reading actually needs.
One more constraint, from a companion study of 3,750 responses across 50 brands and five industries: cross-model agreement on the top-recommended brand was 41.6%. Never pool engines into one score. Score ChatGPT, Gemini and Perplexity separately or you are averaging three different systems into a number that describes none of them.
The eight fields, and the public source that settles each
A field earns its place only if somebody other than the brand can settle it. That single rule cuts the list down fast, and what survives is boring, checkable, and durable.
| # | Field | Public source of truth | Pass condition |
|---|---|---|---|
| 1 | Legal name | State business entity registry, or SEC EDGAR for filers | Engine returns the registered entity name, or a trading name the registry links to it |
| 2 | Founding year | Registration date on the state entity record | Stated year matches the registry, or an earlier date the brand evidences on a crawlable page |
| 3 | Service area | Google Business Profile service area, capped by Google at about 2 hours of driving time | No city or region named that sits outside the profile |
| 4 | Service list | The brand's own service pages plus GBP Services | No invented service, and the top three revenue services all appear |
| 5 | Hours | The live hours field on the Google Business Profile | Open and close times match for the current day |
| 6 | Phone | GBP primary phone | Exact digits, correct area code, no legacy tracking number |
| 7 | Ownership | Officers or registered agent on the state entity record | Named owner matches the record, or the engine declines |
| 8 | Licensing status | State licensing board lookup, for example Utah's Division of Professional Licensing, or the NPI Registry in healthcare | License number and active status match |
Rank them by consequence, not by how often they break. Licensing status and phone are the two that cost real money the day they are wrong, because one invites a regulatory complaint and the other routes a buyer to a dead line. Founding year is the one clients care about most and the one that matters least.
Field 2 needs an honest caveat, because I have seen it fudged. A registration date is a proxy, not a birth certificate. Businesses reorganise, buy predecessors, and operate for years as sole proprietors before incorporating. If the brand claims an earlier founding, the audit accepts it only when the evidence sits on a page a crawler can reach. If the evidence lives in a filing cabinet, the engine was never going to find it, and the field fails.
Field 7 has a jurisdictional hole worth naming. Plenty of states do not publish officers for privately held entities. Where the registry is silent, the correct score is not "pass," it is "unadjudicable in this jurisdiction," and the field drops out of the denominator for that brand. Shrinking the denominator honestly beats inflating the score.
The two fields I cut, and why the exclusions carry the credibility
I wanted ten. I could only defend eight.
Price band, excluded. There is no public authority on what a service costs. The only source is the brand's own pricing page, and the brand is the party under audit. Score it and you have built a mechanism where the audited party grades its own paper.
Stated differentiator, excluded. "Fastest response in the county." "Most experienced team in Southern Utah." These are positioning claims. They are not false, they are unfalsifiable, which is worse for a metric. An engine that repeats your differentiator has not been accurate, it has been persuaded, and that belongs in a recommendation reading rather than a citation or accuracy reading.
Keep both rows visible on the scoring sheet, greyed out and stamped excluded. The visible exclusion is the part that makes the other eight believable. A scoring instrument with no declared out-of-scope is almost always a sales instrument.
This is the same discipline that separates a defensible client report from a vanity dashboard, which is why the exclusions belong in what you actually send the client, not just in your working file.
The run protocol
Freeze the prompt set
Write one prompt per field, phrased the way a buyer would ask. Save them verbatim. This is the same asset as your standing AI visibility prompt set and it must not drift between runs.
Capture the source of truth first
Pull the registry record, the GBP export, and the licence lookup before you touch an engine. Screenshot each with a date. The adjudicator has to exist before the verdict.
Run five repeats per prompt per engine
Eight fields times five repeats times three engines is 120 observations. Fresh session each time, no memory, no personalisation, logged-out where the engine permits it.
Score each observation in three states
Correct, incorrect, or not stated. Never collapse the third into the second.
Flag grounding separately
For every correct answer, record whether a cited source actually supports it. This is a second column, not a modifier on the first.
Compute per engine, never pooled
Report three accuracy rates and three grounding rates. Averaging them hides the engine that is actually hurting you.
One hundred and twenty observations sounds heavy until you run it. It is roughly ninety minutes of work for a single-location business, and it is the only version of this that produces a number you can put next to the same number in October.
Scoring in three states, not two
The distinction between wrong and silent is where most audits leak credibility, and the research says which way the leak runs.
The Tow Center at Columbia tested eight generative search tools across 1,600 queries and found they returned incorrect answers to more than 60 percent of them, with Perplexity at 37 percent wrong and Grok 3 at 94 percent. The behavioural finding matters more than the headline rate.
ChatGPT, for instance, incorrectly identified 134 articles, but signaled a lack of confidence just fifteen times out of its two hundred responses, and never declined to provide an answer.
Read that as an audit design instruction. Engines are structurally biased toward answering, so "not stated" will be rarer than your intuition says and "incorrect" will be more common. If your audit is returning a pile of blanks, suspect your prompts before you celebrate.
| Incorrect | Not stated | |
|---|---|---|
| What happened | The engine asserted a wrong value | The engine omitted the field entirely |
| Usual root cause | A stale third-party record outranks your own | The fact is not on any retrievable page |
| First move | Correct the upstream record, then wait for recrawl | Publish the fact in plain text on a crawlable page |
| How fast it clears | Weeks to months, because the bad source persists | Days to weeks, once the page is indexed |
| Risk if you misclassify it | You publish a fact that was never the problem | You chase a phantom bad citation |
The grounding column has its own literature and it is older than most people assume. Nelson Liu, Tianyi Zhang and Percy Liang measured four generative search engines for EMNLP 2023 and found a mere 51.5% of generated sentences are fully supported by citations, with 74.5 percent of citations supporting the sentence they were attached to. Three years on, Oumi's 67 percent claim-support figure sits in the same neighbourhood. The grounding problem has not been engineered away, so measure it as a standing column rather than as a one-off curiosity.
## What a fixed score still does not tell you The honest limits are not a disclaimer paragraph. They change how you read the output. The error rate is not yours alone. The European Broadcasting Union and the BBC evaluated over 3,000 responses across four assistants, 22 public service media organisations, 18 countries and 14 languages, and found 45% contained at least one significant issue, with sourcing the biggest single cause at 31 percent. Gemini alone hit 76 percent. Some proportion of your client's failed fields is baseline engine behaviour that no amount of publishing will fix.
See how the audit fits the method→This research conclusively shows that these failings are not isolated incidents. They are systemic, cross-border, and multilingual, and we believe this endangers public trust.
There is also no first party number to check yourself against. No engine publishes your accuracy rate. Everything here is sampled, and sampled measurement carries error bars whether or not the tool draws them, which is the standing caveat on every AI visibility tracking tool on the market.
None of the experts have special access to the internal workings of Google's local search algorithm.
He wrote that about local search, and it travels. Nobody auditing AI answers has privileged access either. The eight-field score is defensible because its inputs are public and its procedure is repeatable, not because it sees inside anything.
Fixing what the audit finds, in order
The repair order is fixed, because the cheap fixes are also the ones that unblock the expensive ones.
Start with retrievability. Google states that to appear in its generative features, a page must be indexed and eligible to be shown with a snippet. A stray nosnippet tag can zero out every field at once, which is why a crawler access check runs before an accuracy audit, not after it.
Second, put the eight facts in plain text on a page. Not in an image, not in a PDF, not only in markup. Google says outright that structured data is not required for generative AI search, and Ahrefs' controlled test on 1,885 pages found no meaningful citation uplift from adding schema. Mark it up anyway for classic search hygiene, but do not sell it as the accuracy fix. It is not, and the evidence against schema as an AI lever is now hard to argue with.
Third, correct the upstream records: the state registry, the licence board, the Google Business Profile, and the directories that syndicate them. This is unglamorous entity maintenance and it is where the incorrect state usually gets its material, which is the practical half of any real entity SEO programme.
Fourth, and slowest, work the third-party corpus. Across 15 SaaS brands, 84% to 93% of AI citation weight sat on third-party sites rather than the brand's own domain. If a review site, an old press release or a competitor's comparison page carries a wrong fact about you, your own site correcting it is necessary and often not sufficient. That is the same mechanism behind why brand mentions out-correlate backlinks, running in reverse.
Audiences are already doing this cross-checking themselves. BrightLocal's 2026 research found 45% of consumers have used AI tools to find local business recommendations, up from 6 percent the year before, and that 97 percent of AI users at least sometimes cross-check those recommendations against real reviews. A wrong fact in an AI answer does not just misinform, it fails the verification step the buyer runs next.
Re-running it, which is the entire point
Same eight fields. Same prompts, word for word. Same three engines. Same five repeats. Sixty to ninety days later.
Report four numbers per engine: correct, incorrect, not stated, and grounded. Movement in the incorrect column is the only one that proves a fix landed, because a field moving from incorrect to not stated means the bad source stopped winning but nothing replaced it. That is progress, and it is not victory.
Do not celebrate a swing smaller than the noise floor. With 34.8 percent of variance coming from resampling alone, a two-field improvement on a forty-observation-per-engine sample is inside the error bar. Say so in the report. Clients trust the practitioner who names the noise, and they eventually stop trusting the one who does not.
If you want the fuller apparatus this sits inside, the method is written up here, and the rest of the instrument set lives in the insights index. The next artifact to build against this one is a worked audit on a real brand, where these eight fields get run end to end and the failures get published alongside the passes.
Frequently asked questions
What is brand accuracy in AI answers?
It is the share of objectively checkable facts about a brand that an AI engine states correctly. Measured properly it covers a fixed field set, each field settled by a public source outside the brand's control, sampled multiple times per engine so the reading can be repeated later.
Why only eight fields?
Because only eight survive the test that somebody other than the brand can adjudicate them. Legal name, founding year, service area, service list, hours, phone, ownership and licensing status all have a registry, a board or a Google Business Profile that settles the dispute without the brand's opinion.
Why exclude price band and stated differentiator?
Neither has a public adjudicator. Price is settled only by the brand's own page, so scoring it lets the audited party grade itself. A stated differentiator like fastest or most experienced is unfalsifiable, which makes it unscoreable. Both stay visible on the sheet, marked excluded.
How many times should I run each prompt?
Five per prompt per engine. A 2026 variance decomposition of 12,933 LLM brand responses found within-prompt resampling accounts for 34.8 percent of variance, and that a repeat past the fifth reduces relative-error variance by roughly 0.0003. Five is where extra runs stop paying.
Should I average the score across ChatGPT, Gemini and Perplexity?
No. Cross-model agreement on the top-recommended brand was 41.6 percent in a 3,750-response study, so pooling averages three different systems into a number describing none of them. Report one accuracy rate and one grounding rate per engine, always.
What is the difference between an incorrect answer and a missing one?
Incorrect means the engine asserted a wrong value, usually because a stale third-party record outranks your own. Not stated means the fact was never retrievable. The first is fixed upstream at the source, the second by publishing the fact in plain text on a crawlable page.
Does adding schema markup fix wrong facts in AI answers?
There is no evidence it does. Google states structured data is not required for its generative AI features, and Ahrefs' controlled test across 1,885 pages against 4,000 controls found no meaningful citation uplift. Mark up your pages for classic search hygiene, not as an accuracy remedy.
How long before a corrected fact shows up in AI answers?
Missing facts tend to appear within days to weeks once the page carrying them is indexed. Incorrect facts take longer, often weeks to months, because the bad upstream source has to be corrected and then recrawled before the engine stops preferring it.
Sources
- Oumi. Oumi's study finds 50% of AI Overviews contain facts not supported by cited sources (2026-04)
- Search Engine Land. Google AI Overviews: 90% accurate, yet millions of errors remain (2026-04)
- Tow Center for Digital Journalism, Columbia Journalism Review. AI Search Has a Citation Problem (2025-03)
- European Broadcasting Union and BBC. News Integrity in AI Assistants (2025-10)
- Nelson F. Liu, Tianyi Zhang and Percy Liang, Findings of EMNLP 2023. Evaluating Verifiability in Generative Search Engines (2023-10)
- Dmitrij Zatuchin, arXiv. Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (2026-07)
- Dmitrij Zatuchin, arXiv. Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models (2026-06)
- Google Search Central. AI Features and Your Website (2025-12)
- Google Search Central. Optimizing your website for generative AI features on Google Search (2026-07)
- Ahrefs. Does schema help AI citations? A controlled study of 1,885 pages (2026-05)
- Google Business Profile Help. Guidelines for representing your business on Google (2026-07)
- Google Business Profile Help. Edit your Business Profile (2026-07)
- US Securities and Exchange Commission. EDGAR Full Text Search (2026-07)
- Utah Division of Professional Licensing. Licensee Lookup and Verification System (2026-07)
- Centers for Medicare and Medicaid Services. NPPES NPI Registry (2026-07)
- BrightLocal. Local Consumer Review Survey: AI and trust (2026-03)
- Aleyda Solis, Orainti. SaaS AI Search Optimization: where citation weight actually sits (2026-07)
- Whitespark. Local Search Ranking Factors 2026 (2025-11)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →