The local AI visibility study protocol, published first
I have not collected a single data point for this study. This page is the entire method, fixed on 29 July 2026, published before capture opens so that anyone can run it and contradict me.
- Every AI visibility study circulating in this industry publishes results without a protocol, which makes the findings unreproducible and unfalsifiable.
- This page publishes the protocol first: four engines, five verticals, six markets, twenty frozen prompts, eight runs per prompt, 3,840 captures, and a five field scoring rubric.
- No result appears anywhere on this page, because no data has been collected. The protocol was fixed on 29 July 2026.
- Mention, citation and recommendation are scored as three separate fields, because collapsing them is the most common measurement error in this category.
- The publication rule binds me: results publish against this protocol unchanged, deviations get logged, and a null or a failed reliability check publishes anyway.
I have not collected a single data point yet
Every AI visibility study I can find publishes results first and method later, or never. I am doing the opposite. What follows is the complete protocol for a local AI visibility study, fixed on 29 July 2026 and published before capture opens. There is no finding anywhere below this line, because no data exists yet.
That is not modesty. A protocol published after the results cannot be checked. Once you have seen the numbers you can choose the engine list that flatters them, drop the market where the figure went the wrong way, and quietly rewrite the scoring rule so that a passing mention counts as a recommendation. None of that requires bad faith. It requires only that the method was still negotiable when the data arrived.
A defensible AI visibility study publishes its protocol before capture: the engine list, the capture environment, the geography, the prompt construction, the runs per prompt, the scoring rubric, the inter rater procedure, and the blank capture sheet. If any one of those is decided after the data lands, the result is not falsifiable. This page is that protocol, fixed on 29 July 2026, with zero observations collected.
Science solved this problem decades ago and marketing never adopted the fix. The Center for Open Science puts the case plainly: preregistration exists because "the same data cannot be used to generate and test a hypothesis, which can happen unintentionally and reduce the credibility of your results." Our industry has no equivalent norm, no registry, and no cost for retrofitting a method to a headline.
The gap is not that marketing lacks AI search data. There is a great deal of it. The gap is that there is no way to tell a result from a selection, and nobody has been made to care.
What I found when I audited the studies already circulating
Before writing this I read the AI visibility studies and methodology guides currently ranking for this topic. They sort into three tiers, and the tiers are not where you would expect.
The academic preprints do it properly. Julius Schulte, Malte Bleeker and Philipp Kaufmann published Don't Measure Once in April 2026, arguing that AI search visibility has to be characterised as a distribution rather than a single point.
The inherent probabilistic nature of AI search changes this paradigm. Answers can vary across runs, prompts, and time, making one-off observations unreliable.
Pratyush Kumar's Ranqo preprint is the more instructive case. It reports more than 100,000 responses across 100 plus brands and five engines between March and May 2026, and then does something almost nobody in the commercial space does: it publishes seven numbered protocols it has designed but not yet run, and states of its own trajectory analysis that every trajectory is "an observational baseline, not a treatment effect." A paper that markets a product still drew the causal line in the right place.
The small honest practitioner studies are better than their reputation. The 2026 AI visibility study from Tristar Marketing Solutions discloses four engines, 80 queries, three businesses, a Middle Tennessee geography and a May 2026 window, then admits that one run per query makes the output "a sample rather than a verdict." I think the design is badly underpowered. I also think its self assessment is more rigorous than most six figure vendor reports.
The methodology explainers are the worst offenders. These are the pages that assert rigour without ever operationalising it. Rampify's AI visibility research guide is the cleanest specimen: it names six requirements including fresh context sub agents and aggregation across runs, and names no engine versions, no sample size, no geography, no scoring rubric and no inter rater step. Nothing in it is wrong. Nothing in it can be replicated, which means nothing in it can be contradicted either.
Ranked by how often it is missing, here is what the category actually omits. Inter rater procedure is absent almost universally. Runs per prompt is usually absent. Capture environment is usually absent. Geography is often absent. Engine version strings are almost never recorded. That ordering matters, because the first two are the ones that decide whether the number means anything.
No metric is auditable without access to the observations behind it.
The question, and the only three comparisons I am allowed to make
The study question, fixed: for a local service business in a United States market, how often does a leading AI assistant name it, cite it, and recommend it, and how accurate is what the assistant says about it?
Three pre-specified confirmatory comparisons, and only three:
- Recommendation rate differs by engine.
- Recommendation rate differs by vertical.
- Recommendation rate differs by market size band.
Everything else this dataset can support will be reported and explicitly labelled exploratory. That distinction comes straight from the preregistration literature, and it is not decoration. A design with twenty analysis cells and no pre-specified comparison will always surface something that looks significant. Declaring the three in advance is the only thing that stops the study from finding whatever it wants.
Why local, and why now. Sterling Sky's Joy Hawkins measured Google's AI powered local packs surfacing 5,943 unique businesses against 18,330 in traditional three packs across the same 322 markets, roughly a third as many winners. SOCi's index of nearly 350,000 locations found only 1.2 percent recommended by ChatGPT against 35.9 percent appearing in Google's local three pack. BrightLocal's consumer panel found 45 percent of consumers have used AI tools to find local business recommendations, up from 6 percent the prior year. Demand is arriving faster than supply of visibility, and the compression is happening without a single agreed measurement standard underneath it.
If you have a business profile that ranks really well, even if you don't see a drop in ranking, your calls from the business profile are going down over time.
The instrument, part one: engines, environment, geography
Engines
Four, all captured through the consumer interface, never through an API.
| Engine | Surface | Account state | Recorded per run |
|---|---|---|---|
| ChatGPT | chatgpt.com web UI | Logged out where permitted, otherwise a clean account with memory off and Temporary Chat on | Displayed model label |
| Google AI Mode | google.com AI Mode | Signed out browser profile | Interface variant |
| Google Gemini | gemini.google.com | Clean account, personalisation off | Displayed model label |
| Perplexity | perplexity.ai | Logged out, default model | Displayed model label |
The API substitution is the shortcut I am refusing, and it is worth saying why. Metehan Yesilyurt's teardown of how these tools collect data draws the distinction cleanly: API collection often queries a different model than the one consumers actually use, without web search, personalisation or interface specific behaviour. Cheaper capture would buy me a number about a model. I want a number about what a person in that city is shown. Anyone building or buying in this space should read what the tracking tools actually do before trusting a dashboard figure.
Capture environment, fixed
- Fresh browser profile per market block, cookies and storage cleared between blocks
- No custom instructions, no uploaded files, no memory, no prior turn in context
- Web access enabled, default model, default settings, no parameter tuning
- One prompt per conversation, zero follow up turns
- Full page HTML and full page screenshot archived for every run, filename carrying the UTC timestamp
- Model or interface version string recorded per run
- Refusals, clarifying questions and empty answers recorded as runs and never discarded
The last rule is the one people cheat on without noticing. Every discarded refusal inflates every rate that follows it, and no published study I read stated what it did with them.
Geography
Six United States markets, two in each of three size bands. Location is set and verified before every block, and the verification method is recorded on every row.
This is not a formality. Google's own documentation states that location is estimated from device location, from home and work addresses saved to the account, from prior activity, and from IP address, which is "roughly based on geography." Typing a city name into the prompt is therefore not a geography control. It is a string in a prompt. Both go on the sheet: the city named in the prompt, and the verified location context of the capture session.
The instrument, part two: prompts, runs, and the arithmetic
Four prompt archetypes, held constant across five verticals with only the service noun swapped. Written before capture, frozen, never edited mid study.
P1 Need
I need an emergency plumber in {market}. Who should I call?
P2 Comparative
Who are the best plumbing companies in {market}?
P3 Constrained
Which plumber in {market} can come out today and works on tankless water heaters?
P4 Criteria
What should I look for when hiring a plumber in {market}, and which local companies meet that bar?
SparkToro's work with Gumshoe.ai is the reason these are frozen. Across 142 human written prompts asking for the same kind of recommendation, average semantic similarity was 0.081. Two agencies tracking what they both call the same query are very often not tracking the same thing. Freezing four archetypes does not fix that problem. It makes the choice visible, which is the most anyone can honestly claim. There is more on how to build a prompt set that survives that critique.
SE Ranking's Yevheniia Khromova recommends 20 to 40 tracked prompts spread across journey stages, and warns that "Tracking more prompts doesn't compensate for tracking the wrong ones." Four archetypes across five verticals gives 20 distinct prompt strings, at the floor of her range. That is a deliberate choice: run depth buys precision within a cell, prompt breadth buys coverage across cells, and for a first study I want precision first.
Runs per prompt: eight.
And the arithmetic, computed before capture rather than after:
| Quantity | Value |
|---|---|
| Engines | 4 |
| Verticals | 5 |
| Markets | 6 |
| Prompts per vertical | 4 |
| Runs per prompt | 8 |
| Total captures | 3,840 |
| Primary analysis cell | Engine by vertical |
| Observations per cell | 192 |
| Number of cells | 20 |
At a measured rate of 10 percent, 192 observations give a 95 percent Wilson interval of roughly 6.5 to 15.1 percent. At 50 percent it is roughly 43 to 57 percent. That is the precision this design buys, it is arithmetic rather than a finding, and it is exactly why I will not be reporting a two point gap between two engines as a result. If you want the longer argument on this, I wrote it up separately as how big an AI visibility sample has to be.
The scoring rubric, and the inter rater step almost nobody runs
Five target fields, scored independently on every run.
| Field | Definition | Values |
|---|---|---|
| MENTION | Target business name appears in the answer prose | yes / no |
| CITATION | A URL on the target domain appears in the cited sources | yes / no |
| RECOMMENDATION | Target is presented as an option to use, not merely referenced | yes / no |
| POSITION | Ordinal position of the target's first mention among named businesses | integer or null |
| ACCURACY | Name, phone, address, hours and primary service, each scored separately | correct / incorrect / absent |
Mention, citation and recommendation are three fields rather than one for a specific, evidenced reason. Semrush and Kevin Indig's ghost citations study found that 61.7 percent of AI citations were ghost citations, where a page was used as a source but the brand name never appeared in the answer text at all. Any study that collapses those three into a single visibility score is producing a number that cannot be compared to any other study's number, including its own from last quarter. That distinction gets its own treatment in citations versus recommendations, and the accuracy half in brand accuracy in AI answers.
Answer level fields recorded alongside: count of distinct businesses named, total citation count, whether a map or local module rendered, refusal flag, clarifying question flag. Cited source URLs are captured in full so the source mix can be analysed later using the approach in how to analyse AI citation sources.
The inter rater procedure
- Two coders independently score a random 20 percent subsample, which is 768 runs, drawn with a published random seed.
- Krippendorff's alpha is computed per field.
- Alpha at or above 0.80, the field publishes.
- Alpha between 0.67 and 0.79, the rubric is revised once, coders are retrained, the entire subsample is recoded, and both alphas publish.
- Alpha below 0.67, the field is dropped from the published results and the failure is reported as a finding in its own right.
Those thresholds are not mine to negotiate. An alpha at or above 0.80 is "generally considered a satisfactory level of agreement," and the 0.67 to 0.79 band is "often considered the lower bound for tentative conclusions." Below that the published guidance calls it "indicative of poor agreement among raters."
A prediction, on the record, before any coding happens. MENTION and CITATION will clear 0.80 comfortably, because both are close to mechanical. RECOMMENDATION is the field that will struggle, because deciding whether a business was presented as an option to use rather than merely referenced is a judgment call. If it fails the gate, the headline metric of this entire study dies and I publish that instead of quietly loosening the definition.
None of the experts have special access to the internal workings of Google's local search algorithm.
## The statistics, the stopping rule, and the publication rule All proportions report as 95 percent Wilson score intervals. Never Wald. Brown, Cai and DasGupta's 2001 paper in Statistical Science showed that the coverage failures of the Wald interval are far more persistent than practitioners appreciate, and recommended the Wilson interval or the equal tailed Jeffreys prior interval for small samples. The Wald interval is what most marketing charts are silently using on the rare occasions they show an error bar at all, and it is worst exactly where AI visibility rates live: small samples, proportions near zero. Runs within a prompt are not independent observations, and pretending otherwise would flatter every interval in the report. So two estimates publish side by side. The primary treats the individual run as the observation, giving 192 per cell. The conservative estimate treats the prompt by market cell as the observation, giving 24 per cell and much wider intervals. Both appear in the results. If they point in different directions, the conservative one leads the writeup.
Book a call→- Capture window is 14 consecutive days, opening and closing on stated dates, never extended to chase a number
- Any engine that materially changes its interface or default model inside the window is flagged and reported separately, never silently pooled
- Every departure from this page is entered in a deviation log with a date and a reason, published alongside the results
- A null result publishes
- A failed reliability gate publishes
- If capture is blocked or the study cannot be completed, the failure publishes with the reason
This is the section that costs something, and it is the section every commercial AI visibility report leaves out. A stopping rule you can extend is not a stopping rule. A publication rule that only fires on a favourable result is not a rule, it is a content calendar. The whole point of fixing this now, in public, with a date on it, is that I no longer get to choose later.
What this protocol cannot tell you
Six limits, stated before anyone can accuse me of discovering them afterwards.
It is not causal. There is no treatment, no control group, no randomisation. This measures a state of the world, not the effect of any tactic on that world. Kumar's paper draws the same line and defers the causal question to a closed loop randomised protocol it has designed but not run. Nothing in the results will support a sentence beginning with the words "if you do X."
It says nothing about revenue. No public study connects AI visibility to closed business, and this one will not be the first. When somebody quotes you a revenue figure from AI search, ask for the denominator and watch what happens. The honest connection runs through your own first party data, which is a different problem covered in AI search attribution.
The depersonalised capture is a laboratory, not a market. Covered above, restated here because it is the limitation most likely to be quoted back at me.
Five verticals and six markets do not generalise to every local business. They generalise to five verticals and six markets. Anyone extending the numbers past that is doing their own inference, not reading mine.
It is a snapshot of a system that changes weekly. Two weeks of capture describes two weeks. Model updates, interface changes and index refreshes all land inside windows like this one.
It is a sample of engines, not a census. Copilot, Grok, Claude and Apple's surfaces are out of scope for version one, which means any share style figure is a share of these four and nothing more.
A measurement programme that will not state its own limits is selling certainty, and certainty is the one thing this category cannot currently supply. That principle sits underneath the whole Cited Method, and particularly the MEASURE stage, which is manual first for exactly the reasons on this page.
How to break this study
Run it. Everything above is sufficient to replicate the design exactly, and I would rather be contradicted by someone with a method than agreed with by someone without one.
The specific attacks that would damage or falsify the finding, in the order I think they are most likely to land:
- Run sixteen per prompt instead of eight and get materially different rates. That would mean eight runs undersamples the distribution and my intervals are lying about their own width.
- Score RECOMMENDATION with your own two coders and fail to reach alpha 0.80. That would mean the construct is not reliably codeable and my headline metric is not measuring a real thing.
- Capture the same twenty prompts signed in, with real location history, and get different rates. That would confirm the laboratory critique and force a redesign toward personalised capture.
- Repeat the identical window thirty days later and get a materially different answer. That would mean the number has a shorter half life than the report describing it, which would be the most useful finding of all.
Every one of those is a real way to be wrong, which is the test I would apply to anyone else's study before I quoted it. The rest of the evidence work lives in the insights index, including why schema does not appear to lift AI citations, why llms.txt is not a citation lever, and the crawler access audit that has to happen before any of this measurement is worth running at all. If you want this protocol applied to your own markets rather than mine, that is a conversation.
Results are not in. The protocol was fixed on 29 July 2026 and this page is the version of record. When the data lands it gets measured against this, unchanged, or the deviation gets published next to it.
Frequently asked questions
What is a pre-registered SEO or AI visibility study?
A study whose full method is published before any data is collected. The engine list, capture rules, prompt set, sample size, scoring rubric and analysis plan are all fixed in advance, so the findings cannot be retrofitted to the numbers that happened to appear.
Why does publishing the protocol first actually matter?
Because a method decided after the results can be tuned to them, usually without anyone intending it. The Center for Open Science makes the same point: the same data cannot be used to both generate and test a hypothesis without reducing the credibility of the result.
How many runs per prompt does an AI visibility study need?
More than one, and the exact number is a design choice you should state. This protocol uses eight. SparkToro measured under a 1 in 100 chance that two runs of the same prompt return the same brand list, so single run screenshots are noise rather than evidence.
Should AI visibility data be captured through the API or the consumer interface?
The interface, if you are making claims about what people see. API collection often queries a different model version without web search or interface specific behaviour. API capture is cheaper and faster, and it measures a model rather than a market.
What is inter rater reliability and why does an AI study need it?
It is the check that two independent coders scoring the same answers agree. Deciding whether a business was recommended or merely mentioned is a judgment call, so without a reliability figure the scoring is one person's opinion presented as data.
What Krippendorff's alpha counts as acceptable?
Published guidance treats alpha at or above 0.80 as a satisfactory level of agreement, 0.67 to 0.79 as the lower bound for tentative conclusions, and anything below 0.67 as poor agreement between raters. This protocol drops any field that falls under 0.67.
Why report Wilson intervals instead of the usual error bars?
Brown, Cai and DasGupta showed the standard Wald interval has unreliable coverage, and recommended the Wilson interval for small samples. AI visibility rates are usually small samples with proportions near zero, which is precisely where the Wald interval performs worst.
When do the results of this study publish?
After a fixed 14 day capture window, measured against this protocol unchanged. A null result publishes. A failed reliability check publishes. If capture is blocked and the study cannot be completed, that failure publishes with the reason.
Sources
- Julius Schulte, Malte Bleeker, Philipp Kaufmann (arXiv 2604.07585). Don't Measure Once: Measuring Visibility in AI Search (GEO) (2026-04)
- Pratyush Kumar, Ranqo (arXiv 2606.20065). Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines (2026-06)
- SparkToro with Gumshoe.ai. New Research: AIs Are Highly Inconsistent When Recommending Brands or Products (2026-01)
- Center for Open Science. Preregistration (2026-07)
- Brown, Cai and DasGupta, Statistical Science 16(2) 101-133. Interval Estimation for a Binomial Proportion (2001-05)
- Marzi, Balzano and Marchiori. K-Alpha Calculator methodological notes (2024)
- Metehan Yesilyurt. How AI Visibility Tools Actually Collect Data: API vs UI Scraping, Sampling and Methodology Transparency (2026-07)
- Sterling Sky (Joy Hawkins). The State of Local SEO in 2026 (2026-06)
- SOCi, reported by Search Engine Land. AI local visibility report 2026 (2026)
- BrightLocal. Local Consumer Review Survey: AI and trust (2026-03)
- Semrush with Kevin Indig / Growth Memo. The Ghost Citations Study (2026-06)
- SE Ranking (Yevheniia Khromova). How to Choose Prompts to Track for AI Visibility (2026-04)
- Google Search Help. Understand and manage your location when you search on Google (2026-07)
- Tristar Marketing Solutions. 2026 AI Visibility Study: How AI Search Sees Service Businesses (2026-05)
- Rampify. The AI Visibility Research Guide (2026)
- Whitespark (Darren Shaw). Local Search Ranking Factors 2026 (2025-11)
- Ahrefs (Louise Linehan and Xibeijia Guan). Do AI citations lift when you add schema? A 1,885 page study (2026-05)
- Google Search Central. Introducing Search generative AI performance reports in Search Console (2026-06)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →