Measurement Instruments · 14 min read

Building a Prompt Set That Represents Real Demand

Every published resource on this topic is a list of prompts. None of them is a derivation. Here is how the panel actually gets built, weighted, capped and versioned, with the fail state named at every stage.

0.081average semantic similarity between 142 human-written prompts asking for the same recommendationSparkToro with Gumshoe.ai
The short version
  • Two people asking for the same recommendation write prompts that are 8.1 percent semantically similar, so your prompt set is not a sample of demand, it is a construction that determines the number you report.
  • Between 65 and 85 percent of ChatGPT prompts match no traditional search keyword, which means keyword rewriting builds a panel biased toward the queries you already rank for.
  • Intent classes have different base rates (comparative prompts produce 2.4 times the brand mention rate of informational ones), so a flat average across clusters moves when your cluster mix moves, even with performance flat.
  • If you do not declare cluster weights, your prompt counts are your weights. An undeclared weight is still a weight.
  • AI Mode replaces 56 percent of its cited sources every week and ChatGPT Search 74 percent, so the frozen, versioned panel is the only variable in the system you actually control.

Your prompt set is the measurement, not a window onto it

SparkToro collected 142 human-written prompts asking for the same kind of recommendation and measured an average semantic similarity of 0.081 between them. Two people who want the same answer write questions that are roughly eight percent alike.

That number has a consequence this industry has not priced in. If prompts for a single intent are eight percent similar to each other, the forty you picked are not a sample of demand. They are a construction. Change the construction and your visibility score changes with your brand's actual standing held perfectly still.

The short answer

A prompt set is an instrument, not a list. Build it in five stages: derive clusters from observed demand rather than imagination, generate variants against a fixed run budget, weight the clusters explicitly, then freeze and version the panel. The panel is the only variable in AI visibility measurement you fully control, which is exactly why it has to be the one you nail down.

There is academic support for treating prompt choice as a first order variable rather than an implementation detail. A January 2026 paper by Jennifer Haase and colleagues ran 12 LLMs across 10 prompts with 100 samples each (N = 12,000) and found that prompts explained 36.43 percent of the variance in originality scores against 40.94 percent for model choice. Choosing which question to ask was almost as consequential as choosing which model to ask it of.

Honesty check before anyone quotes that at a client. The paper measures creative tasks, not search retrieval, and the effect is metric dependent: on fluency the same study put prompts at only 4.22 percent against 51.25 percent for model choice. I use it as a directional argument about where variance lives inside LLM outputs, not as evidence about AI search specifically. Nobody has published a variance decomposition of AI visibility into prompt, engine and run components. That is the most useful study in this field that does not exist.

The published work has the emphasis backwards. Almost everything written about prompt tracking is about size and run counts, which is the machinery downstream of the panel. Profound's own guide tells readers to start with a list of 100 prompts and notes that customers track anywhere from 100 to 1,000. Kevin Indig's Search Engine Land piece on prompt tracking accuracy is the best of the statistical writing, and even it opens with a worked forty prompt mix presented as a given rather than derived. Size is a solved question. Composition is not, and composition is the part that decides your number.

Stage one: derive demand, do not imagine it

Fail state: imagined demand.

The default construction method is keyword rewriting. Take your top thirty organic keywords, rephrase them as questions, call it a prompt set. Both Parse's prompt set guide and SE Ranking's prompt selection guide lead with a version of it.

It is structurally biased, and there is a number that shows why. Semrush analyzed US clickstream data from a 200 million user panel across more than a billion lines of data covering October 2024 to February 2026 and found that for most of the study period, between 65 and 85 percent of ChatGPT prompts could not be matched to any traditional search keyword in its database.

Keyword rewriting therefore starts you inside the 15 to 35 percent of prompt space that already looks like search. That slice is real and it is growing: the same study found the share of prompts using traditional search language nearly doubled from 18.9 percent to 34.9 percent between October 2025 and February 2026. But if keywords are your only source, you have built a panel systematically tilted toward the queries you already win.

Rank your sources by how close they sit to observed behavior, not by how easy they are to get.

TierSourceWhat it actually is
1Bing grounding queries, Search Console query dataEngine observed retrieval and search language, first party, free
2Sales call recordings, support tickets, live chat logsHuman demand in the customer's own words
3People Also Ask, forum and community threadsHuman phrasing, but filtered through a different surface
4Vendor prompt volume estimatesPanel data plus statistical modeling
5Asking an LLM to generate promptsModel priors wearing a costume

Tier one gets the least attention and deserves the most. Microsoft's AI Performance report in Bing Webmaster Tools publishes grounding queries, which Microsoft describes as the key phrases the AI used when retrieving content that was referenced in AI generated answers. That is engine side retrieval language, it is free, and hardly anyone mines it for panel construction. Microsoft also states plainly that the data shown represents a sample of overall citation activity, which is precisely the disclosure the paid trackers do not volunteer.

Tier four needs its own label. Profound is unusually candid that its prompt volume figures combine real panel prompts with statistical modeling, applying probabilistic extrapolation to correct for demographic and geographic bias. Modeled volume is a legitimate prior. It is not a measurement, and a panel weighted on it inherits somebody else's model.

Tier five deserves a warning rather than a caveat. Asking an LLM to generate your prompt list is the fastest method available and the only one guaranteed to correlate with model priors instead of user demand. You are asking the system under test to write its own exam.

65-85%
of ChatGPT prompts matched no traditional search keyword
0.081
average similarity between 142 prompts for one intent
56%
of AI Mode cited sources replaced every single week

Prompt tracking shouldn't be about capturing every possible query, because that turns into noise real fast.

Gaetano DiNardiPrincipal Consultant, Marketing Advice

Stage two: cluster by intent, because intent classes do not share a base rate

Fail state: over-clustering, or one flat average across everything.

Semrush and Kevin Indig ran 115 prompts across 14 countries and four engines, logging 3,981 domain appearances, and found comparative queries produced a 43.3 percent brand mention rate, 2.4 times the rate of informational queries. Commercial queries hit an 84.4 percent citation rate and a 35.6 percent mention rate.

Most people read that as a content tip. It is a measurement problem. Intent classes carry different base rates, so averaging across them means your headline number moves whenever your cluster mix moves, even when per cluster performance is completely flat. Add eight comparison prompts to a forty prompt panel and your score climbs. Nothing about your brand changed.

That is the most common false trend in generative engine reporting, and it is invisible unless clusters are reported separately. It is also why the way you build a client reporting layer has to expose cluster level movement rather than a single composite.

The cluster count rule follows from the base rate problem: enough clusters that prompts inside a cluster share a similar base rate, few enough that each cluster still holds enough prompts to move meaningfully. Four to six for most businesses. The failure at both ends is symmetrical.

Two ways to get clustering wrong
Under-clusteredOver-clustered
What breaksComposition changes read as performance changesNo cluster holds enough prompts to move above noise
SymptomScore jumps the week you add promptsEvery cluster line is jagged and nothing is significant
Typical shapeOne number for 40 prompts14 clusters of 3 prompts each
FixSplit until base rates inside a cluster convergeMerge until each cluster carries at least 6 to 8 prompts

Cluster boundaries also decide what you can diagnose. Separating citation-shaped intents from recommendation-shaped ones is what lets you tell a citation apart from a recommendation later, and it is what makes source analysis of who gets cited interpretable per cluster instead of mush.

Track by topic cluster, not every single prompt. You'll drown in data if you log every variation.

Melissa PoppVP of Content Strategy and Innovation, RicketyRoo Inc

Stage three: generate variants against a cap, not against your imagination

Fail state: variant explosion.

Surfer's fan-out research gives the only empirical anchor I know of for how many variants one intent actually spawns. Ten thousand keywords produced roughly 33,000 fan-out queries across 173,902 URLs, about 3.3 subqueries per head term. Google's own retrieval expands a question into a handful, not into forty. If you are generating twelve phrasings per prompt, you are modeling something the engine does not do. The mechanics of that expansion are worth understanding before you generate anything, because query fan-out determines what the engine is really retrieving against.

Local panels are where this detonates. Run the arithmetic. Forty base prompts, twelve cities, three phrasings each, five runs a week, three engines. That is 1,440 distinct prompts and 21,600 model calls per week. Nobody runs that. So the set gets quietly trimmed by whoever is doing the work that day, and the composition of your instrument becomes an accident nobody documented.

Invert the calculation. Fix the run budget first, then solve backwards:

cells = clusters x prompts per cluster x city variants x runs x engines

Pick the weekly call budget you will actually fund. Pick the run count (five per week is the common floor, and Storylake argues for ten runs per prompt per engine as a working floor, thirty where budgets allow). Then let the equation tell you how many cities you can afford. A 3,000 call weekly budget across three engines at five runs leaves 200 prompt variants. Five clusters of eight base prompts is forty, which buys exactly five city variants. Not twelve.

Which five matters more than most people think. Places Scout, analyzing Sterling Sky's ranking reports, found AI local packs featured 5,943 unique businesses against 18,330 in regular three packs, roughly a third as many winners. Across the 322 markets Sterling Sky looked at, 88 percent had fewer unique businesses in the AI local pack than in the traditional one. City level outcomes diverge sharply and compress hard, so cities have to be sampled rather than enumerated, and the sample should be stratified by revenue contribution rather than by population or alphabet. Then the city list gets frozen along with everything else. This is the same discipline that makes a local AI visibility study defensible instead of anecdotal.

Variant generation rules that survive contact with a budget
  1. Cap total cells before writing a single variant, not after
  2. Generate 2 to 4 phrasings per base prompt, matching observed fan-out breadth rather than imagination
  3. Sample cities stratified by revenue contribution, then freeze the list
  4. Never let a variant cross a cluster boundary, or you have silently reweighted the panel
  5. Record every variant against its parent prompt ID so you can collapse back to base level

Stage four: weight the clusters, or your long tail outvotes your revenue

Fail state: unweighted average.

Here is the part almost nobody says out loud. If you do not set cluster weights explicitly, your prompt counts are your weights.

Indig's worked example uses 40 seed prompts split 12 brand, 12 category and 16 problem. That is a 30/30/40 weighting whether or not anyone writes it down. It is a defensible split. It is still an undeclared choice that fully determines the headline number, and the reader of your report has no way to see it.

Rank the weight bases by how defensible they are:

  1. Closed revenue attributable to the cluster. Best available, and almost nobody has it, because attribution from AI search is genuinely hard.
  2. Pipeline or lead volume by topic. Good proxy, obtainable from a CRM within a quarter.
  3. Engine observed grounding query frequency. First party, free, sampled by the vendor's own admission.
  4. Modeled prompt volume from a vendor panel. Usable as a prior if you say it is modeled.
  5. Equal weight. Perfectly defensible if declared. Indefensible if it arrives by accident.

I am not going to pretend a validated weight scheme exists. None has been published, and anyone selling you one built it from judgment. The distinction that matters is not weighted versus unweighted. It is declared versus undeclared. Publish the weights next to the score, or the score is not auditable.

One more weighting trap sits above the cluster layer. Semrush's expanded index found ChatGPT cites about 15 sources per response while Gemini cites about 3. Averaging a citation metric across those two engines mixes populations whose slot scarcity differs roughly fivefold. Report per engine, then weight engines by where your buyers actually are. The same logic applies when choosing between AI visibility tracking tools, because most of them average engines by default and do not tell you.

The next iteration of prompt tracking will look less like rank tracking and more like polling: repeated runs, clear sampling rules, confidence intervals, segmented panels, and raw-answer audits.

Kevin IndigGrowth advisor and author of Growth Memo

MEASURE is stage two of the Cited Method, and the frozen panel is the artifact it produces. Access first, then measurement, then mapping.

See the measurement stage
Five stage process diagram showing prompt panel construction from demand derivation through freeze and version, with the failure mode named at each stage
The order is load bearing. Weighting a panel built from imagined demand produces a precise number about nothing.Sources: Semrush clickstream study (200M user panel), Semrush with Growth Memo ghost citations study (115 prompts, 4 engines), SISTRIX (82,619 prompts over 17 weeks).
Use this graphic on your site

Free to republish with a link back to this page. Copy the embed code:

<a href="https://josephtimpson.com/insights/ai-visibility-prompt-set"><img src="https://josephtimpson.com/assets/infographics/ai-visibility-prompt-set.svg" alt="Five stage process diagram showing prompt panel construction from demand derivation through freeze and version, with the failure mode named at each stage" width="1200" style="max-width:100%;height:auto"></a><p>Graphic by <a href="https://josephtimpson.com/insights/ai-visibility-prompt-set">Joseph Timpson</a></p>

Stage five: freeze and version, because the substrate rotates faster than your report

Fail state: unversioned panel.

SISTRIX tracked 82,619 prompts across 1,548,213 snapshots over 17 weeks in six countries. Weekly rotation of cited sources came in at 5 percent for AI Overviews, 56 percent for AI Mode and 74 percent for ChatGPT Search.

Sit with the AI Mode figure. More than half the sources behind an answer are replaced every week. Now change your prompts at the same time and you have two moving parts and one observation. There is no arithmetic that separates them afterwards. The panel is the only variable you fully control, and that is the entire argument for freezing it.

The same study found the rotation is not uniform: 86 percent of prompts have a stable core of domains rotating at close to zero, with the periphery churning at 89 percent a week. That reframes what you should even be measuring. Core membership is the durable signal, raw citation count is mostly periphery noise, and core membership only becomes visible across many weeks of identical prompts. A panel that changes monthly can never detect it.

The version protocol:

How to version a prompt panel so month two is comparable to month one
Step 01

Freeze v1.0 at baseline

Prompt text, cluster, weight, city variants, engines and run count all locked on the day the first measurement runs. Nothing is provisional.

Step 02

Additions create v1.1

New prompts start a fresh baseline for their own cluster only. Existing cluster series continue uninterrupted rather than being restated.

Step 03

Retirements are marked, never deleted

A retired prompt keeps its ID, gains a retirement date and a stated reason, and its historical data stays in the series.

Step 04

Every number carries its version

A chart spanning two panel versions gets a vertical line at the boundary and a footnote saying what changed.

Step 05

Re-derive quarterly with an overlap cycle

Run the old and new panels in parallel for one full cycle and publish both, so the reader can see how much of the step change was method rather than performance.

The change log needs nine fields per prompt: prompt ID, prompt text, cluster, weight, city variants, engines, run count, version added, version retired plus reason. That is the whole artifact. It fits in a spreadsheet, and it is the difference between a measurement and a screenshot.

The overlap cycle in step five is the expensive one, and it is the one everybody skips. It roughly doubles your run cost for a single reporting period. It is also the only thing that stops a panel refresh from producing a fake step change that gets presented as a win. If you are going to skip it, say in the report that you skipped it.

a single citation placement is not a reproducible result but merely a snapshot

Johannes BeusAuthor of the AI citation drift study, SISTRIX

What this method still cannot do

Every instrument has a stated accuracy. Here is this one's, without softening.

No first party share of voice exists inside any engine. Nobody sells you a denominator, because no engine publishes one. Every share of voice figure in this category is a share of your own panel. Say that in the report.

Repeated runs do not stabilize citations the way you would hope. Indig cites AirOps research across 815,000 prompt-page pairs finding that after running the same prompt three times in ChatGPT, only 2.2 percent of citations remain across all three. Freezing the panel controls one source of variance. It does not control the engine.

Prompt volume is modeled, not measured, by the vendor's own account. Weighting on it is inheriting a model you cannot inspect.

The variance decomposition I leaned on measures creative tasks, not search. It supports the direction of the argument. It does not prove it for retrieval, and I would rather flag that than let it pass.

Panel construction does not fix access. If an engine's crawler cannot reach your pages, a perfect panel measures your absence with excellent precision. Google's own documentation states that AI features are rooted in core Search ranking and require a page to be indexed and snippet eligible, which is why the crawler access audit runs before any of this, and why so much of what gets sold as generative engine optimization is built on claims that do not survive checking.

The honest summary: a well built panel converts an unfalsifiable claim into a defensible estimate with a stated method. It does not convert it into a fact. That distinction is what separates an audit you can hand to a client from a screenshot in a sales deck.

AIs do not give consistent lists of brand or product recommendations. If you don't like an answer, or your brand doesn't show up where you want it to, just ask a few more times.

Rand FishkinCo-founder, SparkToro

The whole construction on one page

Run this in order. Each stage consumes the previous stage's artifact, which is why skipping one quietly corrupts everything after it.

Prompt panel construction, five stages
  1. Pull demand from tier 1 and tier 2 sources first: grounding queries, Search Console, sales calls, support tickets. Log the source against every prompt.
  2. Cluster into 4 to 6 intent groups where base rates inside a cluster converge. Never report a single flat average across them.
  3. Fix the weekly run budget, then solve backwards for city variants. Generate 2 to 4 phrasings per base prompt, no more.
  4. Declare cluster weights explicitly, ranked by revenue proximity. If you use equal weights, write that down as a choice.
  5. Freeze v1.0, log nine fields per prompt, and re-derive quarterly with one overlap cycle running both panels in parallel.

The reason this matters more than the size of your panel: sample size answers how many prompts and runs, and it is a solved statistics question once you know what you are sampling. Composition answers what you are sampling, and it is unsolved because nobody publishes their derivation. Two agencies tracking the same brand with the same N and the same run count will report different numbers, and the entire difference will sit in the panel.

Which is also the argument for building the thing you can defend rather than buying the thing that ships fastest. The dashboard is not the instrument. The frozen, weighted, versioned, demand derived panel underneath it is, and it belongs to you.

More on the measurement layer in the insights archive, and the full five stage sequence lives in the Cited Method.

Frequently asked questions

How many prompts should an AI visibility prompt set contain?

Published guidance clusters around 20 to 50 for a starter panel, 50 to 200 for serious tracking, and 200 to 500 at enterprise scale. The number matters less than composition. Solve backwards from your weekly run budget instead, because run count multiplies against prompts, cities and engines.

Can I build a prompt set from my existing keyword list?

Partially, and you should know the bias you are inheriting. Semrush found between 65 and 85 percent of ChatGPT prompts matched no traditional search keyword in its database. Keyword rewriting anchors your panel to the queries you already rank for, so use it as one input among five, never as the only one.

Should I ask ChatGPT to generate my prompt list?

It is the fastest method and the weakest source. An LLM generated prompt list correlates with model priors rather than with what your buyers actually type, which means you are asking the system under test to write its own exam. Use it to find phrasings you missed, never as the derivation itself.

Why does cluster weighting change my AI visibility score?

Because intent classes have different base rates. Semrush measured comparative queries at a 43.3 percent brand mention rate, 2.4 times informational queries. Averaging across clusters means the headline number moves whenever the cluster mix moves, even when performance in every cluster is flat.

How many city variants should a local prompt set include?

However many your run budget affords after you fix clusters, prompts per cluster, runs and engines. For most local businesses that lands at three to five, not twelve. Sample cities stratified by revenue contribution rather than population, then freeze the city list along with the rest of the panel.

How often should I change the prompts I track?

Quarterly at most, and never silently. AI Mode replaces 56 percent of its cited sources weekly and ChatGPT Search 74 percent, so the engine is already a moving target. Changing prompts at the same time leaves two moving parts and one observation, which no amount of arithmetic can separate afterwards.

What is a prompt panel version and why does it need one?

A version is a frozen snapshot of prompt text, clusters, weights, city variants, engines and run count, recorded with a date. It exists so a chart spanning two panel versions can be annotated at the boundary. Without versioning, a panel refresh produces a step change indistinguishable from real performance movement.

Does a good prompt set give me true share of voice in AI answers?

No. No engine publishes a denominator, so every share of voice figure in this category is share of your own panel. A well constructed panel converts an unfalsifiable claim into a defensible estimate with a stated method. It does not convert it into a fact, and reports should say so.

Sources

  1. SparkToro with Gumshoe.ai. New Research: AIs Are Highly Inconsistent When Recommending Brands or Products (2026-01)
  2. Semrush. ChatGPT traffic analysis: Insights from 17 months of clickstream data (2026-04)
  3. Haase, Gonnermann-Muller, Hanel, Leins, Kosch, Mendling, Pokutta (arXiv). Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks (2026-01)
  4. SISTRIX. AI Citation drift: How stable are sources in AI search results? (2026-05)
  5. Search Engine Land (Kevin Indig). How to make prompt tracking much more accurate (2026-06)
  6. Semrush with Growth Memo. The Ghost Citations Study (2026-06)
  7. Surfer (Joshua Hardwick). Query fan-out impact on AI Overview citations (2025-12)
  8. Microsoft Bing Webmaster Blog. Introducing AI Performance in Bing Webmaster Tools (Public Preview) (2026-02)
  9. Sterling Sky (Joy Hawkins). The State of Local SEO in 2026 (2026-06)
  10. Profound. Prompt Volumes (2026-07)
  11. Profound (Nick Lafferty). How to Design Prompts for AI Visibility Tracking in 7 Practical Steps (2026-02)
  12. SE Ranking (Yevheniia Khromova). How to Choose Prompts to Track for AI Visibility (2026-04)
  13. Storylake (Andre Franco). Share of Model: a definition and a measurement protocol you can reproduce (2026-07)
  14. Parse (Dimitry Apollonsky). How to build an AI visibility prompt set (2026-04)
  15. Semrush. Semrush Releases Expanded 2026 AI Visibility Index, Analyzing 126 Million AI Search Prompts (2026-06)
  16. Google Search Central. Optimizing your website for generative AI features on Google Search (2026-07)
Joseph Timpson
Written by
Joseph Timpson

Joseph Timpson has worked in search since 2010 and runs Timpson Marketing out of St. George, Utah. He built The Cited Method, a five stage framework for earning and proving real citations in AI answers, and publishes what does not work alongside what does.

Want me to run this on your site and show you the before and after?

One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.

Book a free consultation