Every AI visibility tool sells you a percentage. Your brand appears in 23% of AI answers. Your rival gets 41%. The number sits on a dashboard, it goes up and down, and it feels like a ranking. That familiarity is the problem, because a ranking is a fact about a page Google published and a share of voice number is a fact about a spreadsheet a vendor built in private.
I went looking for someone who would show their work. The first page of Google for this term returns ten results. Eight of them belong to companies selling AI share of voice tracking. Not one publishes the prompt set behind its own number, and the two that give a formula give different formulas.
That is not a scandal. It is early. The metric is about eighteen months old and nobody has agreed what goes on the bottom of the fraction yet. But it does mean that if you are about to report a share of voice figure to a client or a board, you should know exactly what you are reporting and how much of it is noise.
This guide covers the definition, the formula, the denominator problem, and the piece almost nobody talks about, which is how many times you have to run a prompt before the answer means anything at all. That last number is calculable, and it is much larger than you would guess.
- AI share of voice is your brand’s mentions divided by all brand mentions across a set of AI answers. The formula is trivial. The prompt set is everything.
- No two tools use the same denominator, so two vendors can hand you 18% and 44% for the same brand in the same week and both be right.
- The same prompt does not give the same answer twice. SparkToro found under a 1 in 100 chance of two identical brand lists across repeated runs.
- You need roughly 81 runs per prompt for a number good to plus or minus 10 points, and about 323 runs for plus or minus 5.
- Treat it as a quarterly trend, never a weekly score. A month-on-month move smaller than about 14 points is statistically indistinguishable from nothing.
What is AI share of voice?
AI share of voice is the percentage of AI-generated answers in your category that mention your brand, measured against every brand mentioned in those same answers. If you run 100 category questions through ChatGPT and the answers name brands 400 times in total, and 60 of those are you, your AI share of voice is 15%.
It borrows its name from the old media metric, where share of voice meant your slice of total category advertising spend or press coverage. The logic transfers cleanly. The measurement does not, because a magazine ad either ran or it did not, and an AI answer is generated fresh every time somebody asks.
This is a different question from two neighbors it gets confused with. Prompt volume asks how many people are asking a question in the first place. AI visibility usually means a yes or no on whether you appear at all. Share of voice is comparative. It only exists relative to the other brands in the answer, which is why it is the one executives like and the one that is hardest to compute honestly.
The formula, and the part nobody publishes
Semrush states the calculation plainly in its own guide to the metric, published July 17, 2026. Your AI mentions divided by total AI mentions across all brands in your category, times 100. Its Enterprise product then adds a twist, weighting by the topic’s search volume, which means Semrush’s two products can produce two different numbers for the same brand.
That is the whole formula. It is not where the difficulty lives.
The difficulty lives in the denominator, and Dan Taylor named it exactly in Search Engine Land on June 8, 2026. In classic rank tracking the denominator is a keyword list you chose and can see. In AI search the universe of possible prompts is effectively infinite, so every vendor picks an arbitrary slice of it. Taylor’s phrase for the result is a hidden denominator, and his charge is that these tools “obscure their denominator within proprietary, vendor-defined systems that are almost certainly incomplete.”
Here is what that means in practice. There are at least three different things a tool can put on the bottom of the fraction.
| Denominator | What it measures | Who it flatters | Failure mode |
|---|---|---|---|
| Total brand mentions | Your mentions as a share of every brand named across the prompt set | Brands in crowded categories with one dominant rival | Adding a fourth competitor to the tracked list cuts everybody’s number |
| Answers containing you | The share of answers you appear in at all, ignoring how many rivals appear beside you | Anyone, because the numbers come out much higher | Not a share of anything. Two brands can both score 90% |
| Volume-weighted mentions | Mentions weighted by how often the underlying topic is searched | Brands that win the few high-volume head questions | Inherits every error in the volume estimate, which is itself modeled |
None of these is wrong. They answer different questions. But a report that says “AI share of voice: 31%” without saying which one it used has not told you anything you can act on, and it certainly cannot be compared against a number from a different tool.
Why the same question gives you a different number
Even if you and I agreed on the prompt set, we would still get different answers, because the engines themselves are not stable.
SparkToro ran the experiment properly and published it on January 28, 2026. Rand Fishkin recruited 600 volunteers who ran 2,961 individual queries across 12 prompts on ChatGPT, Claude and Google’s AI Overviews, with each prompt run 60 to 100 times per platform. The headline finding is that there is under a 1 in 100 chance that two runs of the same prompt return the same list of brands. Ask for the same list in the same order and it is closer to 1 in 1,000.
The mechanism is not marketing mystique, it is hardware. Setting temperature to zero does not make a model deterministic. An ICLR 2026 blog post dissecting non-determinism in LLMs traces it to floating-point arithmetic being non-associative on GPUs combined with dynamic batching in inference servers, and states the consequence bluntly. A request’s output depends on server load and on what other users sharing that GPU were doing at that exact millisecond.
Sit with that for a second. Part of your measured AI share of voice is a function of how busy OpenAI’s servers were when your tool ran its checks.
The same SparkToro work contains the reassuring half, and it matters just as much. Across nearly 1,000 headphone responses, the top brands (Bose, Sony, Sennheiser and Apple) appeared in 55% to 77% of answers. The consideration set is stable even though the list is not. Which brands are in the pool is a real signal. Their exact percentages on any given day are mostly weather.
How many runs does it take before the number is real?
This is the question the tool vendors do not answer, and it has an actual arithmetic answer.
Every prompt run is a coin flip. Either your brand gets named or it does not. So the margin of error on a share of voice figure is the standard margin of error on a proportion, and you can solve it for the sample size you need.
Assume your true appearance rate is somewhere around 30%, which is typical for a brand doing reasonably well in a category with a handful of rivals. Here are the run counts you need at 95% confidence.
- Plus or minus 10 points: about 81 runs of that prompt.
- Plus or minus 5 points: about 323 runs.
- Plus or minus 3 points: about 897 runs.
Those are per prompt. If your prompt set has 20 questions in it and you want a figure good to five points, that is roughly 6,500 engine calls per measurement, per engine.
Now do the comparison math, because that is what people actually use the number for. Comparing this month against last month means comparing two independent proportions, and the margin of error on the difference is about 1.4 times the margin on either one. With 81 runs per prompt in each month, a month-on-month move has to clear roughly 14 percentage points before you can say it is real rather than noise.
Fourteen points. Most dashboards report movements of two or three and draw an arrow.
I am not claiming the tools are lying. I am claiming that the sampling cost of a precise number is high enough that almost nobody is paying it, and that the honest response is to stop reading the small movements. If your share of voice went from 22% to 25%, you learned nothing. If it went from 22% to 41% over two quarters, you learned something.
Does rewording the prompt change your share of voice?
Yes, but less than the re-run problem, and this is the one place where the pessimism gets overdone.
Peec AI ran the largest study on this I could find, covered by Malte Landwehr in Search Engine Journal on June 15, 2026. It analyzed 1,754 prompts and 37,804 AI responses across five platforms and 18 sub-verticals. The finding is that human phrasing variation is far tamer than marketers assume. Most human-written prompts clustered tightly together, and brand visibility stayed stable across those variations. It only fell off, by roughly half, once prompts drifted into the bottom similarity band.
Two things in that study do move the number, and both are things you control when you build a prompt set.
- Prompt format: asking for a ranked list surfaces around 20% more brands than an open question. If your set is all “best X” listicle prompts, your denominator is inflated and everyone’s share looks smaller.
- Prompt style: keyword-style prompts produced up to 25% more visibility than conversational ones. A prompt set written by an SEO looks nothing like a prompt set written by a customer.
So the practical rule is that paraphrasing is safe and format is not. Mix ranked-list and open-question prompts deliberately, in whatever ratio matches how your buyers actually ask, and keep that ratio fixed between measurements. Changing your prompt format between quarters will produce a movement that has nothing to do with your brand.
The four numbers a share-of-voice report has to publish
Here is the standard I would hold any tool to, including ours. A share of voice figure is only interpretable if it arrives with four things attached.
- The prompt set: the actual questions, or at minimum the count and how they were chosen. A number without this cannot be reproduced or compared.
- The run count per prompt: how many times each question was asked. This is what sets the margin of error, and it is the number vendors are quietest about.
- The engines and dates: which models, which surfaces, over what window. ChatGPT and Perplexity cite substantially different sources, so a blended figure hides more than it shows.
- The denominator definition: which of the three variants above was used.
If a report gives you a percentage and none of those four, treat the percentage as decoration. That is a strong claim and I will stand behind it. The metric is not fake, but a metric you cannot reproduce is a vibe with a decimal point.
How to measure AI share of voice by hand
You can do a defensible version of this for free in an afternoon, and it is worth doing once so the numbers stop feeling abstract.
- Write 10 to 20 prompts a real buyer would type before choosing in your category. Mix formats deliberately. Some ranked-list asks, some open questions.
- Pick your competitor set and freeze it. Adding a brand later changes every historical number. Write the list down.
- Run each prompt 10 times per engine, in fresh sessions, across ChatGPT, Claude and Perplexity. This is nowhere near 81, so be honest that your margin is wide.
- Tally every brand named, not just yours. The denominator is the whole point.
- Record the date, the run count and the prompt set alongside the percentage. Future you will not remember.
Run it once and you have a baseline. The trouble starts next quarter, when you have to do the whole thing again identically, by hand, from memory. Almost nobody does. That is where the trend dies, and the trend was the only part with real value.
How to measure it inside your agent
This is what we built, so weigh it accordingly.
ContextBolt SEO is a hosted MCP server. You paste one URL into Claude Code, Claude Desktop or Cursor, and then ask SEO questions in plain language. There is no dashboard and no app. It costs $35 a month flat, includes 1,000 credits, and there is a 7-day free trial with a 100-credit cap.
Three of its tools are relevant here. ai_share_of_voice compares you against named rivals across a topic. ai_visibility shows whether you surface at all and which of your pages get cited. ai_answer runs a real prompt through ChatGPT, Claude, Gemini or Perplexity and hands back the literal answer plus every source it cited.
The honest framing matters more on this product than any other we sell. The share of voice and visibility numbers come from a modeled dataset, derived from an index of AI responses rather than a live log of every conversation. It is directionally accurate and good for trends. It is not a ledger of what ChatGPT said to your customers. The ai_answer tool is the ground-truth one, because it actually runs the prompt and shows you the reply. Two layers, and copy that pretends otherwise is selling you the hidden denominator all over again.
What genuinely helps is the memory. Every lookup writes to a ./seo-findings/ folder as markdown on your own machine, so the next check can be compared against the last one and your agent will tell you what moved. That is the part the by-hand method never survives, and it is why this is worth doing in an agent rather than a spreadsheet. The wider workflow, including how this sits next to your Google data, is in AI search visibility with Claude.
Should AI share of voice be your main metric?
No, and the most useful critique of it comes from the same Search Engine Land piece. Dan Taylor argues for three metrics he considers more honest than a blended share number: share of mentions across the sources models actually learn from, share of recommendations when someone explicitly asks for a shortlist, and share of narrative, meaning what gets said about you when you do appear.
That third one is the sharp one. High visibility inside a bad frame is worse than invisibility. Being named in 40% of answers as the cheap option that nobody uses at scale is a number going the right way and a business going the wrong one.
My own read after building tooling for this. AI share of voice is a real metric with a real use, which is telling you whether you are in the consideration set at all, tracked over quarters, against a frozen prompt set and a frozen competitor list. That is genuinely worth knowing and most brands have never checked. What it is not is a weekly KPI, a thing to compare across vendors, or a number to put in a board deck without its four attachments.
Measure it. Publish your denominator. Ignore anything under fourteen points. And spend the time you save on getting cited rather than on watching the percentage.