Every tool in this category hands you a number out of 100 within about ten seconds of meeting you. It is the first thing you see, it is the thing you screenshot, and it is the thing you will be asked about in the next meeting.
It is also the least useful output the tool produces, and the reason is simple. There is no agreed definition of what it measures, no published benchmark to compare it against, and no two vendors who calculate it the same way. You can score 31 in one tool and 68 in another on the same Tuesday, and neither number is wrong.
That does not make the score worthless. It makes it a trend line rather than a grade. Here is how the number is actually built, what a good one looks like when you strip out the marketing, and what to watch instead.
- An AI visibility score summarizes how often AI answers name your brand, usually on a 0 to 100 scale that blends mention volume with how many engines mention you.
- There is no industry benchmark. Any “average score” you are quoted is that vendor’s own user base, measured their own way.
- Two tools will disagree, and both can be right. Different prompt sets, different engines, different time windows, different formulas.
- The score is a trend, the questions are the work. AI visibility tracking in ContextBolt SEO shows the number next to the actual questions, answers and sources behind it, from inside Claude or Cursor, and keeps every run on a board with twelve months of history from day one. $35 a month.
- Compare against rivals in the same run, never against a number from a different tool or a different month’s methodology.
What the score is measuring
Strip the branding off and almost every AI visibility score is built from the same two ingredients.
Volume. How many AI answers, out of the set the tool looked at, named your brand. This is the bulk of the number.
Spread. How many different engines those mentions came from. Being named by one assistant is fragile. Being named by several is a position, because you are not exposed to one model’s next update.
Some vendors add sentiment, or position within the answer, or whether you were cited as a source as well as named. Those refinements move the number by a few points and none of them fix the underlying problem, which is the denominator.
The denominator is the whole ball game. A score is mentions divided by something. That something is a set of questions somebody chose. Change the question set and the score changes, with nothing at all having changed about your brand. Most vendors will not tell you what is in theirs, and the ones running a fixed prompt list are at least honest about the fact that you picked it.
This is the same issue we walked through in AI share of voice, and it applies with more force to a composite score, because the score buries the denominator under a scale as well.
Why two tools disagree
Four independent reasons, and they stack.
| What differs | What it does to your score |
|---|---|
| The question set | Broad category questions dilute you. Narrow buyer questions concentrate you. Same brand, different number. |
| The engines counted | A tool counting only ChatGPT and a tool counting six engines are measuring two different populations. |
| The time window | Twelve months of answers versus a current-state index gives different totals from the same source. |
| The scale | Linear scales spread the top out. Log scales spread the bottom out, so small brands see more movement. |
Semrush’s own guide to measuring AI share of voice walks through the first three of those and lands on the same conclusion, which is that the prompt set decides the number.
There is a fifth reason underneath all of them, and it is the one that should make you cautious about any single reading. AI answers are not deterministic. SparkToro’s testing found that two runs of the same prompt rarely return the same brand list. Run the same question twenty times and you get a distribution, not a value. Any tool showing you one decimal place is showing you precision it does not have.
Search Engine Land’s write-up on which AI share of voice metrics matter lands in the same place from the agency side. The composite is the least actionable thing on the screen.
So what is a good score?
The honest answer is that nobody knows, and I would be suspicious of anyone who tells you otherwise.
There is no published benchmark because there is no shared methodology to benchmark against. When a vendor tells you the average score in your industry is 34, they are telling you the average among companies who bought their product and configured it their way. That is a customer base, not an industry.
What you can do instead is three comparisons that hold up.
Against your competitors, in the same run. This is the only fair one. Same tool, same questions, same window, same day. If you are at 18 and the category leader is at 46, that gap is real, because everything except the brand was held constant.
Against yourself last month. Same setup, one variable changed, which is time. This is where the score earns its keep, and it is why a tool that backdates your history is worth more than one that starts counting the day you sign up.
Against the count, not the score. “Named in 4 of my 12 buyer questions” is a fact. “Score 31” is a summary of that fact with information removed. Where a tool gives you both, use the count in any decision and keep the score for the chart.
If you want a rough sense of the shape rather than a benchmark: a brand nobody has written about scores zero, a small company with a handful of comparison pages scores somewhere in the teens to thirties, and the household name in a category sits near the ceiling because the scale compresses at the top. That is a description of a curve, not a target.
How ours is calculated
We publish a score in our free AI visibility checker, so it would be poor form to complain about opaque scoring and not show our own.
It is two terms. Mention volume on a log scale, plus a bonus for engine spread.
score = 22 x log10(mentions + 1) + 8 x (number of platforms, capped at 6)
The result is clamped to 0 to 100, and zero mentions returns a hard zero rather than a floor. In practice that means:
| Mentions found | On 1 platform | On 2 platforms |
|---|---|---|
| 0 | 0 | 0 |
| 5 | 25 | 33 |
| 100 | 52 | 60 |
| 1,000 | 74 | 82 |
| 5,000 | 89 | 97 |
Two things follow from that, and both are limitations rather than features.
The log scale flatters the bottom. Going from 5 mentions to 100 is a twenty-fold increase and moves the score 27 points. Going from 1,000 to 5,000 is a five-fold increase and moves it 15. That is deliberate, because early movement is what a small site needs to see, but it means a 25 and a 52 are much further apart in reality than they look.
Spread is worth a lot, on purpose. Eight points per platform is a large share of the range. Being named by two engines rather than one is genuinely more durable than doubling your mentions on one, and the score says so loudly.
Our aggregate count covers ChatGPT and Google AI answers, the two engines with an answer log large enough to count across. Perplexity, Claude and Gemini publish nothing to count, so they can only be asked live, one question at a time. Any score claiming to count them monthly is modeling that number rather than measuring it.
Three questions to ask any vendor
Before you take a score seriously, ask the person selling it these three. The answers take a minute and they tell you more than any demo.
“What questions did you measure?” If the answer is a prompt list you chose, fine, and write it down so next quarter’s number is comparable. If the answer is a proprietary set they will not describe, the score cannot be reproduced by you or by anyone, and it is a subscription metric rather than a market one.
“Which engines are counted, and which are asked live?” These are different things and vendors blur them constantly. Counting means measuring across an index of many answers. Asking live means running one prompt and reading the reply. A tool can honestly say it “covers” six engines while counting two of them, and the two numbers behave completely differently over time.
“What time window is each figure from?” A twelve-month total and a current-state reading, drawn from the same index, will not match. If two figures on the same screen disagree and nothing explains why, that is usually the reason, and a dashboard that does not label the window is asking you to reconcile numbers it never reconciled itself.
A vendor who answers all three plainly is worth more than one with a prettier chart. This is also the fastest way to work out whether you are looking at a measurement or a marketing asset.
What to watch instead
The named list. Which brands appear in the answers for your category, in order. This is the reading list. Whoever leads is being retrieved constantly and their pages are public.
Your cited pages. Which of your URLs get pulled in as sources, and how often. Your most quoted page is your template, and most sites are surprised by which one it turns out to be.
The individual verdicts. For each buyer question, are you named, cited without being named, or absent? Those three states need three different fixes, and a score averages all three into one digit. The middle one is the nastiest and the least discussed, which is why it has its own post.
The market line. How many answers are being written about your topic at all, month over month. Flat mentions in a growing market is a loss, even though your number went up. This is the comparison most dashboards leave out and it changes how you read every other figure on the page.
The practical version
Look at the score once a month. Look at the questions every time you publish something.
If you want both in the same place, AI visibility tracking inside ContextBolt SEO puts the number next to the questions, the full answers, the sources and twelve months of history from the first run, so the score is a headline over evidence rather than a thing on its own. It runs from Claude, Cursor or Codex for $35 a month alongside keyword research, SERPs, backlinks and your own Search Console, which is less than most dedicated trackers charge for the visibility half by itself. We compared eight of those in best AI visibility tools, including the ones that do this better than we do.
Whatever you use, write down the questions you measured. A year from now, the score will be meaningless without them.