Skip to content
Writing
the measurement

Counting mentions is easy. Counting them honestly is not.

Every decision in this method makes the resulting number smaller. That is not modesty — each one exists because the alternative produces a figure that falls apart the first time somebody checks it.

Published 2026-08-20 | 9 minute read

Share of voice in AI answers is the proportion of a category’s answers in which a brand is named, adjusted for where it is named and how commercial the question is. That sentence hides six decisions, and the decisions are the whole method.

One: the question set is the measurement

A set written to flatter a brand will flatter it. Questions have to be the ones a buyer types, in their words, spread across the intents that actually occur — problem-shaped, category-shaped, comparison-shaped, and local if the category has a geography.

Freeze the wording. A question reworded between readings has started a new series and ended the old one, and it usually happens without anybody noticing.

Two: match whole words, never substrings

This is the largest single source of fake numbers in the category. We measured a brand whose name is also an ordinary word returning 2,420 mentions from an index while the brand’s own domain returned zero on the same query. Three orders of magnitude, and the error always flatters the customer, which is why it survives.

Three: resolve aliases before you count

A brand that appears under several variants is counted as several entities, and the parts do not add up. In a recent run a business appeared under three spellings of itself, each landing separately — including the one answer that recommended it first for its speciality, which never registered against the brand at all.

Alias splitting under-reports, which means it is the one error nobody complains about. Check it before you trust a low number, not after.

Four: exclude what the buyer never sees

Reasoning models return a private draft alongside the reply, and that draft names brands the reply then drops. Counting it reports visibility in a place no customer reads. Exclude it, and record how many items you dropped so the exclusion is auditable.

Five: a failed call is not a zero

A call that errored is excluded from both sides of the ratio, and the coverage is declared next to the number. Scoring failures as absences turns every provider outage into a competitor’s good week. An answer that named nobody is also not a loss: it means the question did not discriminate, and a set full of those is a set that needs rewriting.

Six: every named brand counts

Including ones nobody listed. A share of voice that improves because you stopped tracking a rival is corrupt, and it is easy to produce by accident. In one run, six of the ten most-named entities in a category were not on anybody’s competitor list at the start.

Weighting, and why counts alone mislead

Two brands named in all four answers are not equally visible if one is always first and the other is always fifth. In a run we published, two brands were both named in four of four and their weighted share still differed by eleven points, because one landed higher in every list. That gap is the entire reason position enters the calculation.

Commercial intent is the second weight. Being named in "what is an issue tracker" is worth less than being named in "best issue tracker for a small team", and a method that treats them equally lets a brand look healthy on questions nobody buys from.

When a difference is reportable

Only when the confidence intervals of the two readings do not overlap, which requires at least three runs per reading. In a category we measured, a brand at 17.6% weighted share carried an interval from 15.1% to 20.4% across ten samples with nothing changing. Any tool reporting a four point move on that surface is reporting the width of its own noise.

A single measurement dressed as precision is how credibility gets spent.
The rules this product is built under

What the number is for

Not a scoreboard. The figure exists to make two things comparable: your category over time, and you against the set of brands the answers actually name. Everything else on the report — the stored answers, the cited sources, the discovered competitors — is what makes the figure checkable, and a share of voice published without them is an assertion in a percentage costume.

questions

The short answers

What is share of voice in AI answers?

The proportion of a category’s answers in which a brand is named, adjusted for where in the answer it appears and how commercial the question is. The adjustments matter: two brands named in every answer can differ by eleven points of weighted share because one lands higher in every list.

How do you calculate AI share of voice?

Ask a frozen set of buyer-shaped questions across the engines that matter, count whole-word brand matches in the reply only, resolve name variants first, exclude failed calls from both sides of the ratio, count every brand named including unlisted ones, then weight by position in the answer and by the commercial intent of the question.

Why not just count mentions?

Because a mention in fifth place is not a mention in first, and a mention on a definitional question is not a mention on a buying question. Raw counts also hide alias splitting and substring noise, both of which move the number by orders of magnitude.

What is the biggest source of wrong numbers?

Substring matching on brand names. We measured a brand that is also an ordinary word returning 2,420 mentions from an index while the brand’s own domain returned zero on the same query.

Should the reasoning trace be counted?

No. Reasoning models draft names the final reply then drops, and the draft is not what a buyer sees. Exclude it and record how many items were dropped so the exclusion can be audited.

How do you handle a failed engine call?

Exclude it from both sides of the ratio and declare the coverage next to the figure. Scoring failures as absences turns every outage into a competitor’s good week.

How many runs before the number means anything?

Three is the floor for reporting change, because there is no confidence interval below that. Ten samples gives a usable band — in one category we measured, a brand at 17.6% carried an interval of 15.1% to 20.4%.

What makes a share of voice figure trustworthy?

That it opens. Every figure should resolve to the question that produced it, the engine, the model, the date and the full answer text. A percentage with none of that underneath is an assertion in a percentage costume.

Find out in under ten minutes

One free check, no account and no card. You get the answers the assistants gave, not a score.

Every figure quoted here comes from a real question put to a real assistant, saved with the date and the assistant that answered it. This piece was written to answer the search "ai share of voice".

free scan

are you in the answer?

Five real questions against a live answer engine with web search on, each asked ten times. That is fifty real calls, so the report lands in under ten minutes rather than while you wait. You see every question before anything runs, and you get the answer text, who got named and the sources the engine read, not a score.

your site/what you sell/where your buyers are/the questions