Skip to content
Writing
ai recommendations

You asked and it named you. They asked and it did not.

Both answers are real. The variance is a property of the thing, not a fault in your check — and it is the reason a single look at your own screen tells you almost nothing.

Published 2026-08-20 | 5 minute read

You typed the question, the assistant named your business, and you closed the tab reassured. A customer typed something close to it an hour later and got two other companies. Neither of you saw a glitch. Assistants do not return a fixed answer to a fixed question, and until you know which parts move and which do not, every check you run is a coin you cannot read.

Four things vary, and only one is about you

  • The generation itself. Ask the same question twice in the same clean session and the names can differ. This is the model composing, not remembering.
  • Your account. History and stated preferences lean the reply toward what you have talked about before, which for a business owner is their own business.
  • Your location. A local question resolves against where the request appears to come from, and that is the single largest swing on any category with a geography.
  • The web underneath. Engines with search on read pages that changed since yesterday, so an answer can move because a listicle was updated by somebody else entirely.

Only the last one is a change in the world. The first three are the measurement instrument moving, and mistaking one for the other is how a quarter gets spent on the wrong thing.

How much does it actually move

This is where the category is mostly assertion. Everyone says answers vary; almost nobody publishes the size of it, because publishing it requires asking the same question many times and paying for every call.

In a run we did against a local hospitality category, the leading brand was named in 91 of 165 answers. Its weighted share of voice came out at 17.6%, and across ten samples the confidence interval ran from 15.1% to 20.4%. That is the shape of the problem in one line: a rival could report a five point improvement, believe it, and be looking at nothing at all.

Five points of movement inside a five point interval is not an improvement. It is the same reading, taken twice.

What does not vary

Enough to be worth measuring, which is the good news. Across repeated asks the same handful of names keep coming back, and the ordering of the top two or three is stable far more often than it is not. What moves is the tail: the fourth and fifth names, and whether a marginal brand is included at all.

So a brand that is never named across a well-built question set is not unlucky. And a brand named in nearly every answer is not having a good week. Both of those are findings. It is the middle of the distribution where a single reading is worthless.

Measuring through the noise

  • Ask from a clean context every time. No account history, and the location set to your customers rather than to you.
  • Freeze the wording. A question reworded between readings has started a new series and ended the old one, usually without anybody noticing.
  • Repeat. Three runs is the floor for saying anything about change; ten gives an interval tight enough to act on.
  • Compare intervals, not points. If the bands overlap, you have not measured a difference, whatever the two numbers look like side by side.

None of this is exotic statistics. It is the ordinary discipline of measuring anything noisy, applied to a surface most tools currently report as though it were a rank tracker.

questions

The short answers

Why does ChatGPT give different answers to the same question?

Four causes. The model composes rather than recalls, so repeat asks differ. Your account history and location both bend the reply. And engines with web search on read pages that change underneath them. Only the last is a change in the world; the other three are the instrument moving.

Why does ChatGPT name my business when I ask but not when my customer asks?

Almost always your account. Chat history and location make the reply friendlier to you than a stranger gets, which is why checking from your own logged-in session is the most common way to conclude you are visible when you are not.

How many times do I need to ask before the answer means anything?

Three is the floor for saying whether something changed, because there is no confidence interval under three runs. Ten samples gives a band tight enough to act on. In one of our runs a brand at 17.6% weighted share carried an interval from 15.1% to 20.4% across ten samples.

Does the whole answer change, or just parts of it?

Mostly the tail. The top two or three names are stable far more often than not; what moves is the fourth and fifth, and whether a marginal brand is included at all. That is why consistent absence and consistent presence are both real findings, and the middle is where a single reading is worthless.

Is a change of five percentage points significant?

Not on its own. If the confidence intervals of the two readings overlap you have not measured a difference, however far apart the two numbers look. A tool that draws an arrow after one rerun is drawing variance with a direction on it.

Find out in under ten minutes

One free check, no account and no card. You get the answers the assistants gave, not a score.

Every figure quoted here comes from a real question put to a real assistant, saved with the date and the assistant that answered it. This piece was written to answer the search "why does chatgpt give different answers".

free scan

are you in the answer?

Five real questions against a live answer engine with web search on, each asked ten times. That is fifty real calls, so the report lands in under ten minutes rather than while you wait. You see every question before anything runs, and you get the answer text, who got named and the sources the engine read, not a score.

your site/what you sell/where your buyers are/the questions