Skip to content
Writing
ai recommendations

Four assistants, four different answers.

They disagree with each other considerably more than they disagree with themselves — which means a reading from one of them is a reading about one of them, and your buyers are not all in the same place.

Published 2026-08-20 | 5 minute read

Ask the same buying question of four assistants and you get four overlapping but distinct shortlists. Some of that is the model. Most of it is what each one reads before answering, and how much it reads.

What separates them in practice

  • Whether web search is on by default, and how many pages it pulls. This is the largest single difference and it changes month to month.
  • How heavily the answer leans on aggregators and review sites versus primary pages.
  • How many names the answer is willing to give. Some default to three, some to six, and the length of the list decides whether a marginal brand appears at all.
  • How much of the reasoning is discarded before the reply. Reasoning models draft names they then drop, and only the reply is what a buyer sees.

That last one has a measurable cost. We exclude the private draft from every count and record how many items we dropped, because counting it reports a brand as visible in a place nobody reads.

What this means for measuring

A single-engine reading is a reading about that engine. It is not wrong; it is narrow, and it is usually reported as though it were general.

The practical floor is two engines with different retrieval behaviour, and four is where the picture stops moving much. In a run we did across two of them, the same brand was absent from almost every answer on one engine and named as the first option for its speciality on the other. Either engine alone would have produced a confident, wrong summary.

The disagreement is the finding, not a problem with the measurement. A category where all four agree is a category with a settled answer, and that is worth knowing too.

Which one should you care about

The honest answer is the one your buyers use, and almost nobody knows that number for their own category. In the absence of it, weight by reach and measure at least two. What you should not do is pick the assistant that gives you the friendliest answer and report that.

questions

The short answers

Do different AI assistants recommend different businesses?

Yes, and they disagree with each other considerably more than they disagree with themselves. The main causes are whether web search is on, how heavily each leans on aggregators, and how many names each is willing to list.

Which AI assistant should I track?

The one your buyers use, which almost nobody knows for their own category. Failing that, measure at least two with different retrieval behaviour. One engine is a reading about one engine.

Why does one assistant name me and another does not?

Usually retrieval rather than judgement: they read different pages, and the length of the list each is willing to give decides whether a marginal brand appears at all. We have measured a brand absent from nearly every answer on one engine and first for its speciality on another.

Do reasoning models change the answer?

They add a private draft that names brands the reply then drops. That draft is not what a buyer sees, so counting it reports visibility that does not exist. We exclude it and record the number of items dropped.

Is it enough to check one assistant if it is the biggest?

Only if you are willing to describe the result as being about that assistant. Reported as general visibility it is the single most common way a confident wrong summary gets produced.

Find out in under ten minutes

One free check, no account and no card. You get the answers the assistants gave, not a score.

Every figure quoted here comes from a real question put to a real assistant, saved with the date and the assistant that answered it. This piece was written to answer the search "which ai recommends businesses".

free scan

are you in the answer?

Five real questions against a live answer engine with web search on, each asked ten times. That is fifty real calls, so the report lands in under ten minutes rather than while you wait. You see every question before anything runs, and you get the answer text, who got named and the sources the engine read, not a score.

your site/what you sell/where your buyers are/the questions