Skip to content
doing it yourself

Yes, you can ask it yourself. that is not the hard part.

The hard part is that the answer changes when you ask again, it is not the answer your buyer gets, and it cannot be compared to the one you got last month. We measured all three, by hand, before building anything.
what the check found

157 companies came back ranked. The one paying attention was in none of them.

Fifteen real buying questions, three engines, thirty answers. Verified across every word of captured text: zero occurrences of the brand doing the measuring. Not ranked low. Absent from the consideration set.

The two questions closest to a home fixture were the clearest. Asked which agencies combine exactly the three services it sells, ChatGPT and Gemini each returned a shortlist, and the two lists had not one firm in common. Neither included the brand. ChatGPT even offered to extend its list to Latin America, with the door held open, and it still did not surface.

That is the finding a single manual check does give you, and it is worth having. Go and get it: the free scan does exactly this and costs nothing. Everything below is about the next question, the one that decides whether anything changes. Is that real, and is it moving?

one check is a sample of one

Two runs of the identical question, minutes apart, agree on 36% of the brands they name

Same engine, same model, same market, same wording. Only the clock moved. Some pairs of runs shared nothing at all. Everything that differs between two answers like that is the instrument, not the market.

36%

Agreement between two runs of one question, minutes apart. Some pairs: zero.

65 of 82

Brands whose swing between repeats was at least as large as their own average share. For four in five, the noise is bigger than the number.

4.3 pts

The smallest change two measurements can tell apart. Under that, a chart is drawing noise with a confident line.

So the check you ran told you one of the answers, not the answer. Run it again next month, see a different set, and you cannot say whether your market moved, your work paid off, or the dice landed differently. Repetition is not thoroughness here. Without it there is no measurement at all, and every number built on a single ask is a number that will contradict itself the next time anybody looks.

not the answer your buyer gets

A VPN changes your address. it does not change who the engine thinks you are.

Your account is inside the answer

Gemini was given exactly one fact: “I run a seed stage AI company.” It came back describing the specific product, the payment-recovery use case and the region, none of it in the prompt. That is account data, not inference. Turning off Memory does not stop it, because a separate surface keeps Gmail, Drive, Calendar and the rest feeding it. Disconnecting them and asking again gave zero overlap with the personalised run.

A browser location is not a market

Running from a Colombian locale leaked into the results: Colombian and Puerto Rican LinkedIn subdomains in the citations, a Philippine storefront in another, Google’s own links resolving to a Colombian domain. Locale, timezone and the location on the account all travel with the question. Change the exit node and every one of them stays exactly where it was.

You are not told which model answered

The consumer interfaces expose an intelligence setting, not a model id. In the study the browser and the API were not even the same model family or tier on one engine. Two checks a month apart can differ because the vendor changed the default underneath you, and nothing on the screen would say so.

One engine is not “AI”

Of 138 citations captured, exactly two URLs appeared on more than one engine. Of 157 brands ranked, three were named by all three. The engines are reading different internets, so winning on the one you happen to use tells you close to nothing about the others.

same question, same hourChatGPTGeminiPerplexity
FirstDemand CurveNoGoodMetaflow AI
SecondNoGoodRZLTNoGood
ThirdOmniscient DigitalDaydreamKalungi
Brands named9613

Three engines, one question, one hour, three different winners.

what doing it properly takes

About ninety hours a month, by hand

Not a rhetorical number. Covering what a mid tier covers, meaning forty questions across three engines, sampled often enough to produce a band and refreshed weekly, is roughly 3,600 answers a month. At ninety seconds each to type it, wait for it, read it, and write down who was named and in what order, that is close to ninety hours. Every month, or the series has a hole in it.

Cut it to something a person would actually sustain, say ten questions across three engines with three repeats, and it is a couple of hours a month. Three repeats is the bare floor below which a band cannot be computed at all, and you still have no brand matching, no denominator, no history and no evidence trail. You have a spreadsheet of impressions.

And that assumes the accounts cooperate. In the study Gemini failed twice in fifteen, once on a permission gate and once on a three-minute hang, and Perplexity stopped at question two behind a signup wall. Minor on their own. They matter because every one of them is a hole you have to notice, and telling a hole apart from a real absence is the entire job.

what you get instead

Asking is the cheap part. the measurement is the product.

A number with a band, or no number at all

Every question is asked several times in one run, and the figure arrives with the spread those repeats produced. Below half the samples the percentage is withheld entirely and the screen says why. You are never handed a number precise enough to act on and too noisy to be true.

Movement only when it is movement

The wobble was measured, so anything smaller than it is reported as no change. That is the difference between a dashboard that always has news and one you can take to a board.

Per engine, never averaged

Three of 157 brands were named by all three engines. An average across them is arithmetic performed on unrelated things, so the number is reported per engine, and a blend is labelled as one when it happens.

The pages that are actually being read

This is the part you cannot get by asking. The engines assemble these answers out of a small set of listicles, and the study caught it cleanly. On one question, all ten of Perplexity’s sources were agency-published listicles, four of the recommended firms had published the source cited for them, and one was cited on thirteen separate claims including claims about its competitors. You get that list, and whether you are on it.

A receipt under every figure

Each mention keeps the passage it came from, its position in the list, the engine and the date. When somebody asks where the number came from, you open it rather than defend it.

A history that cannot be backfilled

There is no archive of what an engine said last month, and nobody can sell you one. The series starts the day it starts, which is the only part of this that gets more expensive by waiting.

what nobody solves

And what we do about each one

These are real and they are not ours to fix. They are also the reason the measurement has to be built rather than eyeballed: every one of them is a way a manual check goes wrong silently.

Nobody can fully depersonalise Gemini

With its connectors off, one prompt in the study refused to answer at all and demanded they be switched back on. So we run it as clean as it can be run, label the reading, and never let it disappear into a single blended number.

The consumer surface does not name its model

We cannot make it. What we can do is record which surface every answer came from, and refuse to pool a model-pinned reading with an unpinned one, so a change you see is not two different instruments being averaged.

Some questions cannot be won by anyone

One prompt in fifteen was definitional on both engines and named zero brands. We say so and take it out of the share, rather than quietly reporting a number measured on a question with no commercial surface.

The instrument still wobbles

Sampling shrinks the uncertainty and reports it. It does not abolish it. We tell you the size of the wobble because knowing it is what separates a measurement from a screenshot, and because a tool that hides it will show you movement that was never there.

start here

Check it yourself first

Genuinely. Ask the engines about your category and see who comes back. If your name is in every answer, you do not need us. If it is not, and in this study 157 companies were named while one was not, the question stops being whether you are visible and becomes which pages are being read, and what it takes to be on them.

Read the method

Every figure here comes from two studies against live engines: a fifteen-question browser capture across ChatGPT, Gemini and Perplexity, and a repeat-sampling run that asked one question ten times on one surface. The method page carries both in full, including the parts that did not work.

free scan

are you in the answer?

Five real questions against a live answer engine with web search on, each asked ten times. That is fifty real calls, so the report lands in under ten minutes rather than while you wait. You see every question before anything runs, and you get the answer text, who got named and the sources the engine read, not a score.

your site/what you sell/where your buyers are/the questions