Yes, you can ask it yourself. that is not the hard part.
157 companies came back ranked. The one paying attention was in none of them.
Fifteen real buying questions, three engines, thirty answers. Verified across every word of captured text: zero occurrences of the brand doing the measuring. Not ranked low. Absent from the consideration set.
The two questions closest to a home fixture were the clearest. Asked which agencies combine exactly the three services it sells, ChatGPT and Gemini each returned a shortlist, and the two lists had not one firm in common. Neither included the brand. ChatGPT even offered to extend its list to Latin America, with the door held open, and it still did not surface.
That is the finding a single manual check does give you, and it is worth having. Go and get it: the free scan does exactly this and costs nothing. Everything below is about the next question, the one that decides whether anything changes. Is that real, and is it moving?
Two runs of the identical question, minutes apart, agree on 36% of the brands they name
Same engine, same model, same market, same wording. Only the clock moved. Some pairs of runs shared nothing at all. Everything that differs between two answers like that is the instrument, not the market.
36%
Agreement between two runs of one question, minutes apart. Some pairs: zero.
65 of 82
Brands whose swing between repeats was at least as large as their own average share. For four in five, the noise is bigger than the number.
4.3 pts
The smallest change two measurements can tell apart. Under that, a chart is drawing noise with a confident line.
So the check you ran told you one of the answers, not the answer. Run it again next month, see a different set, and you cannot say whether your market moved, your work paid off, or the dice landed differently. Repetition is not thoroughness here. Without it there is no measurement at all, and every number built on a single ask is a number that will contradict itself the next time anybody looks.
A VPN changes your address. it does not change who the engine thinks you are.
Your account is inside the answer
Gemini was given exactly one fact: “I run a seed stage AI company.” It came back describing the specific product, the payment-recovery use case and the region, none of it in the prompt. That is account data, not inference. Turning off Memory does not stop it, because a separate surface keeps Gmail, Drive, Calendar and the rest feeding it. Disconnecting them and asking again gave zero overlap with the personalised run.
A browser location is not a market
Running from a Colombian locale leaked into the results: Colombian and Puerto Rican LinkedIn subdomains in the citations, a Philippine storefront in another, Google’s own links resolving to a Colombian domain. Locale, timezone and the location on the account all travel with the question. Change the exit node and every one of them stays exactly where it was.
You are not told which model answered
The consumer interfaces expose an intelligence setting, not a model id. In the study the browser and the API were not even the same model family or tier on one engine. Two checks a month apart can differ because the vendor changed the default underneath you, and nothing on the screen would say so.
One engine is not “AI”
Of 138 citations captured, exactly two URLs appeared on more than one engine. Of 157 brands ranked, three were named by all three. The engines are reading different internets, so winning on the one you happen to use tells you close to nothing about the others.
| same question, same hour | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| First | Demand Curve | NoGood | Metaflow AI |
| Second | NoGood | RZLT | NoGood |
| Third | Omniscient Digital | Daydream | Kalungi |
| Brands named | 9 | 6 | 13 |
Three engines, one question, one hour, three different winners.
About ninety hours a month, by hand
Not a rhetorical number. Covering what a mid tier covers, meaning forty questions across three engines, sampled often enough to produce a band and refreshed weekly, is roughly 3,600 answers a month. At ninety seconds each to type it, wait for it, read it, and write down who was named and in what order, that is close to ninety hours. Every month, or the series has a hole in it.
Cut it to something a person would actually sustain, say ten questions across three engines with three repeats, and it is a couple of hours a month. Three repeats is the bare floor below which a band cannot be computed at all, and you still have no brand matching, no denominator, no history and no evidence trail. You have a spreadsheet of impressions.
And that assumes the accounts cooperate. In the study Gemini failed twice in fifteen, once on a permission gate and once on a three-minute hang, and Perplexity stopped at question two behind a signup wall. Minor on their own. They matter because every one of them is a hole you have to notice, and telling a hole apart from a real absence is the entire job.
Asking is the cheap part. the measurement is the product.
A number with a band, or no number at all
Every question is asked several times in one run, and the figure arrives with the spread those repeats produced. Below half the samples the percentage is withheld entirely and the screen says why. You are never handed a number precise enough to act on and too noisy to be true.
Movement only when it is movement
The wobble was measured, so anything smaller than it is reported as no change. That is the difference between a dashboard that always has news and one you can take to a board.
Per engine, never averaged
Three of 157 brands were named by all three engines. An average across them is arithmetic performed on unrelated things, so the number is reported per engine, and a blend is labelled as one when it happens.
The pages that are actually being read
This is the part you cannot get by asking. The engines assemble these answers out of a small set of listicles, and the study caught it cleanly. On one question, all ten of Perplexity’s sources were agency-published listicles, four of the recommended firms had published the source cited for them, and one was cited on thirteen separate claims including claims about its competitors. You get that list, and whether you are on it.
A receipt under every figure
Each mention keeps the passage it came from, its position in the list, the engine and the date. When somebody asks where the number came from, you open it rather than defend it.
A history that cannot be backfilled
There is no archive of what an engine said last month, and nobody can sell you one. The series starts the day it starts, which is the only part of this that gets more expensive by waiting.
And what we do about each one
These are real and they are not ours to fix. They are also the reason the measurement has to be built rather than eyeballed: every one of them is a way a manual check goes wrong silently.
Nobody can fully depersonalise Gemini
With its connectors off, one prompt in the study refused to answer at all and demanded they be switched back on. So we run it as clean as it can be run, label the reading, and never let it disappear into a single blended number.
The consumer surface does not name its model
We cannot make it. What we can do is record which surface every answer came from, and refuse to pool a model-pinned reading with an unpinned one, so a change you see is not two different instruments being averaged.
Some questions cannot be won by anyone
One prompt in fifteen was definitional on both engines and named zero brands. We say so and take it out of the share, rather than quietly reporting a number measured on a question with no commercial surface.
The instrument still wobbles
Sampling shrinks the uncertainty and reports it. It does not abolish it. We tell you the size of the wobble because knowing it is what separates a measurement from a screenshot, and because a tool that hides it will show you movement that was never there.
Check it yourself first
Genuinely. Ask the engines about your category and see who comes back. If your name is in every answer, you do not need us. If it is not, and in this study 157 companies were named while one was not, the question stops being whether you are visible and becomes which pages are being read, and what it takes to be on them.
Every figure here comes from two studies against live engines: a fifteen-question browser capture across ChatGPT, Gemini and Perplexity, and a repeat-sampling run that asked one question ten times on one surface. The method page carries both in full, including the parts that did not work.