Thicket.
← Back to Journal

Measurement

Half the Vendor Names an AI Visibility Check Returns Don't Survive a Second Run

We had a finding ready to send to three companies: their competitive set had reshuffled over three weeks. Then we ran the same questions again twenty minutes later, and it had reshuffled by exactly as much.

Asked the same buyer-intent question twice in one morning, a language model returns a meaningfully different list of vendors. Only 53% of names appear in both readings. Two readings three weeks apart differ by no more than two readings twenty minutes apart — mean symmetric difference 3.50 against 3.67, sign test p = 1.000 over twelve questions. Absence, by contrast, was perfectly reproducible: the three companies we tested went unnamed in all 78 runs.

Paired dot chart of twelve questions, showing how many vendor names differed between two readings taken the same morning versus two readings taken three weeks apart. The two distributions overlap almost completely.
Each row is one question. Neither colour is consistently further right. That is the whole result.

Why we ran it

We sell a diagnosis: when your buyer asks an AI assistant for a recommendation in your category, does it name you? To make that claim we run the question against a model several times and record every vendor it names, verbatim. Three companies had received that audit from us in early August. Each had asked for it, read it, and gone quiet.

Following up on silence requires something new to say, so we re-ran their exact recorded questions — thirteen of them, three runs each, same model — seventeen to twenty days later. The result looked excellent. Plausible, Fathom and TelemetryDeck had dropped out of the privacy-analytics answer and Umami had entered. Pocket, Cubox and Mailbrew had left the read-later answer, replaced by KTool and Upnext. Nine names had turned over across four questions for one company alone.

The story wrote itself: your category is reshuffling every month and you are outside it in both readings. It is a finding, it is alarming, and it happens to be an argument for buying a monthly monitoring subscription rather than a one-off audit. That last property is exactly why it deserved a control before it deserved an email.

The control

We ran the same thirteen questions a second time, the same morning, roughly twenty minutes after the first batch. Nothing else changed: same model, same prompts, same three runs each. If three weeks of market movement were driving the turnover, the control should barely move.

QuestionNames changed
in 20 minutes
Names changed
in 3 weeks
Lightweight privacy-friendly product analytics04
Hotjar alternatives for session replay20
Open-source session replay tools01
Record and replay mobile app sessions23
Organise and connect research notes with AI62
NotebookLM alternatives for research papers76
Read-later apps that help you finish45
Turn YouTube and podcasts into notes36
AI-native clouds bundling db + inference42
Inference and a database on one cloud05
Alternatives to Vercel and Supabase44
Functions + Postgres + an LLM endpoint124
Mean3.673.50

The control moved more on five questions, the three-week gap moved more on six, and one was a tie. A two-sided sign test gives p = 1.000. There is no effect of elapsed time in this data at all — not a weak one, not a suggestive one. The turnover we were about to report as market movement is entirely reproducible at zero elapsed time.

One question is excluded from the twelve: a definitional query whose extracted "names" were feature phrases rather than vendors. It churned enormously on both arms and would have flattered the result in the same direction, so leaving it out is the conservative choice.

What survives

Absence. All three companies were named in zero of 78 runs that morning, across two independent draws. One of them is zero of 45 across three separate readings on two dates a week apart. In none of our records has a company we screened as absent ever flickered into an answer on a re-draw. Whether you are in the list is a stable property; which other names are in the list, past the top two or three, is not.

The head-versus-tail split is worth stating precisely, because it is what makes the measurement usable at all. On the infrastructure question, CoreWeave was named in all three readings and RunPod in two of three; Together AI, Crusoe, Modal and Lambda were each named exactly once. Quoting the first two names of a list is defensible. Quoting the fifth is quoting a coin flip.

The uncomfortable part

Our own product is priced as monitoring, and monitoring means detecting change. This result says that at three runs per question we cannot detect change in who elseis named. We can detect a client crossing from absent to present, because absence is stable — but that is a slower, rarer, more binary product than a monthly report on category movement, and it would have been easy to keep selling the second one. Nobody had measured the noise floor of this instrument in the twenty-four days since we built it. It shipped with a docstring warning that model output varies between runs, and we then read every three-run output as a fact about the world.

The rule we have adopted, and the one we would suggest to anyone doing this work: a claim that something changed requires a same-session control. One draw against another draw at the same moment. If the control moves as much as your interval does, there is no change to report. It costs one extra batch and it is the difference between a measurement and a story.

What we sent instead

The three follow-ups went out with the retraction in them — the trend line we had planned, why we ran a control, and why we are not going to sell it. What replaced it is the number that held: absent in every reading, never once named. It is a smaller claim than the one we started with. It has the advantage of being true twice.

Method: gemini-flash-latest, three runs per question, 13 questions across three companies, readings on 2 – 15 August and two independent batches on 22 August 2026. One model answering from training knowledge, not live web search; we would expect retrieval-grounded answers to be more stable. Every run, including the raw model output, is recorded in our repository.

Frequently asked

How stable are the vendor names an LLM returns for the same question?

In our test, about half stable. We asked 13 buyer-intent questions of gemini-flash-latest at three runs each, then repeated the identical batch twenty minutes later. Across 12 comparable questions, 50 of 94 distinct vendor names — 53% — appeared in both readings. The other 47% were present in one three-run reading and absent from the next, with no change in the model, the prompt, or anything else. The head of each list was far more stable than the tail: the first two or three names usually returned, the fifth frequently did not.

Does the answer set actually change over weeks, or is that just sampling noise?

On this evidence it is noise. We compared readings taken 17 to 20 days apart against a control pair taken the same morning. The mean symmetric difference — the count of names present in one reading and not the other — was 3.50 names across three weeks and 3.67 names across twenty minutes. A sign test over the 12 questions put the control higher on five, the three-week gap higher on six, with one tie: p = 1.000. Elapsed time contributed nothing we could detect. Anyone reporting a month-over-month shift in an AI answer set without a same-session control is likely reporting the instrument.

So is an AI visibility check worthless?

No, but its useful claim is narrower than the one usually sold. Absence was completely reproducible. The three companies we tested were named in 0 of 78 runs that morning, across two independent draws, and one of them was 0 of 45 across three readings on two separate dates. Not one flickered into an answer even once. Whether a company is in the answer is a stable measurement; which other companies are in the answer, past the top few, is close to a coin flip. Build claims on the first, not the second.

How many runs do you need before a reading means anything?

More than three for anything except presence and absence. Three runs is enough to establish absence, because absence in our data never wavered across any number of draws. It is not enough to establish that a named list changed, because two three-run draws of the same question at the same moment differ by roughly three and a half names on average. To detect a real shift you would need the noise band to be narrower than the effect you are claiming, and we have not yet found the run count where that becomes true. Until you have measured your own noise floor, the honest reporting unit is presence or absence.

What does this mean for AI visibility monitoring products?

Monitoring implies detecting change, and detecting change requires resolving it above the noise. A monthly report that says your competitive set moved — three vendors entered, two left — is describing something our control reproduces at zero elapsed time. That does not make such products useless; it makes one specific output of them unreliable. The output that survives is the binary one: are you named for the questions your buyers ask, or not. We sell that measurement, and we are publishing the limit we found in it because a customer paying for measurement is entitled to know what it cannot resolve.

What are the limits of this result?

It covers one model, gemini-flash-latest, answering largely from training knowledge rather than live web search. It covers 13 questions across three small software companies in analytics, research tooling and cloud infrastructure. It uses three runs per reading, and a higher run count would narrow the noise band — that is the point of the finding, not a defect of it. We have not tested whether ChatGPT, Claude, Perplexity or web-grounded search behave the same way, and we would expect grounded search to be more stable because retrieval anchors the answer. Read this as a demonstration that the noise floor must be measured before a trend is claimed, not as a universal constant.

More from the Journal