Thicket.
← Back to Journal

Decision

We Stopped Selling Our Main Product. Our Own Data Contradicted the Pitch.

A prospect worked out that our product does not do what we said, in one email, by quoting a caveat we had published ourselves. He was right.

We sold a $1,500-a-month content operation as the way to become visible when an AI assistant is asked your category’s buying question. Three separate measurements now say on-domain publishing does not move that number. We have stopped selling it, and cut the price of what replaced it to $199.

Three measurements of whether on-domain publishing makes an AI assistant name you: a customer with four comparison pages at 0 of 24 runs, our own 419 articles at 0 of 12, and a study finding third-party attention predicts it at p=0.0008 while self-published content shows no signal.
Any one of these is arguable. Together they are the same answer from three directions — and we were selling the opposite.

The email that ended it

We had audited a developer-tools company, found it absent from every run, and sent three recommended artifacts. The first was: write a comparison page against each of your two main competitors.

Their CEO replied:

Those pages have been live since July. So artifact 1 already exists, and the model still returns zero — which matches your own caveat that on-domain pages are not the lever.

We checked. Four comparison pages, several hundred words each, robots: index, follow, all four in the sitemap, live for two months. Absent in 0 of 24 runs.

Two things were wrong at once. We had recommended work he finished in July, which means we had read his homepage and his blog and not his site. And the recommendation itself was contradicted by a caveat we had written and he had read.

Three measurements, none of them dependent on the others

sourcewhat was measuredresult
A company we audited4 comparison pages, live 2 months, indexable, in their sitemapnamed in 0 of 24 runs
Ourselves419 published articlesnamed in 0 of 12 runs, flat for a month
Our own studywhich variable predicts being namedthird-party attention, p = 0.0008. Self-published content: no signal

Any one of those is arguable. His site is one site; ours is a young domain; a correlation at p = 0.0008 is not a mechanism. Together they are the same answer arriving from three directions, and we were selling the opposite.

What the replies actually asked for

We had six positive replies to cold outreach. When we finally read all six rather than the two we remembered, the pattern was not subtle:

  • Zero asked us to produce content.
  • Two asked for the measurement, over time. One wrote: “a measurement worth tracking over time, not once.”
  • One declined explicitly on price — “beyond our current monthly commitment.”
  • One asked us for a backlink swap, which is its own kind of answer.

We had been selling a product into a population that asked for it zero times out of six, while asking for something else twice.

The part where we are careful about our own evidence

Our pre-registered falsifier had two clauses: no signed customers, and fewer than five qualified leads. Both were true, so it fired. But only one of them was informative, and we would rather say that than take credit for a clean result.

Zero sales from six replies proves nothing. At a 10% close rate you would see zero about half the time; at 5%, three quarters of the time. That clause was never powered by six conversations. What the six do establish is qualitative, and qualitative evidence does not need a large sample to be worth acting on: nobody asked for the thing we were selling.

What we sell now

The measurement. A fixed set of the customer’s own buyer questions, run on a schedule, with every run drawn twice on the same day.

The control draw is the entire point. We measured the instrument against itself: only 53% of vendor names survive a re-draw twenty minutes later, and two same-morning draws differ as much as draws three weeks apart. So a before-and-after reading without a control cannot tell you whether anything changed. Absence is the stable half — reproducible across every run we have done.

And we say the uncomfortable part in the product description: we can tell you the number, and we cannot yet move it on demand.

The date we will admit we were wrong

One customer shipped category content specifically against our finding and pre-registered the test himself — if the re-runs show movement, even one mention, publishing works; if it stays at zero after consistent publishing, that is a different problem. That reads on 17 September 2026.

If he has moved off zero, on-domain content does work, the original product was right, and this was an overcorrection from a sample of six. We will say so here and reverse it.

The transferable bit

If your product claims a mechanism you can measure, measure it — on yourself, and on the prospects you pitch — before you sell it. We were in the unusual position of selling a remedy whose effect we owned the instrument to check, and we shipped the pitch first and the check second.

The people most likely to notice are your best prospects. Ours did it in one email, using our own published caveat.

The company described is not named: it replied to cold outreach, declined to buy, and did not sign up to be a case study. Figures are from our own records — visibility runs against gemini-flash-latest, three runs per question, and the name-stability figure from two draws of the same thirteen questions twenty minutes apart. Our own visibility reading is published at we audit our own AI visibility, zero included.

Frequently asked

What exactly did you stop selling?

A $1,500-a-month autonomous content operation: daily articles, the full AEO/GEO layer, a newsletter, analytics. It was sold as the way to become visible when an AI assistant is asked your category's buying question. We still run that machine on our own properties. We no longer sell it as a remedy for AI invisibility, because we cannot show it produces that outcome.

What convinced you it does not work?

Three measurements that do not depend on each other. First, a company we audited had four comparison pages live for two months — indexable, in their sitemap, several hundred words each — and was named in zero of twenty-four model runs. Second, our own site has published 419 articles and is named in zero of twelve runs, flat for a month. Third, in our own study of which companies get named, the variable that tracked it was third-party attention at p = 0.0008; self-published material did not track it at all. Any one of those is arguable. Together they are the same answer from three directions.

Isn't six replies too small a sample to change a business on?

For the sales conclusion, yes, and we say so. Zero sales from six replies is unremarkable — at a 10% close rate you would expect zero about half the time. That clause of our own falsifier was never powered. What the six replies did establish is qualitative and does not need a large n: not one of them asked for a content operation, two asked for the measurement over time, and one declined explicitly on the monthly price. We are not claiming statistical proof that nobody wants content ops. We are reporting that our data contradicts the mechanism we were selling, which is a different and much cheaper thing to establish.

What are you selling instead?

The measurement, at $199 a month — down from $1,500. A fixed set of the customer's own buyer questions, run on a schedule, with every run drawn twice on the same day. The control draw is the point: we found that only 53% of vendor names survive a re-draw twenty minutes later, so a before-and-after reading without a control cannot tell you whether anything changed. Absence, by contrast, is reproducible. We sell the number and its error bars, and we say plainly that we cannot yet move it on demand.

Why cut the price rather than keep it and change the pitch?

Because the price was set for a service that produced content daily, and a measurement is not that. The only pricing evidence we had was a prospect declining with the words 'beyond our current monthly commitment' — so we priced from what it costs us to run, which is a handful of model calls and the discipline to run them identically every time, rather than from what comparable dashboards charge. Charging less for something we can evidence seemed better than charging more for something we cannot.

How will you know if this was the wrong call?

It has a date. One customer shipped category content specifically against our finding and pre-registered the test himself: if the re-runs ever show movement, even one mention, publishing works; if it stays at zero after consistent publishing, that is a different problem. That reads on 17 September 2026. If he has moved off zero, on-domain content does work, our original product was right, and this decision was an overcorrection from a sample of six — and we will say so and reverse it.

What is the general lesson?

If your product claims a mechanism you are able to measure, measure it — on yourself, and on the prospects you pitch — before you sell it. We were in the unusual position of selling a remedy whose effect we had the instrument to check, and we shipped the pitch first and the check second. The people most likely to notice are your best prospects: the one who caught us did it in a single email, by quoting a caveat we had published ourselves.

More from the Journal