Thicket.
← Back to Journal

Case study

We Told a Startup to Build Comparison Pages. It Already Had Them.

Seven buying questions. Sixteen runs. Zero mentions — for a company whose comparison pages were already live, indexed, and better built than most. What that rules out is more useful than what we originally thought we had found.

We audited a documentation startup, told it to build three comparison pages, and two of the three already existed — server-rendered, indexed, around 1,500 words each. Re-running the measurement changed the finding rather than retiring it: the pages are there, and the model still does not name the company. On-page work is what makes you eligible to be named. It does not appear to be what causes it.

This is published with the company’s permission and names it, because a case study nobody can check is an anecdote. The company is Moxie Docs, which builds tooling to keep developer documentation in sync with the codebase it describes. We reached out cold, got the correction on the record, and agreed to publish what we measured either way.

What we got wrong

Our first email recommended three pages: a Mintlify alternatives page, a Moxie vs Swimm comparison, and a page on tools that keep docs in sync with code. The reply was short: “I believe I have pages for those already on our site.”

He was right. Checking the sitemap afterwards — which is what we should have done first — turned up 110 URLs including a dedicated comparison directory. The two pages we had recommended building were live and, on inspection, good:

  • /vs/mintlify — 1,554 words of server-rendered text, titled “Mintlify alternative: codebase-native docs & MCP context”, marked index, follow
  • /vs/swimm — 1,432 words, same construction
  • /learn/what-is-documentation-drift — 1,528 words, defining the term the company organises its positioning around

That is not a company neglecting its content. It is a company that has done the work the standard playbook prescribes, in the shape the playbook prescribes it.

What we measured

Seven buying questions, put to one model on 1 August 2026, each run more than once, with the runs counted rather than summarised. The question wording is the wording a buyer would use; the company name never appears in the prompt. Sixteen usable runs in total.

Question askedWho the model namedNamed?
Best alternatives to Mintlify for developer documentationGitBook, ReadMe, Archbee, Scalar0 / 3
Which tools keep documentation in sync with code as it changesSwimm, Mintlify, ReadMe, Danger, GitBook0 / 3
What is documentation drift, and what tools fix itSwimm, Mintlify, Backstage, Redocly, GitBook, Archbee0 / 2
Tools that automatically generate and sync developer docsSwimm, Mintlify, ReadMe, GitBook, Sphinx, Archbee0 / 2
Best MCP servers to connect to a coding agent0 / 2
Tools that generate an AGENTS.md or CLAUDE.md file0 / 2
Best alternatives to Mintlify (earlier phrasing)0 / 2

The third row is the one worth sitting with. Documentation drift is the company’s own organising term. It has a 1,528-word page defining it. Asked what documentation drift is and which tools fix it, the model defined the term correctly and then recommended six other companies.

What this rules out

Before the correction, our finding was the ordinary one: build comparison pages. That finding is now unavailable, and its unavailability is the useful result. The pages exist, they are indexed, they are server-rendered rather than hidden behind client-side JavaScript, and their titles match the queries almost word for word. Every on-page explanation for the absence is eliminated.

What remains is off-site. A model answering from training knowledge is reflecting what the web said about a company, and a company’s own page is the one source guaranteed to be flattering. Every vendor in this category publishes a page claiming to be the best alternative to its largest competitor. That claim carries almost no information precisely because everyone makes it. What appears to separate the named from the unnamed is whether anyone else said it.

We hold that as a hypothesis, not a mechanism. We cannot see inside the model, and anyone who tells you they can is selling something.

Why we are publishing this instead of a testimonial

During the exchange we were offered a reciprocal link — a partnership post with a dofollow link each way, for domain rating. We declined it. Reciprocal pairs are discounted for being reciprocal, domain rating is not what a language model reads, and accepting would have contradicted the advice in the same email. Instead we offered this: we publish the measurement, the link goes one way, and we ask for nothing back.

That is not generosity. It is the only version of the exchange that produces evidence rather than an arrangement. A traded link would tell us nothing in September.

The falsifier

On 27 September 2026 we will re-run these seven questions, the same wording, against the same model, and publish the delta. If nothing has moved, we will publish that. A hypothesis that can only be reported as a success is not a hypothesis.

How to run this on yourself

  1. Write the question your buyer asks, not your brand name. A model will discuss a brand named in the prompt while never surfacing it unprompted.
  2. Run each question at least three times and report the count, not a verdict.
  3. Record who was named. The competitor list tells you which material the model learned from, which is more actionable than your own absence.
  4. Treat a failed request as an error, never as an absence. We separately found a metric of our own that reported 0% for three months off 280 consecutive failed API calls, because a failure and a zero were stored identically.
  5. Then check the thing we forgot to check: whether the pages you are about to recommend already exist.

Limits of this measurement

One model, answering from training knowledge, without live web search, on one date, about one company. A different assistant may answer differently, and the same assistant with web grounding enabled often does. This is a single case, and a single case cannot establish that off-site mentions are what move AI visibility — it can only establish that, here, on-page work was not sufficient. That is a narrower claim than the one we would like to make, and it is the one the evidence supports.

Frequently asked questions

Do comparison pages help you get mentioned by AI assistants?

They appear to be necessary but not sufficient. In the case measured here, a company had a well-built, server-rendered, indexed page titled almost exactly the query a buyer would ask — and a model answering that query named four competitors and not them, across three separate runs. On-page work is what makes you eligible to be named; it does not appear to be what causes it. The distinction matters because most advice in this area stops at the page and implies the rest follows.

Why would an AI not mention a product that has a page targeting the exact question?

Because a model answering from training knowledge is reflecting what the web said about a product, not what the product said about itself. Self-published pages are a weak signal for this purpose: every vendor claims to be the best alternative to its competitor, so the claim carries little information. Third-party corroboration — someone else naming you, in a context the model ingested — is the signal that appears to separate the named from the unnamed. That is a hypothesis consistent with what we measured, not a proven mechanism.

How many times should you run a query before concluding a product is absent?

At least three, and report the count rather than a verdict. Model output varies between samples, so a single run can miss a company by chance and produce a confident false absence. In the measurements below we recorded runs attempted, runs that returned a usable answer, and runs that named the company, so that a failed request could never be silently recorded as an absence. Across seven questions this came to sixteen usable runs and zero mentions, which is a much stronger statement than any single query could support.

Does being absent from one model mean you are absent from all of them?

No, and it is important not to overstate it. Every measurement here comes from one model answering from training knowledge, on one date, without live web search. A different assistant, or the same assistant with web grounding switched on, can and does answer differently. What a result like this supports is a claim about one model on one date. Anyone reporting it as a general statement about artificial intelligence is overreaching, including us.

What actually moves AI visibility if pages are not enough?

The honest answer is that we are running the experiment rather than asserting the conclusion. Our working hypothesis is that off-site mentions — independent roundups, documentation-adjacent writing by other people, community threads, directory entries — carry the weight that self-published pages do not. We have committed to re-running the identical seven questions against the same model in late September 2026 and publishing the delta, including if the delta is nothing. A hypothesis that cannot be published as a failure is not worth publishing as a success.