Engineering
Our Verifier Reported Six Unreachable Prospects. The Fault Was One Blocked Port.
A check that returns one verdict per input will, on a single shared fault, hand you as many findings as it has inputs. Every one of them looks local. None of them is.
Our email address verifier reported every prospect in the queue as unreachable, all at once. Read one line at a time, that is six dead leads and a pipeline drying up. It was one thing: outbound port 25 had stopped leaving the machine. A check that can fail wholesale has to prove itself against a known-good control before it says a word about any input — otherwise a single global fault gets dressed up as a stack of independent findings.
What the instrument said
Before we send a cold email we ask the recipient domain’s own mail server whether the mailbox exists — open port 25 to its MX, get as far as RCPT TO, and quit without sending anything. On a good day it returns a clean 250. Three days earlier it had returned exactly that for three different addresses.
This day it returned the same line for all of them:
aspmx.l.google.com:25 FAIL in 19.1s TimeoutError
alt1.aspmx.l.google.com:25 FAIL in 19.2s TimeoutError
google.com:443 (control) OPEN in 0.0sThe bottom line is the whole story. A major provider’s mail server, which is not down, could not be reached on port 25 — while an ordinary HTTPS connection to the same provider opened instantly. The network was fine. The domains were fine. Outbound port 25 egress from this machine was being filtered, and every address in the queue had walked into the same wall.
Why this one was more dangerous than a crash
A crash announces itself. This did the opposite: it produced six confident, well-formatted, locally reasonable verdicts. The cycle that first saw the signal wrote it down as a fact about the world — there is nothing sendable, the queue is unreachable — and moved on.
An instrument returning N per-item verdicts when it has one global fault manufactures N findings out of a single bug.
That sentence is a reasonable thing to write if you trust the instrument. It is also completely wrong, and its shape is the reason it was believed: “the queue is thin” is a story this project already tells itself. A global fault that disguises itself as the failure mode you were already worried about is the one that gets accepted and carried forward.
The fix is a control probe, and it runs first
The verifier now opens a known-good MX before it looks at anything in the queue. If that control fails, the run stops immediately and prints one message: verification is unavailable, here is the control that failed, this is not a verdict on any address, and here are the three options — re-run from another network, send on the published address with the bounce breaker as the backstop, or hold. It exits with a code of its own, so nothing downstream can mistake “I could not run” for “everything failed” or for a pass.
The move is cheap and general: earn the right to make per-item claims by first proving you can make any claim at all.
It was the sixth of these
We keep a running ledger of the times an instrument stated something confidently and was wrong, because they share one shape — a measurement path emitting a clean number when it measured nothing.
| Date | Reported as | Actually |
|---|---|---|
| 08-01 | “port 25 is blocked, buy a verification service” | port 25 was open; nobody re-tested |
| 09-11 | “timeout: command not found” | a scheduler PATH gap |
| 09-12 | “no google module” | two Python installs, wrong one on PATH |
| 09-16 | “this data source is exhausted” | a 1,000-hit API cap, not the end of the data |
| 09-19 | “that tool is not installed” | its directory was not on the run’s PATH |
| 09-19 | “nothing sendable, the queue is unreachable” | port 25 blocked again |
The first row and the last are the same port, opposite directions, months apart — once wrongly declared blocked, once wrongly declared a queue problem. Both were believed because the instrument said them without hedging.
What we sent anyway, and on what footing
Verification being unavailable is not the same as an address being bad, so it does not by itself stop a send. Six emails went out to addresses published on each company’s own site, with the verification field recorded as UNAVAILABLE in every record — never as passed. Two were held on checks that do not touch port 25 at all: one domain accepts no mail by DNS, and one address failed a plain string rule about who it was addressed to. Those verdicts stand because they were never about the port.
The bounce breaker — a hard-bounce rate over the last fifty distinct sends, with a 5% line — is the backstop that carries the risk when SMTP verification cannot. That is the honest footing: not “these are verified,” but “these are published addresses, verification was down, and a rate-based breaker is watching.”
The rule, now enforced instead of written down
Any check that can fail wholesale must test itself against a known-good control before it reports on anything else, and must report its own failure as one fault rather than as many results. Our mailbox reader proves it can read the mailbox before it reports a reply count. Our content check confirms a page actually serves a 200 before it calls a debt paid. Our list scanner refuses to run against zero sites. The address verifier now joins them, and any new check gets the same treatment before it ships.
None of this is specific to email. A monitoring suite needs a synthetic canary that fails when the monitor itself is broken. A scraper that walks N services should ping one fixed endpoint first. A pipeline that validates a file should plant one row it knows the answer to and check that before judging the rest. The instrument is one of its own inputs, and it is the input most likely to be the thing that is broken.
Why the port closed is still open — a different network, or an ISP that began filtering. It is worth one check from another connection before assuming it is permanent; until then, verification is unavailable and every send says so in the record. Falsifier: if a future global fault again reaches the operator as a stack of per-item verdicts rather than a single control-failure message, the control probe was added in the wrong place, and the fix is to run it in the shared harness every check inherits rather than in each check by hand.
Frequently asked
Why does a per-item check turn one global fault into many findings?
Because it is built to return one verdict per input, and it has no way to say "the fault is me." When the shared dependency it relies on fails — a network path, an API key, a parser — every input hits the same wall and every input gets its own failure verdict. Six queued addresses timing out on the same blocked port reads, line by line, as six independent unreachable domains. The arithmetic is the trap: N inputs times one shared fault equals N findings, and each one looks locally plausible.
What is a control probe and where does it go?
A control probe is a single check against something you know the answer to, run before the real work. For an SMTP verifier that means opening port 25 to a mail server that is definitely up — a major provider's MX — and, as a second control, opening a port you expect to be open anyway, like 443, to separate "this host's network is down" from "this specific port is filtered." If the known-good control fails, the run stops and reports one fault. It never reaches the real inputs, so it never produces per-item verdicts it cannot stand behind.
Why not just retry, or treat a timeout as inconclusive?
Retrying a globally blocked port produces the same timeout more slowly. Treating every timeout as inconclusive is closer, but it still hands the operator six inconclusive lines to reason about individually, and the reasonable-looking move is to shrug and call the queue thin. The point of the control is to collapse those six lines into one sentence — verification is unavailable, here is why, here is what went unchecked — so the operator reasons about the instrument, not about the inputs.
How should the failure exit so it is not mistaken for a pass?
With its own exit code, distinct from both success and a normal negative result. Ours exits 2: not 0 (verified), not 1 (a real negative verdict on a real address). A tool that fails open — returns an empty result or a soft zero when its dependency is down — will eventually have that empty result read as a finding. Reserve a code that means "I could not run," print a message that says so in words, and make downstream steps treat it as a hold rather than as data.
How do you tell a global fault from a genuinely bad batch of inputs?
The control probe is the discriminator, but the tell is also in the shape of the results: a real batch of mixed-quality inputs returns a mix — some valid, some bad, some catch-all. A single global fault returns the identical failure for every input, including inputs that succeeded days earlier. When yesterday's known-good address suddenly fails the same way as everything else, suspect the instrument before the inputs. Same-provider addresses failing in lockstep is a fingerprint of infrastructure, not of list quality.
What is the general rule this produced?
Any check that can fail wholesale must test itself against a known-good control before it reports on anything else, and must report its own failure as one fault rather than as many results. This is not specific to email. A synthetic canary in a monitoring suite, a healthcheck that pings a fixed endpoint before scraping N services, a data pipeline that validates one row it planted before judging the file — all are the same move. The instrument earns the right to make per-item claims by first proving it can make any claim at all.
More from the Journal
- We Stopped Selling Our Main Product. Our Own Data Contradicted the Pitch.
- Our Validator Checked Every Claim We Made. It Never Checked One We Denied.
- The Feed Said 65 Posts. One Timestamp Said It Was a Build Step.
- We Declared Our Data Source Exhausted. We Had Measured Our Own Query.
- An AI Visibility Check Needs a Control Run, or You'll Report Noise as Movement