Thicket.
← Back to Journal

Engineering

Our Script Said Three Data Sources Had Zero Rows. They Had 16,589.

One function returned the same empty list for “there is nothing here” and for “I could not look.” Every number built on top of it had the same problem.

A fetch function that returns an empty list when a request fails makes the failure look exactly like a query with no results. Ours did that on three data sources in one run, and the summary said “0 rows” for each. The real counts were 6,230, 5,966 and 4,393. The fix is to return a status with the data and to print UNKNOWN, never zero, for anything that was not read.

What happened

We were building a list of companies from certificate transparency logs, the public record of every TLS certificate issued. The script asks crt.sh for every certificate under a top-level domain and keeps the distinct apex domains. On its first run against three TLDs it reported:

.health 0 rows
.care 0 rows
.dental 0 rows

That looks like a finding about those TLDs: few companies there, or none using certificates the log sees. It was close to being written down that way. What had actually happened is that crt.sh was rate-limiting us, and all three requests failed.

The function

This was the whole fetch function:

def fetch(tld, timeout=60):
    url = f"https://crt.sh/?q=%25.{tld}&output=json"
    try:
        raw = urlopen(Request(url, headers=UA), timeout=timeout).read()
        return json.loads(raw)
    except Exception as exc:
        print(f"{tld} FAILED: {exc}", file=sys.stderr)
        return []

The failure was logged. It went to stderr, and then the function returned the same value a real empty result would have returned. The caller counted the list, got zero, and printed zero on the line that gets read.

An empty list from a failed request and an empty list from a query with no results are the same object. They mean opposite things.

Nothing downstream could tell them apart, because the information had been thrown away inside the function. A log line one line above a result that contradicts it only helps someone who already suspects the result.

The fix

def fetch(tld, timeout=60, retries=2):
    """Returns (rows, status). NEVER conflate them."""
    url = f"https://crt.sh/?q=%25.{tld}&output=json"
    last = ""
    for attempt in range(retries + 1):
        try:
            raw = urlopen(Request(url, headers=UA), timeout=timeout).read()
            if not raw.strip():
                last = "empty body (rate limit?)"
                time.sleep(4 * (attempt + 1)); continue
            return json.loads(raw), "ok"
        except Exception as exc:
            last = f"{type(exc).__name__}: {exc}"
            time.sleep(4 * (attempt + 1))
    return [], f"FAILED ({last})"

Four changes, and the first matters most:

  1. The status travels with the data. The caller gets (rows, status) and has to look at the status before it can count anything.
  2. An empty body is suspect. A whole TLD with no certificates is unlikely. A 200 with nothing in it is more likely a rate limiter than an answer.
  3. Retries with backoff, so a rate limit gets a chance to clear before we give up.
  4. The report says UNKNOWN. A source that could not be read prints FAILED (reason) — UNKNOWN, not zero on its own line, and the run ends with a list of every unread source and a note that the pool is incomplete by that much.

What the re-run found

RunRows reportedWhat it meant
First run0 / 0 / 0three failed requests
Re-run, fixed6,230 / 5,966 / 4,39316,589 rows read

Reduced to distinct apex domains, those rows held 1,379 domains that the first run had silently dropped. The list would still have looked normal, just smaller, and nothing in it would have shown the gap.

This keeps happening, in different places

This is not the first time a failure has shown up in this project as an empty result or a zero. Each time the tool involved was different:

  • An API’s 1,000-hit cap read as “this data source is exhausted” (write-up).
  • A blocked network port read as six unreachable prospects (write-up).
  • An extractor that found no names read as “no vendor was mentioned” (write-up).
  • A mailbox reader pointed at the wrong inbox reported zero replies for four weeks.

The common cause is code that fails open: when it cannot measure something, it returns the value that means “nothing”. Zero, an empty list and None are all legitimate answers, which is why they get believed.

The rule

Never return an empty collection or a zero to mean “I failed.” Return the data with a status, or raise. At the reporting layer, a source you could not read is UNKNOWN. Print that word, not a number.

A quick way to find the pattern in your own code is to search for return [], return {}, return 0 and return None inside except blocks. Each one is a place where a failure can be reported as a result.

Scope: one script, one data source, one run. The three row counts and the 1,379 figure come from our own logs of 2026-09-28. We have not checked whether crt.sh always returns an empty body when it rate-limits or sometimes returns an error, so the fix handles both. Falsifier for the rule: if a future status file reports a zero that turns out to be an unread source, the rule was applied in one script instead of in the shared fetch layer that every script uses, and that is where it needs to move.

Frequently asked

Why is `except: return []` a bug and not just a shortcut?

Because the caller can no longer tell two outcomes apart. An empty list means "I asked and there was nothing." After `except: return []`, it can also mean "I never got an answer." Code downstream counts the rows, gets zero, and reports zero, and nothing about the zero shows that it was never measured. The shortcut saves one branch in the fetch function and moves the ambiguity to every place that reads the result.

The failure was logged. Isn't that enough?

No. Ours logged it. The old function printed a FAILED line to stderr and then returned an empty list, and the summary line printed "0 rows" for that source. People read the summary. A log line that disagrees with the result it sits next to only helps someone who already suspects a problem. The status has to travel with the data, onto the line that gets read.

What should the function return instead?

The data and a status together: `(rows, status)`, a result type, or an exception that the caller has to handle. The form matters less than the rule that no code path returns an empty collection to mean failure. Ours now returns `(rows, "ok")` on success and `([], "FAILED (reason)")` otherwise, and the caller checks the status before it counts anything.

What if the server returns 200 with an empty body?

Treat it as suspect, not as an answer. Rate limiters and overloaded APIs sometimes return a successful status with nothing in it. For a source where an empty result is unlikely, such as a certificate log for a whole top-level domain, an empty body counts as a probable rate limit. Back off, retry, and if it is still empty, report it as unread.

How should a partial read be reported?

Name what is missing and say that the total is incomplete by that much. Our script now prints "UNKNOWN, not zero" on each source it could not read, and ends with a line listing every unread source and saying the pool is incomplete. A total with a hole in it is still usable if the hole is labelled. Without the label it is a wrong number.

Is retrying the fix?

It is part of it. A re-run with backoff read all three sources in our case. But retries run out, and when they do the function is back to choosing between an empty list and an honest failure. Retry to recover as much as you can, and report whatever you could not recover as a failure.

More from the Journal