I wrote a kill gate that couldn’t be killed

Context: this happened on our own build, before any data was collected. Nothing shipped. I am writing it up because the mistake is more instructive than the thing being built.

We are building a tool that detects a specific failure: a required CI check that never ran but reported success anyway. Before starting, I did what I would tell any client to do — I wrote down, in advance, the result that would make us abandon the project.

Writing the criterion first is the whole point. Decided afterward, it becomes a rationalization. Decided before, it is a commitment.

Here was mine, roughly: run the tool against twenty mature repositories; if at least 30% show at least one finding that survives manual verification, proceed. Otherwise stop.

That looks disciplined. It is not. It could not have failed.

01

Three ways it was rigged, none of them deliberate

I counted a category I had not defined. The criterion counted findings labeled ABSENT. There is no such classification in the system. I wrote a threshold against a label that did not exist, which means at scoring time someone would have had to decide what counted — after seeing the data.

“Survives manual verification” verified the wrong thing. It confirmed that a check really was skipped. But CI systems skip jobs constantly and correctly: a documentation change does not need the integration suite. Nearly every repository would have produced confirmed skips. The criterion measured a normal, healthy behavior and counted it as evidence for my hypothesis.

The repositories were not chosen in advance. I specified how to pick them — but one of my own criteria was “at least three distinct check names,” which requires inspecting a repository before selecting it. A person who wants the project to continue, inspecting candidates, will pick differently than one who doesn't. The selection bias sat inside the control.

Any one of these would have been enough. Together they guaranteed the answer.

02

How it was caught

Not by me. A colleague reviewing the plan wrote back, in effect: this gate probably cannot fail, and listed the three reasons above. He then waited for a ruling rather than proceeding.

Two details matter. He flagged it before any data existed — so the fix could not be shaped by what the data showed. And he raised it as a blocking question rather than adjusting quietly, which is why there is a record of it at all.

The revised criterion is materially harder to clear: named classifications only, a minimum presence rate, manual confirmation that no intentional filter explains the gap, and the repository list committed and hashed before the first scan.

The principle

Falsifiability applies to your decisions, not just your tests

I write about tests that would pass whether or not the feature works. The lesson generalizes further than I had applied it. A decision criterion that would be met regardless of the answer is the same defect, one level up. It looks like rigor. It produces a predetermined outcome with a paper trail.

And notice where it happened: in a project about checks that cannot fail, written by someone with that idea actively in mind, in a document whose whole purpose was to prevent self-deception.

Being careful did not catch it. A second reader who was expected to push back, and had nothing to lose by doing so, caught it in one pass.

So alongside the question I ask of every green check — what would I have seen if this were broken? — here is the one for every plan: what result would have made me stop? If you cannot name it concretely, in advance, and show that it was reachable, you have not set a criterion. You have written down your conclusion.

This is the gap we hunt.

Plumbline runs an independent verification audit on AI-built software — not just whether your checks pass, but whether they’d fail if something were broken, and whether they ran at all. Through a relay you control, zero access to your systems.

Book a 20-min fit call
← All field notesplumblinehq.ai