Every green light was real. The red one, we never turned on.
I sell verification. I write about the ways a green check lies. And a few sessions ago, my own build reported a feature as proven, wrote that into the permanent record, and it wasn’t true — and here’s the part that should bother you as much as it bothers me:
Not one statement in that report was false.
Seven true greens
At the close of the session, the build reported this. Every line of it was accurate.
- 601 tests passing, zero failures
- Type checker: 0 errors
- Linter: clean
- Commit authorship verified
- Pushed clean; local matches origin
- Working tree clean
- Live proofs run against the real database — six scripts, real PASS lines
That is not a lazy report. That last line especially: those proofs weren’t mocks. They opened real connections to a real database under the correct restricted role and did real work. The kind of end-to-end evidence I spend my professional life telling people to demand.
On the strength of it, the feature was recorded as proven — in the state document, in the decision log, in the handoff. The next session would inherit that as ground truth.
It was wrong. And you cannot find the error by auditing that list, because there is no error in that list.
The check that was missing
The build has a preflight gate. Its own header says, in the file: run before every dispatch, merge, and report. It runs a handful of checks nothing else runs — including one that compares what the state document claims is the current commit against what the repository actually has, and fails if they’ve drifted apart.
That check would have caught this. It was correct. It was blocking. It would have failed loudly.
It never ran. Not before the work, not before the commits, not before the report. Across the entire session, the gate was invoked zero times — and nobody noticed, because the report was full of other, true green.
Seven honest checks stood in for the one that mattered. The report reads like comprehensive verification. It is strictly weaker than the single thing that wasn’t there. And the two failures hiding that day lived only inside the gate that wasn’t run — so no amount of scrutiny of the reported items could have surfaced them. There was nothing to catch.
What was actually broken (and why it’s the boring part)
Two things, and I’ll be quick, because the bug is not the story.
A database bug that’s a coin flip. The system talks to Postgres through a connection pooler in transaction mode, where many client connections get multiplexed onto fewer real backend connections. The driver names its prepared statements from a counter that restarts at zero on each client connection — but prepared statements live on the backend. So when two clients each count up to the same name and the pooler happens to route them to the same backend, they collide and the query dies. Whether it collides depends on how the pooler feels that second.
This is an ordinary, documented footgun with an ordinary fix, and we’d already written the fix. It was committed. It was wired into exactly one place that opens database connections — and three other places never called it. One of those three was the proof scripts.
And a gate that had quietly disarmed itself. The state-drift check finds the claimed commit by pattern-matching a line in the state document. A prior commit — and I want you to sit with which one — changed that line’s formatting just enough that the pattern stopped matching. It was the close-out commit: the ritual whose entire job is to certify the state document. The ceremony that certifies the record broke the check that validates the record, and never looked back at itself.
Two things I’d never seen named
The code that touches reality was the least-governed code in the repository. Those six proof scripts — the only code in the entire project that opens a socket to the real database, the sole source of truth about whether any of this works at the boundary — were not type-checked, not linted, not covered by any quality gate. Everything else was. The code we trusted most to tell us the truth was the code we checked least.
Go look at your own repo. The scripts/ folder. The smoke/ directory. The one-off that hits staging. I’d bet it’s outside the CI that governs everything else, and I’d bet it’s the only thing you own that talks to production reality.
And the proof itself was nondeterministic — in the passing direction. Because the collision is a dice roll, the proof scripts could connect to the real database, do real work, and legitimately pass — then fail on the identical bytes the next day. Which is exactly what happened. One session ran them and got real PASS lines. The next ran the same code against the same database and got the collision.
A check that flakes red is annoying but honest — it nags you until you look. A check that flakes green is a liar, and worse, it gets written down. That lucky pass went into the permanent record as proof. One observation of a nondeterministic check carries almost no information, and we treated it as though it carried all of it.
How it was caught, including the part I’d rather not tell you
It was caught by the least clever thing imaginable: running it again, on a different day, against the real system, and comparing the result to what the record claimed. The proofs failed. The record said they’d passed.
Then — and this is what separates a bug fix from a finding — not stopping there. Asking how did this get reported as green? Opening the preflight script and reading it rather than reasoning about it. Finding the broken pattern-match. Then running git log -S to find the exact commit that broke it, which is how we learned it was the close-out ritual itself. Then grepping the session transcript — the verbatim record of what actually executed — to answer whether the gate had been run and ignored, or never run at all.
Now the part I’d rather not tell you. My first post-mortem confidently asserted that the broken check had silently passed — failed open. It was a plausible mechanism. It was also wrong: the check failed loudly, printed [FAIL], and would have blocked. I hadn’t read the script. I had inferred a story.
A post-mortem written specifically to warn against reasoning instead of measuring — by someone actively trying not to do that — did exactly that thing. That’s not carelessness. That’s the pull of the failure mode, and it is strong enough to catch you while you are writing about it.
The invocation gap
I’ve written before about three ways a green check lies: it hit a stand-in instead of the real thing; it only passed on state a clean rebuild can’t reproduce; or it would have passed even if the feature were broken. All three are properties of a check that ran. All three answer the question: this went green — why did it lie?
This is a fourth thing, and it doesn’t fit. The check was real. Correct. Falsifiable. Blocking. It simply never executed — and its absence was perfectly concealed by the presence of checks that did.
Call it the invocation gap. Its defining property, and the reason it’s the nastiest of the four: you cannot catch it by examining the checks. Every check is fine. You have to examine whether they ran — and almost nothing in your tooling tells you that. A skipped job and a passing job are visually indistinguishable in nearly every dashboard we use. Filtered out by a path rule, excluded by a matrix condition, quietly disabled six months ago — and you read “no red” as “green.”
So alongside the question I ask of every green check — what would I have seen if this were broken? — there’s now a second one:
And if you’re building with agents, this is the one to internalize. Everything in that seven-line report was true, specific, articulate, and verifiable. That’s exactly what makes it dangerous. An agent tells you what it ran. It does not tell you what it didn’t run. An inventory of completed checks is not a verdict, and a fluent, honest, complete-sounding report is the best hiding place an omission has ever had.
The fix is unglamorous, like all the real ones: make the gate refuse to be skipped rather than asking to be remembered. Make the absence of a check as loud as its failure. Assert the checks ran, not just that they passed.
A plumb line only tells you the truth when you actually hang it. Left in the drawer, it never says a word — and a wall nobody measured looks exactly like a wall that’s true.
Plumbline runs a one-week Drift Audit on AI-built software — not just whether your checks pass, but whether they’d fail if something were broken, and whether they ran at all. Through a relay you control, zero access to your systems.
Book a 20-min fit callGet the next one.
New writing on verifying AI-built software — the next piece when it’s ready. No spam.