We fixed the bug. Then we made it fail on purpose.
Third in a series on the gap between a passing check and a working system. The first two were about tests. This one is about a fix — and the moment a clean, green deploy told me absolutely nothing about whether the fix worked.
The green deploy that proved nothing
A service in one of our builds talks to an external system using a short-lived credential. It has a recovery path: when the credential goes stale in the middle of an operation, the code is supposed to quietly mint a fresh one and retry. That path had failed once already — a stale credential threw a hard error and killed a job partway through. So we fixed it, and shipped the fix.
The deploy went clean. Health check returned 200. The service was up and serving. Every signal green.
And not one of those signals touched the thing we’d just changed. The recovery path only runs when a credential is actually stale — and a fresh deploy has a fresh credential. So the code we fixed never executed. A green deploy here means “the service starts,” not “the fix works.” Those are different claims, and we’d only checked the first one.
Here’s the trap, stated plainly: the fix could have been completely wrong — minting nothing, retrying nothing — and every check would have come back exactly this green. When the credential isn’t stale, a working recovery path and a broken one produce identical output. The success state and the broken state are the same green until something is actually broken.
So we made it fail
The only way to know whether the fix worked was to manufacture the exact condition it was built for. So we backdated the credential to months-stale and drove a real write through the service.
Before the fix, that’s the precise failure that had killed the job earlier — reproduced on demand, the same hard error. After the fix, the same forced-stale write came back clean: the service minted a fresh credential, retried, and self-healed. That contrast — the identical input failing before and recovering after — is the proof. The green deploy never was.
A fix for a failure you never reproduced is a guess
This is the falsifiability gap again, pointed at a fix instead of a test. A fix for a failure path is only proven once you’ve made the failure happen and watched the fix handle it. Until then, “it deployed green” tells you the building didn’t burn down on its own — not that the sprinkler you just installed works, because you never lit a fire.
And the trap is structural, not careless. A failure path is, by definition, the path that doesn’t run in the normal case. So the normal case — the deploy, the smoke test, the happy-path request — sails through green whether your fix is right or wrong. You cannot tell them apart from the outside until you force the failure.
The discipline that catches it is blunt: for any fix to a failure path, don’t ship on a green deploy. Reproduce the original failure first — force the stale state, kill the connection, send the malformed input, whatever the fix was for. Confirm it still breaks the way it used to. Then confirm the fix makes it stop. The forced failure is the negative case that gives the green its meaning.
A plumb line only tells you the wall is true because gravity is pulling the other way. A fix only tells you it works once you’ve let the failure pull against it. Make it fail first — or the green is just a guess that hasn’t been tested yet.
Plumbline runs a one-week Drift Audit that finds where your AI-built software is green but unproven — including the fixes that deployed clean but never actually ran. Through a relay you control, zero access to your systems.
Book a 20-min fit callGet the next one.
New writing on verifying AI-built software — the next piece when it’s ready. No spam.