Short, technical write-ups on the gap between a passing check and software that actually works — the failures that hide there, and how to find them before they ship.
Three builds with AI agents doing the coding — the operating model, the record, and the gates, with the failure that forced each one.
A required check named “Build & Test Passed or Skipped” is not a mistake — it is a rational workaround for a platform limitation, and it quietly permits the thing it appears to prevent.
I set a go/no-go criterion for a project built to detect checks that cannot fail. The criterion could not fail.
The engine computed the right answer and nothing ever called it. No test noticed, because no test looked at what the user actually receives.
If verification is only worth paying for when it finds a bug, then a clean result is a refund. That framing is wrong.
Seven true checks in the status report. Not one false statement. And the one gate that would have caught the bug never ran.
A clean deploy told us nothing about whether our fix worked — because the code we changed only runs when something is already broken.
Code that worked, but only on state nobody wrote down. It passes every test until the first clean rebuild.
A passing test is only evidence if a broken system would have failed it. Three real cases — and the one question to ask of every green check.
New writing on verifying AI-built software — the next piece when it’s ready. No spam.