Zero type errors. 325 passing tests. Every document blank.

Context: a clinical build of ours, pre-production, synthetic data only. No real users were affected. Published because the failure shape is common and almost nothing catches it.

The build reported healthy for several sessions running. Type checker: zero errors. Engine suite: 325 tests, all passing. Application tests: green. By every signal available, the work was done.

Then someone opened a generated document — the artifact a real user would receive at the end of the workflow — and the section that should have contained their personalized plan was empty. Not wrong. Empty.

01

The defect

The engine had a function that built the plan. It worked. It was tested. It was correct.

The application never called it. The field it was supposed to populate was hardcoded to an empty list, behind a // TODO comment written weeks earlier and forgotten. Every document produced downstream rendered that empty list faithfully.

Nothing was broken. The engine was right; the tests were right; the renderer was right. The wire between two correct things was never connected, and no part of the system was responsible for noticing.

02

Why 325 passing tests were blind to it

The engine tests called the engine directly and asserted it returned the right plan. It did. Those tests would pass forever regardless of whether anything in the product ever invoked that function.

The application tests exercised the surrounding logic with fixtures — and the fixtures supplied a plan, because that is what fixtures do. They encode the author's belief about the shape of the data. Nobody wrote a fixture representing “the field is empty because the app forgot to fill it,” because nobody imagined it.

So both halves were verified, and the seam between them was verified by neither. This is the proxy-coverage gap in a form that is particularly hard to see: the tests hit real code, not mocks. They just hit it from a direction no user ever takes.

03

The class of bug mutation testing cannot catch

Deliberately breaking a line and checking whether tests notice is a powerful technique. It is also blind here. There was no wrong line to break. The defect was an absence — a call that was never written. Mutating the engine's logic would have produced failing engine tests and proved nothing about whether the application used it.

Absence is invisible to any technique that works by perturbing what exists.

The principle

Test the artifact the user receives

The only thing that caught this was a human looking at the output. That is not a scalable control, but it points at one that is: render the real artifact in a test and assert something about its contents.

Not “the function returns a plan.” Not “the renderer handles a plan correctly.” Instead: run the actual pipeline end to end and assert that the produced document contains at least one item, that it reflects the inputs given, that the fields a user depends on are populated.

Every test asked whether a component was correct. None asked whether the thing we hand the user is worth handing them.

That test is unglamorous and slightly slow and it would have caught this on day one. It would also catch the next version of this bug, which will not be the same missing call — it will be a different seam nobody thought to name.

If your system produces something a person eventually reads — a document, a report, an invoice, a summary — ask what would happen if it came out blank. In most systems: nothing. Nothing at all.

This is the gap we hunt.

Plumbline runs an independent verification audit on AI-built software — not just whether your checks pass, but whether they’d fail if something were broken, and whether they ran at all. Through a relay you control, zero access to your systems.

Book a 20-min fit call
← All field notesplumblinehq.ai