How to build software with AI agents without lying to yourself

A project I worked on had a verification gate. It ran on every push, it had a flag that ran the full proof corpus, and it exited zero.

It had never once run the proof corpus. Not in the entire history of the project.

A branch-ordering defect meant the function returned before it reached the work it existed to do. Every green result that flag had ever produced was a green result for doing nothing at all. Nobody noticed, because nothing ever failed.

That is not a bug story. The examples in this piece come from three builds I have run with AI agents doing the coding: a regulated healthcare platform over about a hundred and twenty sessions, a verification tool, and the system being built now to publish this website. Across all three, one category kept recurring — a check that reports success without doing its work. Not bad code. Bad evidence.

I have not counted it against every other defect class, so I will not call it the most common. I will say it is the one that kept costing us, and the one that was hardest to see.

What follows is the operating model we ended up with. Every rule in it exists because something went wrong first, and I have tried to keep the incident attached to the rule, because a rule without its incident gets optimized away by the next person who finds it inconvenient.

01

The problem nobody warns you about

A chat-based AI architect has no memory between sessions. Each new conversation starts from zero, and the long conversations that do the real work eventually run out of room. Without a deliberate handoff, three things happen, and they compound.

Knowledge evaporates. Why a decision was made. Which approaches were tried and abandoned. Which instruments are unreliable. What the owner approved and in what words. None of it survives a new chat window.

Snapshots rot. A handoff that says "tests: 3,894 passing" is wrong the moment anyone commits. The next architect reads it, believes it, and plans on fiction.

Errors repeat. The same misdiagnosis costs the same round trips every session. We spent hours, more than once, debugging what looked like a syntax error and was actually a paste too long for a terminal to receive intact.

The fix is one sentence, and everything else in this piece is a consequence of it:

The repository is the memory. The chat is disposable. A handoff is a fast index into the truth, never a substitute for it.
02

The operating model

Three roles, and the boundaries between them are the point.

The owner executes, makes business and legal calls, and approves anything protected. The architect — the chat AI for that session — reads state, plans, writes instructions, reviews work, merges, and writes the record. Lanes are coding agents in isolated working copies. They build exactly what an authorization permits, and they never merge, never push, and never touch a file outside a list they were given.

The architect never executes directly. Each turn it emits one block of commands; a human runs it and pastes the raw output back. This is slower than letting the AI drive, and it buys three things: a human sees every side effect before it happens, every action leaves a log, and — the one that matters most — the architect is forced to reason from measured output instead of from what it assumed.

Every block carries the same contract: refuse to run unless the hostname is the expected machine; timestamp at top and bottom; an end-marker outside the guard so a truncated paste is visible; everything written to a log file, because terminal scrollback is not a record; and side effects gated on expected state — refuse unless the branch is at the commit we think it is. That last one exists because a block pasted twice by accident will merge on the first run and look like unexplained drift on the second.

Blocks stay under about fifty-five lines. Longer pastes drop characters over a terminal, and dropped characters look exactly like a syntax error in code you just wrote. We lost real time to that before we measured it.

03

Two phases, and the first one writes no code

Every unit of work runs in two phases, and the split is the single highest-yield practice in the whole system.

Phase one: trace and halt. The agent edits nothing. It answers a numbered list of questions about the codebase, and every answer carries the command that proved it. Then it stops and commits only its report.

Its real job is not to gather information. Its job is to correct the architect before any code is written. In one session, phase-one reports corrected the architect's plan five separate times.

Every set of instructions ends with the same line: if this directive is wrong, say so rather than complying. An agent that complies with a wrong instruction has done exactly what you asked and wasted the session.

Phase two: build under an authorization. The architect writes a document that names the agent's own findings back to it, states the rulings, lists the exact files it may touch, and requires something specific: a failing test, captured by deliberately reverting the behavior, plus a control proving the test is not vacuous.

That requirement is the heart of it. A test that has never been observed to fail is not a test. It is a green square whose meaning nobody has established.

04

The record that actually survives

These files live in the repository, not in any conversation. Every session reads them at the start and writes them at the end.

A state file holding the current truth: a table of hard facts — the tip commit, the test count, the migration version, what each deployed service is running — plus one narrative section per session. A decision log, numbered, where every ruling carries its reasoning and every owner approval is recorded verbatim, including a one-word reply. An approval is only as good as the description it was given on, and six months later nobody remembers the description. A lessons file where every lesson is stated as a rule, with the incident that taught it. A protected-files list naming what cannot change without explicit approval, and why. A pin file holding the exact number of tests that must pass.

The habit that turned out to matter most is the smallest one. When the architect is wrong, it names the error out loud the moment it is caught — a short tag, in the conversation, before moving on. At the end of the session, that list of named errors gets collected into a single lesson describing the pattern they share.

That last step is what makes it valuable. Thirteen scattered incidents across these builds are thirteen anecdotes. The same thirteen, read together, were all one thing: checking a stand-in for a fact instead of the fact. A summary endpoint instead of whether the request succeeded. A file's permission bits instead of whether the other account could actually read it. A tool's stored credential instead of the one the program was using. Naming it did not stop it. The AI architects on these builds made the same error at least six more times after the rule was written down, including the one that wrote it. What naming changed is that it became catchable: a pattern with a name is something another reader can look for. Several of the recent ones were caught within a day by a second reader, one by a check that refused to run, and one only when it failed on the machine it was run on.

05

Gates, and the rules they taught us

One script is the authority, and every other check defers to it. The rules below are the ones we got wrong first.

A gate must never return before the work it exists to do. That is the story this piece opened with. Check the branch order. Then run the gate's own work as a measurement, once, and look at what it produced — not at its exit code.

The gate is never bypassed. No skip flags, no override variables. The close-out procedure asserts the override variables are unset before it will start, because a bypass used once becomes a bypass used every time there is a deadline.

Every check needs a control. A zero result is meaningless until you have shown the check can return non-zero. We flagged a repository as having no build step because the pattern could not match the way the file was written; the check was working perfectly and measuring nothing.

Validate instruments in the environment they will run in. A check written in one shell and run in another is a different check. A pattern tested against a made-up example is not tested.

Roll back on the real proof, not the unit tests. A redesigned component passed its unit tests, then failed twice against real data it had not anticipated. The merge was rolled back automatically. Unit tests said yes; the world said no; the world wins.

06

The handoff

At the end of a session the architect writes a pack for its successor. The temptation is to write a summary. A summary is the wrong artifact, because it invites trust it has not earned.

The rules that make a handoff safe:

Every number cites the block that measured it. Not "tests pass" but "3,945 passed, cold run, at this commit." Every snapshot is marked verify, do not trust. Estimates are labeled as estimates and never mixed with measurements. Correct the previous handoff explicitly when it was wrong, by name. Say what must not be claimed yet — a handoff that lets your successor over-promise to the owner has failed, however accurate the rest of it is.

And the first thing a new session does is not read the pack. It runs a read-only block that re-derives the facts from the machine itself, then compares. Any disagreement is resolved in favor of the measurement.

One more, learned expensively: at the start of every session, confirm each open item with the owner. Two items our handoffs had called "the single highest-value unlock" had been finished two months earlier. Nobody had asked. Two months of planning around work that was already done.

07

The principles, compressed

  1. The repository is the memory; the chat is disposable.
  2. Measure before you trust — including your own previous handoff.
  3. Every snapshot cites the measurement that produced it.
  4. The agent is usually right. Ask for its measurement, not its opinion.
  5. A gate that refuses is working. Read why before re-running it.
  6. Record approvals verbatim. An approval is only as good as the description it was given on.
  7. Name your errors in public, the moment you catch them.
  8. A proof that passes vacuously is worse than no proof. Every check needs a control.
  9. Validate instruments in the environment they will run in.
  10. Write the handoff for someone who will trust it — because they will, and the only defense is telling them not to.
08

If you start on Monday

The full system took many sessions to build, and building it up front would be building remedies for failures you have not had yet. Start with the parts that carry the most knowledge per hour.

Day one: the record files — state, decisions, lessons, protected paths. Commit them. That alone fixes knowledge evaporation, and it is an afternoon's work.

Day one: the block contract, and the rule that a measurement beats the handoff.

Day one: name your errors. This costs nothing and compounds faster than anything else on the list.

Week one: a pin file, and a check that the suite reaches it.

When you start running agents in parallel: the two-phase split, and reviewing from the branch rather than the working copy.

Later, and only after the incident that proves you need it: the authority gate, the pre-close checklist, the automated close-out. Write the incident into the gate's own comments, so the next person to find it inconvenient knows what it cost.

The point

None of this is about AI

Every rule here would improve a team of humans. What changes with agents is the rate. An agent produces plausible, fluent, confident output at a speed no human review process was designed for — and plausible output is exactly what this discipline exists to catch, because an agent tells you what it did. It does not tell you what it did not do.

The gate in the opening paragraph reported success for the entire life of the project. Everything downstream of it was built on that signal. Whether the code was fine, nobody could say — and that is the point. The evidence was hollow, and hollow evidence is invisible by construction: it looks exactly like the real thing until the day it matters.

Ship fast if you like. Just make sure that when something goes green, you know what it would have taken to go red.
This is the gap we hunt.

Plumbline runs an independent verification audit on AI-built software — not just whether your checks pass, but whether they’d fail if something were broken, and whether they ran at all. Through a relay you control, zero access to your systems.

Book a 20-min fit call
← All field notesplumblinehq.ai