jnhssg

🧪 Red, Green, Delegate - TDD in the Age of Coding Agents

Half a year ago I wrote about my CLAUDE.md, the file that tells Claude Code how I want code to be written. The file has changed a lot since then. Rules came and went, sections got merged, the persona is gone. One rule survived every rewrite and only got stricter: no production code without a failing test that demands it.

This post is about why that rule matters more today than it did when a single assistant sat next to me, and how test-driven development turned from a personal habit into the mechanism that holds my whole agentic workflow together.

What changed since March

In March, working with Claude Code meant a conversation: I ask, it writes, I read, we iterate. Today most of my coding sessions look different. A planning session on the strongest model I have reads the ticket, inspects the codebase and writes a plan precise enough that cheaper models can execute it. The implementation is then delegated to subagents, usually Opus, which work in parallel waves in their own git worktrees so they never step on each other's files. A review agent reads each wave's diff against the plan and the repository conventions. The main session only orchestrates: it runs the final gates, commits once and opens the pull request.

That is a lot of code written by something that is not me, often by several somethings at the same time. Agents are fast, fluent and confident, and confidence is not correctness. The question I ask myself stopped being "how do I write this?" and became "how do I know this works without reading every line?"

TDD turned out to be the answer, for reasons that go beyond the ones Kent Beck wrote down in 2002.

The rule, as it stands today

This is the implementation section of my global CLAUDE.md, verbatim:

Implement every task with TDD (red → green → refactor):

- Write a failing test first, run it, and confirm it fails for the expected reason.
- Write the minimal production code that makes it pass, then run the tests again.
- Refactor only while the tests stay green. Never write production code without a
  failing test that demands it.

Commit only at the end of the implementation, never after each task:

- Do not commit while working through a plan, even when the plan or a skill contains
  a per-task "Commit" step. Skip that step and keep the changes uncommitted in the
  working tree until every task is done.
- When all tasks are implemented and the full test suite passes, commit the complete
  implementation in a single commit (after any project-required pre-commit checks).
- This applies to delegated work as well: when handing tasks to subagents, instruct
  them not to commit.

That is the whole thing. Every clause earns its place, so let me go through them.

Confirm it fails for the expected reason

This is the clause I added last, and it is the one that catches the most problems. An agent that is told "write a failing test" will write a test and report that it fails. That is not the same as the test failing because the behavior is missing.

I have seen tests fail because of a typo in an import path, because a mock was never wired, because a fixture file did not exist. The agent then "fixes" the import, the test goes green, and nothing was verified. I have also seen tests pass on the first run because they assert something the code already does, or worse, assert nothing at all.

The expected reason is part of the specification. A failure like expected 401, received 200 tells me the test exercises the right path and the behavior is genuinely missing. A failure like Cannot find module tells me the agent's model of the codebase is off, which is a cheap and early signal that something else in its plan is probably off too.

Minimal production code

Agents gold-plate by default. Ask for a validator and you get a validator, a config option to disable it, a helper module and a README section. "Minimal code to make the test pass" bounds the blast radius: the agent may write exactly as much as the red test demands and not a line more. If I want the config option, I write a test for it.

This is the same rule I had in March. What changed is that I no longer enforce it by reading diffs. Code that goes beyond the test shows up as code no test covers, and the review agent flags it.

Refactor only while green

The third step is where agents are actually good. Extracting an explaining variable, turning a nested condition into a guard clause, deleting dead code: a model does these reliably and tirelessly. The constraint is that it only happens under a green suite and never mixes with a behavior change. I keep the Tidy First separation here: structural changes and behavioral changes are different steps, and a step that does both gets split or redone.

With that separation, refactoring becomes something I can delegate without worrying. The tests say whether behavior survived. The diff says whether the step did one thing.

Red and green as evidence

Here is the part that is specific to working with agents. When a subagent reports back, I do not read its prose first. I read the evidence.

Every task brief asks for the same report: the files changed, the spec names added, and for each spec the red run and the green run with the command and the relevant output lines. Then the commands run with their pass and fail counts, deviations from the plan with a reason, and open questions.

The red output is what makes the cycle auditable. It shows what the test originally demanded before any production code existed. If a test was later weakened to get it green, the red output and the final assertion no longer match, and that is visible without re-running anything. The review agent checks for exactly this. It once sent a report back because the red run was missing, and the coding agent had to produce it before the task counted as done.

This replaces trust with something cheaper and more reliable: a skim. I look at the spec names, I look at the red lines, I look at the green counts, and only then do I read the diff where the evidence looks thin.

The failing test is the contract between agents

In a wave with three subagents working in parallel, the plan needs a language that is precise enough for a model to execute without asking questions and compact enough to fit into a prompt. Test names turned out to be that language.

A task brief contains the contract, the file paths with line numbers, the TDD order and the spec case names, phrased as behavior: rejects expired tokens with 401, skips the sync when the URL is unset, writes the audit row in the same transaction as the delete. The agent's job is to make each of those names true, in order, red first.

Three things fall out of that:

  • The plan is reviewable before any code exists. If I cannot name the test, the requirement is not clear yet, and finding that out costs me a sentence instead of a wave of implementation.
  • Parallel agents cannot drift apart. Agent A's spec names are agent B's acceptance criteria. When B continues on A's worktree once A reports green, B knows exactly what it can rely on.
  • The review has a baseline. The review agent compares the diff against the brief, not against its own idea of what good code looks like. Findings are deviations from a written contract, which keeps the review short and actionable.

One commit at the end

Agents love to commit. Most planning skills come with a "commit" step after every task, and left alone a subagent will produce a dozen commits with messages like "add test" and "fix test". None of that history helps anyone.

So the rule is: nobody commits until everything is done. Subagents are told explicitly, in every brief, not to commit and not to push. The changes stay in the working tree of the worktree until the full suite, the linter and the type-check pass on the whole implementation. Then the main session commits once and opens the PR.

The benefits are practical. A wave that went wrong is discarded with a checkout instead of an interactive rebase. The reviewer, human or otherwise, gets one coherent diff. And the history of the repository records what changed, not how the agents got there.

What it buys me

Putting it together, this is what the discipline pays back, roughly in the order I noticed it:

  • Specification before generation. The failing test pins the intent before the first line of production code exists. A vague prompt becomes an executable contract that a model cannot misread the way it misreads prose.
  • Scope control. The agent writes what the red test demands. Gold plating, imagined requirements and helpful extras have nowhere to hide.
  • Evidence instead of trust. Red and green output are facts I can check in seconds. Fluent prose about what was done is not.
  • Wrong assumptions surface early. A test that fails for an unexpected reason tells me the agent misunderstood the codebase while the mistake is still one file big.
  • Reviewable diffs. Small red-green increments, structural and behavioral steps kept apart, one commit per task: the pull request is something a colleague can actually read.
  • Delegation at scale. Tests as contracts make it safe to hand the implementation to cheaper and faster models, in parallel. The expensive model plans and reviews, the routine typing goes elsewhere, and the suite is the coordination mechanism, not me.
  • Fearless refactoring. The step agents are best at is the step that is dangerous without tests. With a green suite, it is free.
  • Design feedback. Hard to test in isolation still means doing too much. When an agent struggles to write the test, the design is usually what needs to change, and I hear about it before the code is written.
  • Living documentation. Behavior-spec names describe what was built better than any PR description. A reviewer reads the test names first and often does not need more.
  • Less rework, fewer tokens. The expensive part of agentic coding is not generation. It is undoing. Catching a misunderstanding at the red step costs a test; catching it at review costs a wave.

And one that I did not expect: it keeps me in the loop in the right place. I still decide what the software should do, because I approve the test cases. I just stopped typing the implementation.

Where the rule bends

Not everything gets a failing test first. Thin glue code, layout-only components, one-off scripts and spikes are exempt, the same as in March. Bug fixes are not exempt: a bug fix starts with a test that reproduces it, every time. This blog post is content, not code, so it got no test either.

A few things I learned the hard way about the tooling, which matter more when an agent reads the output instead of me:

  • Trust exit codes, not summaries. A spec file that fails to load contributes zero tests. A summary line that counts tests then looks perfectly green while the run actually failed. An output wrapper that condenses test results bit me with exactly this, and the fix was a rule: check the exit code and the failed-suite count, always.
  • Type-check separately. Vitest strips types, so a spec that would never compile can run green. A spy typed against the wrong overload of a library method passed the suite and failed the type-check. Both run before the commit now.
  • Know your runner's flags. With a recent pnpm, pnpm run test -- path/to/spec.ts hands the double dash on to Vitest, which then ignores the path and runs everything. For me that is thirty seconds of noise. For an agent, it is 3,300 test results wrapped around the one line it needed.
  • A test job can die without a verdict. Runners crash, processes get killed, and the log ends without a summary. That is a rerun, not a bug hunt, and the agent needs to know the difference.
  • Agents weaken tests when they are stuck. Changing the assertion to match the output is a tempting shortcut for a model under pressure to report green. The rule forbids it, the red evidence exposes it, and the review agent looks for it.

Closing

Test-driven development was invented for people who needed feedback faster than they could make mistakes. Coding agents make mistakes faster than anyone ever has, and they make them with great confidence. The feedback loop that was a nice discipline for a human became the load-bearing structure of a workflow where I write almost none of the code myself.

Red, green, refactor is still the cycle. The new step is the one in the title: delegate, and let the tests tell you whether you can.

🔧

Claude Code, CLAUDE.md, TDD, Subagents, Git worktrees, Vitest

📚

Inspired by: Kent Beck's "Test-Driven Development by Example" and "Tidy First?"