Nothing more, nothing less.
You and the agent agree the scope as a few plain-language scenarios. The agent builds it test-driven from the outside in. Mutation testing deletes anything nobody asked for, and refactoring happens at every step, never saved for the end.
This is how I work today, not a finished spec. The hooks in particular are experimental.
1. Plan with the agent, then sign off
You explain the requirement in as much detail as it needs. The agent explores the code first and reports risks, impacts and unknowns. The output is a short plan file with two things in it:
- Scenarios in given/when/then form. You sign these off. They are the contract for what gets built.
- Gotchas: things worth knowing before you start, such as a rate limit or an odd legacy path.
Scenarios are grouped into steps. A step holds one to three scenarios, takes one working session, ends green with something of value, and can be its own pull request. The first step goes thin through every layer, like a tracer bullet. Later steps build each layer out a little at a time. You manage the branches and start each step. The agent never starts the next one itself.
2. Test-driven, not tests-first
The scenarios are a guide, not a list of tests to write. Take the first scenario and write it as a real test at the real user boundary: an HTTP request, a rendered page or a CLI call, through the real stack. Then let the tests drive everything after it.
"Tests first" means writing a batch of tests, then the code, then tidying at the end. "Test-driven" means each test result decides the next small move, and you refactor at every step.
3. The double loop: listen to the failure
The outer loop is the failing acceptance test. Its message tells you which unit or role to test next. The inner loop is the failing unit test. Its message tells you the smallest line of code to write next.
This is the rule I'd keep if I had to drop everything else: read the test output to decide what to build next, not what the agent thinks should be delivered. When the error says "no such route", you add the route and nothing else. Then you run the tests again.
4. Mutation testing as the scope gate
After each double loop goes green, or after four unit tests, whichever comes first, run mutation testing on only the lines changed since the last run.
A surviving mutant is a question: was this behaviour asked for? The measure is the signed-off scenarios. If none of them needs the behaviour, the code goes. That includes defensive code the agent added on its own, such as a null check, a retry or input validation. If it matters, it becomes a scenario in a later step.
Other people use mutation testing to score how good the tests are. Here it deletes code, and it's what enforces "nothing more".
5. Refactor as you go, with the tests frozen
Refactor only when everything is green, and change only production code. Tests change only during red to green. After a refactor, every test must still pass unchanged. A rename or a move that touches a test is fine. A change to what a test asserts is not.
How each loophole is closed
| What goes wrong | What closes it |
|---|---|
| The agent builds more than was asked | Mutation testing on changed lines, measured against the signed-off scenarios |
| The agent writes a test to protect code nobody asked for | Every test must trace back to a scenario |
| The agent delivers less: a hard-coded return, a stub, only the happy path | The outer test runs through the real stack, and a person starts each step |
| The agent quietly changes the scope | It may fix wording, data, order and detail inside a scenario. Adding or removing behaviour anyone can observe needs your sign-off |
| The agent changes tests so they pass | Tests are frozen during refactor, and a hook checks it (experimental) |
What's new here, and what isn't
Outside-in development and the double loop aren't mine. They come from Dave Farley's acceptance-test-driven development and from Growing Object-Oriented Software, Guided by Tests. What I've added is:
- Mutation testing used to decide scope, so it deletes code instead of scoring tests.
- Letting the test output decide the next design move, not the agent's idea of the feature.
- Refactoring at every step with the tests frozen, rather than at the end.
- A signed-off plan whose steps are each one session and one pull request.
- A strict flow, checked by hooks.
Get the next part
Each part of the flow will get its own write-up and video, and the skills will be in a public repo. Leave your email and I'll send one message per release.