Nothing more, nothing less.

You and the agent agree the scope as a few plain-language scenarios. The agent builds it test-driven from the outside in. Mutation testing deletes anything nobody asked for, and refactoring happens at every step, never saved for the end.

This is how I work today, not a finished spec. The hooks in particular are experimental.

1. Plan with the agent, then sign off

You explain the requirement in as much detail as it needs. The agent explores the code first and reports risks, impacts and unknowns. The output is a short plan file with two things in it:

Scenarios are grouped into steps. A step holds one to three scenarios, takes one working session, ends green with something of value, and can be its own pull request. The first step goes thin through every layer, like a tracer bullet. Later steps build each layer out a little at a time. You manage the branches and start each step. The agent never starts the next one itself.

2. Test-driven, not tests-first

The scenarios are a guide, not a list of tests to write. Take the first scenario and write it as a real test at the real user boundary: an HTTP request, a rendered page or a CLI call, through the real stack. Then let the tests drive everything after it.

"Tests first" means writing a batch of tests, then the code, then tidying at the end. "Test-driven" means each test result decides the next small move, and you refactor at every step.

3. The double loop: listen to the failure

The outer loop is the failing acceptance test. Its message tells you which unit or role to test next. The inner loop is the failing unit test. Its message tells you the smallest line of code to write next.

This is the rule I'd keep if I had to drop everything else: read the test output to decide what to build next, not what the agent thinks should be delivered. When the error says "no such route", you add the route and nothing else. Then you run the tests again.

4. Mutation testing as the scope gate

After each double loop goes green, or after four unit tests, whichever comes first, run mutation testing on only the lines changed since the last run.

A surviving mutant is a question: was this behaviour asked for? The measure is the signed-off scenarios. If none of them needs the behaviour, the code goes. That includes defensive code the agent added on its own, such as a null check, a retry or input validation. If it matters, it becomes a scenario in a later step.

Other people use mutation testing to score how good the tests are. Here it deletes code, and it's what enforces "nothing more".

5. Refactor as you go, with the tests frozen

Refactor only when everything is green, and change only production code. Tests change only during red to green. After a refactor, every test must still pass unchanged. A rename or a move that touches a test is fine. A change to what a test asserts is not.

How each loophole is closed

What goes wrongWhat closes it
The agent builds more than was askedMutation testing on changed lines, measured against the signed-off scenarios
The agent writes a test to protect code nobody asked forEvery test must trace back to a scenario
The agent delivers less: a hard-coded return, a stub, only the happy pathThe outer test runs through the real stack, and a person starts each step
The agent quietly changes the scopeIt may fix wording, data, order and detail inside a scenario. Adding or removing behaviour anyone can observe needs your sign-off
The agent changes tests so they passTests are frozen during refactor, and a hook checks it (experimental)

What's new here, and what isn't

Outside-in development and the double loop aren't mine. They come from Dave Farley's acceptance-test-driven development and from Growing Object-Oriented Software, Guided by Tests. What I've added is:

Get the next part

Each part of the flow will get its own write-up and video, and the skills will be in a public repo. Leave your email and I'll send one message per release.