Use mutation testing to delete code, not to score tests

Ask an AI agent for one behaviour and you often get that behaviour, plus a null check, a retry, some input validation and a test for each. None of it is wrong on its own. All of it is code you now own and didn’t ask for.

Mutation testing is usually sold as a way to score your tests. It changes your code in small ways (flips a condition, removes a line) and checks whether any test notices. A change no test notices is a surviving mutant, and the usual reading is that your tests are weak.

I read it the other way round.

A surviving mutant is a question

When a mutant survives in code the agent just wrote, I don’t ask “which test should I add?” I ask: was this behaviour asked for?

The measure is the scenarios I signed off before any code existed. If a scenario needs the behaviour, it needs a test, and the mutant is telling me one is missing. If no scenario needs it, the code goes.

That includes the defensive code the agent added by itself. If it really matters, it becomes a scenario in a later step and earns its way back in with a test first.

Keep it small enough to run

The run is on changed lines only, after each outer loop goes green, or after four unit tests, whichever comes first. A whole-project mutation run is the reason people say “that’s not practical”. A run on the lines you just changed is a different thing.

What it depends on

This only works if the scenarios exist first. Without a signed-off list of what was asked for, a surviving mutant is just an opinion. If you don’t have a plan file, use the acceptance criteria on the ticket as the stand-in.

The full flow is on the method page.

Get the next one