Development How-to

Pin legacy behaviour before an agent refactors it, then grade the tests with mutation testing

An agent's first test suite for old code usually checks what you'd expect, not what the code does. Mutation testing shows the difference in about a second.

A steel vernier caliper clamped on a small metal block with a chipped corner, teal halftone dots gathered around the chip

You’re about to refactor a shipping calculator that nobody on the team wrote and nobody wants to touch. You ask an agent for tests first, it hands back four neat ones, they pass, and you start the refactor with a green light that means almost nothing. Four tests can’t pin forty decisions.

The fix is two steps. First, ask for characterization tests: tests that record what the code does today, bugs included, rather than what it should do. Second, grade those tests with mutation testing, which breaks the code in small ways and counts how many breakages your tests notice. That grade is the part an agent can’t talk its way around. I ran both steps in a scratch repo this morning, and the numbers below are what it printed.

The module and the usual first attempt

Here’s the scratch module. It’s invented, but it has the traits that make legacy code hard to test: magic numbers, a special case buried in the middle, and behaviour no spec mentions.

def shipping_cost(weight_kg, country, express=False, subtotal=0.0):
    """Legacy shipping calculator. Nobody remembers why it works this way."""
    if weight_kg <= 0:
        return 0.0
    if subtotal >= 100 and country == "GB" and not express:
        return 0.0
    base = 4.99 if weight_kg <= 1 else 4.99 + (weight_kg - 1) * 1.5
    if weight_kg > 20:
        base = base + 15
    if country != "GB":
        base = base * 2.5
    if express:
        base = base + 7.5
    return round(base, 2)

Ask an agent “write tests for shipping_cost” and you’ll typically get tests of the behaviour a reasonable person would expect. Here’s the kind of suite I mean, which I wrote to stand in for that:

from shipping import shipping_cost

def test_light_uk():
    assert shipping_cost(0.5, "GB") == 4.99

def test_international_is_dearer():
    assert shipping_cost(0.5, "US") > shipping_cost(0.5, "GB")

def test_express_costs_more():
    assert shipping_cost(2, "GB", express=True) > shipping_cost(2, "GB")

def test_free_over_100():
    assert shipping_cost(2, "GB", subtotal=150) == 0.0
$ python -m pytest -q
....                                                                     [100%]
4 passed in 0.01s

Every test reads like a requirement, and none of them would catch most refactoring mistakes. Two of them use inequalities, so changing 2.5 to 3.0 for international orders sails through. Nothing touches the 20 kg surcharge. Nothing checks the free-shipping threshold at 100 exactly.

Grade the suite before you trust it

Mutation testing tools take your source, make one small change at a time (flip a <= to <, swap a + for a -, change a constant), and run your tests against each variant. A variant your tests fail on is killed. One they pass on survived, which means your suite can’t see that change.

I used mutmut , a BSD-3-Clause Python tool. Its README says it needs a system with fork support, so on Windows that means WSL. Setup is one install and a few lines of config:

# pyproject.toml
[tool.mutmut]
source_paths = ["shipping.py"]
pytest_add_cli_args = ["-p", "no:cacheprovider"]
pytest_add_cli_args_test_selection = ["tests/"]
pip install mutmut pytest
mutmut run
mutmut results

Against the four-test suite, mutmut generated 45 mutants. It killed 23 and 22 survived. The survivor list looks like this (trimmed):

$ mutmut results
    shipping.x_shipping_cost__mutmut_3: survived
    shipping.x_shipping_cost__mutmut_5: survived
    shipping.x_shipping_cost__mutmut_8: survived
    ...
    shipping.x_shipping_cost__mutmut_45: survived

Out of 45 single-character breakages, 22 would have gone through. If the agent’s refactor introduced any of them, the suite would have stayed green. That’s the number to show your agent.

Write the characterization suite

The method has two moves: probe the code, then pin what it returned. The probe is a throwaway script that calls the function with inputs chosen around every branch and boundary. You can write it yourself, or have the agent write it from this prompt:

Read shipping_cost in shipping.py. Do not change it.

Write a throwaway script, probe.py, that calls it with inputs that
exercise every branch and both sides of every boundary (weights at 0,
1, just over 1, 20, just over 20; each country case; express on and
off; subtotal at 99.99 and 100; combinations of express and free
shipping). Print each input with the return value.

Run it and show me the output. Then turn the output into a parametrized
pytest file where the expected values are exactly what the function
returned. Do not correct anything you think is a bug. Instead, list
every result that looks wrong at the end, separately.

The last instruction is the one that matters. An agent that sees a number it dislikes will quietly adjust the expectation to what it thinks the code should say, and now your safety net disagrees with production. You want the net to match production first. Opinions come later and go in a separate list.

Here’s the probe output from my scratch run, trimmed to the interesting rows:

(0, 'GB') 0.0
(-1, 'GB') 0.0
(1, 'GB') 4.99
(1.01, 'GB') 5.0
(20, 'GB') 33.49
(20.01, 'GB') 48.51
(0.5, 'US') 12.48
(2, 'GB', True) 13.99
(2, 'GB', False, 100) 0.0
(2, 'GB', False, 99.99) 6.49
(2, 'GB', True, 500) 13.99
(1, 'gb') 12.48

Several of these would make a reviewer stop. A negative weight ships free. Express orders never get the free-shipping threshold, even at 500. And a lowercase "gb" is charged as international, so the same parcel to the same country costs 12.48 instead of 4.99. Those are the things you’d have “fixed” while refactoring, and each fix would have changed what customers pay.

The pinned file the agent produces is plain pytest. This is the shape, using pytest.mark.parametrize, which runs the test once per tuple:

import pytest
from shipping import shipping_cost

# Pinned from what the code returned when probed. Not from a spec.
CASES = [
    ((0, "GB"), 0.0),
    ((-1, "GB"), 0.0),
    ((1, "GB"), 4.99),
    ((1.01, "GB"), 5.0),
    ((20, "GB"), 33.49),
    ((20.01, "GB"), 48.51),
    ((0.5, "US"), 12.48),
    ((2, "GB", True), 13.99),
    ((2, "GB", False, 100), 0.0),
    ((2, "GB", False, 99.99), 6.49),
    ((2, "GB", True, 500), 13.99),
    ((1, "gb"), 12.48),
    # ...the other rows from the probe
]

@pytest.mark.parametrize("args,expected", CASES)
def test_pinned(args, expected):
    assert shipping_cost(*args) == expected

Keep the comment. In six months someone will see ((1, "gb"), 12.48) and file a bug, and the comment tells them the value was recorded on purpose.

Run the grade again

My pinned file had 16 cases. I removed the four happy-path tests, confirmed the 16 passed, and ran mutmut again:

$ python -m pytest -q
16 passed in 0.02s
$ mutmut run
...
45/45  killed 43  survived 2
$ mutmut results
    shipping.x_shipping_cost__mutmut_2: survived
    shipping.x_shipping_cost__mutmut_19: survived

From 23 killed to 43. Then the useful part: look at the two survivors. mutmut show prints the diff for a mutant:

$ mutmut show shipping.x_shipping_cost__mutmut_2
-def shipping_cost(weight_kg, country, express=False, subtotal=0.0):
+def shipping_cost(weight_kg, country, express=False, subtotal=1.0):

$ mutmut show shipping.x_shipping_cost__mutmut_19
-    base = 4.99 if weight_kg <= 1 else 4.99 + (weight_kg - 1) * 1.5
+    base = 4.99 if weight_kg < 1 else 4.99 + (weight_kg - 1) * 1.5

The first changes a default from 0.0 to 1.0. No input can tell them apart, because the only thing the code does with subtotal is compare it to 100. The second changes <= 1 to < 1, but at exactly 1 kg the else branch computes 4.99 + 0 * 1.5, which is 4.99 either way. Both are equivalent mutants: different source, identical behaviour. No test can kill them, and you shouldn’t try.

That’s the stopping rule. Survivors you can’t distinguish from the original are noise. Survivors where the mutated code returns something different are missing tests. Read each diff and decide which kind it is. If you want the agent to do the sorting, hand it the survivor diffs and ask it to either write a test that kills each one or explain why none can exist, and check its explanations yourself.

Make it yours

The idea doesn’t depend on Python. I only ran mutmut, so if you’re on another stack, find that ecosystem’s mutation tester and check its setup against its own docs.

Slow suites are the main cost. Mutation testing runs your tests once per mutant, so a suite that takes a minute makes 45 mutants take most of an hour. Point it at one module, as the source_paths line above does.

If the function has side effects or reads a clock, pin its inputs first. A characterization test on code that calls datetime.now() is a flaky test with extra steps. Inject the clock or freeze it, then pin.

For output that’s a document rather than a number (a rendered email, a generated config), store the full output as a snapshot file and compare against it. Same idea, bigger values.

What goes wrong

Characterization tests pin bugs. That’s by design, and it’s also the biggest risk: someone treats the green suite as endorsement. Keep the “looks wrong” list the agent produced and file it somewhere your team will find it.

They also only cover the inputs you probed. A probe that skips a country code, a unicode string or a float with awkward rounding leaves that behaviour unprotected, and mutation testing can’t flag a branch of behaviour your mutants never touch. It grades the tests against the code that exists, not against the inputs real customers send. If you have production logs, sample real inputs into the probe.

Mutmut left a mutants/ directory behind in my scratch repo, and plain pytest then failed to collect with an “import file mismatch” error, because it found a second copy of the test files. Deleting mutants/ fixed it. Add it to .gitignore.

And don’t bother if the code is about to be deleted, or if it’s already well covered by tests someone trusts. This earns its cost on code that’s old, untested and about to be changed.

The first fifteen minutes

Pick one function you’re afraid to refactor. Install mutmut, point source_paths at its file, and run it against whatever tests exist today to see the survivor count. Then give your agent the probe prompt above, with the instruction to pin results rather than fix them, and run mutmut again.