Make Claude Code run your tests before it's allowed to say it's done
A Stop hook turns 'please run the tests' from a request into a rule. Here's a working setup and the trap to avoid.
Telling Claude Code to run the tests in CLAUDE.md is a request. It usually works, and sometimes it doesn’t. If you want it to work every time, Claude Code has a mechanism for that: a Stop hook, a script that runs when Claude thinks it’s finished and can refuse to let it stop.
Anthropic’s best-practices guide puts the underlying problem plainly: Claude stops when the work looks done. Without a check it can run, “looks done” is the only signal available, and you become the verification loop. A Stop hook is the deterministic version of that check. It blocks the turn from ending until your script is satisfied.
A hook that runs your tests
Register the hook in .claude/settings.json for one project, or ~/.claude/settings.json for all of them:
{
"hooks": {
"Stop": [
{
"hooks": [
{ "type": "command", "command": ".claude/hooks/require-green-tests.sh" }
]
}
]
}
}
Then the script. This one is ours, built from the documented fields, so run it against a scratch branch before you trust it. Swap npm test for your own command.
#!/bin/bash
INPUT=$(cat)
# Already continuing because of this hook: let Claude stop.
if [ "$(echo "$INPUT" | jq -r '.stop_hook_active')" = "true" ]; then
exit 0
fi
# Nothing changed, so there's nothing to verify.
if git diff --quiet && [ -z "$(git ls-files --others --exclude-standard)" ]; then
exit 0
fi
if ! OUT=$(npm test --silent 2>&1); then
REASON=$(printf 'Tests are failing. Fix them before finishing.\n\n%s' "$(echo "$OUT" | tail -40)")
jq -n --arg r "$REASON" '{decision: "block", reason: $r}'
fi
exit 0
When the script prints {"decision": "block", ...}, Claude doesn’t stop. The reason becomes its next instruction, so the last 40 lines of test output go straight back into the loop.
The two details that matter
Stop hooks fire whenever Claude finishes responding, not only when a task is complete. The docs are explicit about that. Ask Claude a question about the codebase and the hook still runs. The git diff check above is there so a pure Q&A turn doesn’t trigger a full test run.
The second is the loop trap. Claude Code overrides a Stop hook after it blocks eight times in a row with no tool call from Claude in between, and ends the turn with a warning. The documented fix is the stop_hook_active check at the top of the script. It has a cost: if the fix attempt after a block also fails, Claude gets to stop with red tests. If your suite genuinely needs more than eight rounds to converge, you can raise the limit with the CLAUDE_CODE_STOP_HOOK_BLOCK_CAP environment variable instead.
When a script isn’t enough
Two other hook types exist for judgment calls. A prompt hook asks a model whether the work is done. An agent hook spawns a subagent that can actually run commands. This is the documented example for tests:
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "agent",
"prompt": "Verify that all unit tests pass. Run the test suite and check the results. $ARGUMENTS",
"timeout": 120
}
]
}
]
}
}
The docs mark agent hooks as experimental and recommend command hooks for production, which is why the script above is the one to start with.
Hook or CLAUDE.md?
CLAUDE.md instructions are advisory. Hooks are deterministic. Use the file for preferences and the hook for anything you’d be annoyed to find skipped. If you only need a check for one long unattended run, the guide also describes a /goal condition, which has a separate evaluator re-check after every turn.
Next step: add the hook on a throwaway branch, break a test on purpose, and watch what Claude does with the failure output.