Stop telling your AI agent to 'fix the bug'. Make it reproduce it first
A bug prompt in three parts, and a rule for when to throw the whole session away.
“Fix the login bug” is a terrible prompt for a coding agent, and a frustrating one to debug when it goes wrong. The agent will read some code, find something plausible, change it, and tell you it’s fixed. Sometimes it is. Sometimes it has patched a bug you didn’t have.
The cure is boring and old: reproduce the failure before changing anything. It works for agents for the same reason it works for people. A bug you can trigger on demand can be checked. A bug you can only describe can only be argued about.
The three-part bug prompt
Anthropic’s Claude Code best-practices guide recommends giving the symptom, the likely location and what “fixed” looks like, and its own example ends with “write a failing test that reproduces the issue, then fix it.” That’s the shape to copy, in any agent:
Users report that login fails after the session times out.
Symptom: after 30 minutes idle, the next request returns 401 and the
user is sent to the login page instead of getting a refreshed token.
Where to look: the token refresh logic in src/auth/.
First, write a failing test that reproduces this. Run it and show me
the failure output. Don't change any source code until the test fails
for the reason described above. Then make the smallest fix that turns
the test green, and run the full auth test suite.
Each piece does something. The symptom stops the agent guessing at the problem. The location keeps it from reading half the repo. The failing test is the check the agent can run, which matters because, as the same guide says, Claude stops when the work looks done, and without something it can run, “looks done” is the only signal available.
Why the red test is the important part
Watch the first failure output. It answers a question you’d otherwise have to take on trust: is the agent looking at the bug you’re looking at?
Three things tend to go wrong at this stage, and all of them are cheap to catch before any fix exists:
- The test fails for a different reason than the bug, such as a missing fixture or a bad import. The agent then fixes the test setup and calls it done.
- The test passes immediately. Either the bug isn’t where you thought, or the test doesn’t exercise it. Both are useful to know.
- The agent can’t reproduce it at all and proposes a fix anyway. That one is a refusal to accept evidence, and you should stop it.
If you can’t hand the agent a reproduction, say so in the prompt and ask it to find one: “Find the smallest input that triggers this and show it.” That’s an investigation task, so scope it tightly. The same guide warns that unscoped “investigate” requests make Claude read hundreds of files and fill its context.
Ask for the root cause, not the quiet
A failing build gives agents an easy way out: make the error stop. The guide’s advice is to ask for the fix to address the root cause and not suppress the error, and to verify. If a diff adds a try/catch, a skipped test, or a loosened assertion, read it twice.
Know when to start over
Here’s the rule worth borrowing. If you’ve corrected the agent more than twice on the same issue, the session is full of failed approaches, and they bias what comes next. Anthropic’s guidance is to run /clear and write a better opening prompt that includes what you’ve learned. In practice that means pasting in the reproduction, the two things that didn’t work, and why.
A clean session with a sharper prompt usually beats a long one with accumulated corrections. It feels wasteful to throw away forty turns. It’s still faster than the forty-first.
Next step: take the last bug you fixed by prompting and rewrite the prompt in this shape. If you couldn’t write the failing test, that’s the part of the bug you hadn’t understood yet.