A separate evaluator agent is what makes long agent runs worth the money
Anthropic's harness splits planning, building and testing across agents. The cost jump was 20x, and the reasons it paid off are specific.
Ask one agent to build a whole application and two things tend to go wrong. It loses the thread as its context fills, and it grades its own work generously. Anthropic’s engineering team wrote up a harness design for long-running application development that attacks both, and the numbers are honest about what it costs.
The two failures
The first is coherence. As the window fills, models drift, and Anthropic describes a behavior it calls context anxiety: an agent that thinks it’s near its limit starts wrapping up early.
The second is self-assessment. Asked to evaluate their own output, agents tend to praise it confidently even when quality is mediocre, and they miss bugs that are straightforward to verify.
Three agents with separate jobs
The fix is to split the roles.
The planner takes a brief prompt of one to four sentences and expands it into a detailed product spec. It focuses on deliverables and avoids prescribing implementation, so an early technical mistake doesn’t cascade.
The generator builds features iteratively. Before each chunk of work, it agrees a contract with the evaluator about what “done” means.
The evaluator tests the running application with Playwright, clicking through the UI and checking database state. It grades against four criteria (design quality, originality, craft and functionality) with hard thresholds. Fail one, and the work goes back with a detailed bug report. A sample from the post reads like a QA ticket: the rectangle fill tool lets you click and drag, but only places tiles at the drag’s start and end points. Fail.
Reset or compact?
For one model, Sonnet 4.5, context resets worked better than compaction. A reset clears the conversation completely and gives a clean slate, at the cost that the handoff artifact has to carry enough state for the next agent to continue. Compaction summarizes in place. Opus 4.6 largely removed the anxiety, which let Anthropic run builds in a single continuous session.
The lesson isn’t “always reset”. It’s that handoff files are the real interface between agents, and you should write them as if the reader has no memory, because it doesn’t.
What it costs
| Task | Setup | Time | Cost |
|---|---|---|---|
| Retro game maker | Solo agent | 20 minutes | $9 |
| Retro game maker | Full harness | 6 hours | $200 |
| Digital audio workstation (simplified harness) | Planner, generator, evaluator | 3 hours 50 minutes | $124.70 |
That’s about 20 times the cost for the game maker. Anthropic’s argument is that the harness is justified when the task sits beyond what the model does reliably alone, and that the solo version’s output was visibly worse.
Still imperfect
The post doesn’t oversell. Even the improved harness missed usability problems. In the audio workstation, the workflow didn’t make it clear that you should build sprites and entities before populating a level. The authors say the output showed the limits of the model’s QA ability: small layout issues, unintuitive interactions and bugs in deeper features that went undiscovered. Tuning the evaluator took several rounds.
What to take from it
The sentence worth keeping: every component in a harness encodes an assumption about what the model can’t do on its own, and those assumptions are worth stress testing. Newer models may not need a component you built for an older one.
For your own work, the transferable parts are modest. Have something other than the author check the result. Define “done” before building. Check against the running thing, not the code. And price it first. We saw the same pattern in the C compiler experiment : the harness mattered more than the individual agent.
Next step: the next time you run a long agent task, put the acceptance criteria in a file before it starts, and have a second session test the result against that file.