Assistants Analysis

What 16 parallel agents building a C compiler teach about running agents at scale

Anthropic's engineering team spent nearly $20,000 on it. The lessons about tests, locks and duplicate code transfer to much smaller jobs.

Anthropic’s engineering team handed 16 Claude instances a goal, a shared repo and a test suite, and let them write a C compiler in Rust. The result, described in an engineering post , is a roughly 100,000-line compiler that can build Linux 6.9 on x86, ARM and RISC-V, plus projects like SQLite, postgres and redis. It took nearly two weeks, close to 2,000 Claude Code sessions, 2 billion input tokens and 140 million output tokens, for just under $20,000.

You’re not going to build a compiler. The parts worth taking away are the mechanics, because they apply to any job where more than one agent works on one codebase.

Locks in plain git

Coordination was simple. An agent claimed a task by writing a text file into a current_tasks/ directory. If two agents tried to claim the same task, git’s own synchronization forced the second to pick another. When an agent finished, it pulled upstream, merged other agents’ changes, pushed and removed its lock. A loop restarted it on the next task with no human in between.

No scheduler, no message bus. For small teams of agents, files and git are enough, and every step is inspectable in the history.

Write the test harness for the agent, not for you

This is the central lesson. The author says that Claude works autonomously on whatever problem it’s given, so the task verifier has to be nearly perfect. And they kept having to remind themselves that they were writing the harness for Claude, not themselves.

In practice that meant:

  • Don’t print thousands of useless lines. Log the detail to a file, and print little, because output eats the context window.
  • Provide a fast mode, a 1% or 10% random sample of the tests, so an agent can iterate without running everything.
  • Use a known-good oracle where you can. Here GCC served as a reference to compare against when compiling the Linux kernel.

That last point is a general trick. If you’re migrating or rewriting something, a reference implementation turns “is this right?” into a diff.

One giant task defeats parallelism

When the work was “compile the Linux kernel,” every agent hit the same bug, fixed it, and overwrote the others’ fixes. Parallel agents need work that decomposes into independent pieces. The oracle approach above was how the team split a monolithic task into smaller ones.

Give one agent the cleanup job

The author notes that LLM-written code often re-implements functionality that already exists, so they assigned one agent to coalesce duplicate code. If you’ve reviewed a large agent-generated diff, you’ve seen this. A standing “find the duplicates” role is cheaper than catching them one at a time.

What it couldn’t do

The post is candid about limits. There’s no 16-bit x86 code generator, so that part calls out to GCC. The assembler and linker were the last pieces Claude automated and are still somewhat buggy. The generated code is less efficient than GCC’s output with all optimizations disabled, and the Rust is reasonable but nowhere near what an expert would write.

The author’s closing caution is the one to keep: for autonomous systems, it’s easy to see the tests pass and assume the job is done, when that’s rarely the whole story. A harness this good was possible because the problem has an enormous existing test suite. Most of your code doesn’t.

Next step: if you plan to run more than one agent against a repo, write down how each will pick a task, how it will prove it’s finished, and what stops two of them touching the same file.