Models Analysis

That 2-point benchmark lead might just be a bigger server

Anthropic measured how much infrastructure settings move agentic coding scores. It's more than many leaderboard gaps.

Every model launch comes with a table of coding benchmark scores, and the gaps are often a few points. Anthropic’s engineering team published a study that should make you read those gaps more carefully. In its infrastructure noise analysis , it found that how much compute an eval runs on can move agentic coding scores by more than the leaderboard gap between top models.

What they tested

The team ran Terminal-Bench 2.0 under six resource configurations, from strict enforcement of each task’s stated limits (1x) to no cap at all. Same models, same tasks, different headroom.

The infrastructure error rate, meaning runs that failed because of the environment and not the agent’s work, fell as headroom grew:

  • 1x, strict limits: 5.8%
  • 3x headroom: 2.1%
  • Uncapped: 0.5%

Between 1x and 3x the success rate stayed within noise (p = 0.40), so extra resources there only removed spurious failures. Beyond 3x, they changed what agents could solve. Uncapped runs scored about 6 percentage points higher than 1x (p < 0.01).

One example from the post: on a task like bn-fit-modify, a model’s default approach is to install a full Python data science stack. That fails under tight memory and works under generous memory. Whether the agent “can do the task” depended partly on the machine.

SWE-bench showed the same pattern at smaller scale. Across 227 problems with 5x variation in RAM, the highest allocation scored only 1.54 percentage points higher.

What it changes about reading scores

Anthropic’s own guidance is that leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched. That cuts against a lot of marketing. A model release that leads by two points on an agentic benchmark may be telling you something about the harness.

The practical takeaways, for people who read these tables to choose tools:

  • Check whether a score names its resource limits and kill thresholds. If it doesn’t, treat it as a rough indicator.
  • Be wary of cross-vendor comparisons where each side ran its own harness. We noted the same concern in our Opus 5.5 piece , where the vendor’s own numbers included comparisons against other labs’ models.
  • Remember that agentic evals measure the model and its environment together. A tool that gives agents more room to work may score higher without being better at reasoning.

If you run your own evals

You probably should, with a small set of tasks from your own repos. This study suggests two rules. Fix the resources and write them down, so a score from March means the same as one from October. And run each task more than once when the result matters, because infrastructure variance and model variance both show up as noise.

If your internal comparison between two tools differs by two points, you haven’t learned much. If it differs by twenty on tasks you care about, you have.

Next step: the next time a release leads with a benchmark chart, look for the line that says how the evaluation ran. If you can’t find it, discount the gap.