Models Analysis

Claude Opus 5.5 costs 20% less per token and Anthropic says it codes better

The numbers behind the release, what changes in practice, and why Anthropic's own benchmark caveat is worth reading.

Anthropic released Claude Opus 5.5 on September 22. The model id is claude-opus-5-5, it’s on the Claude Platform, AWS, Google Cloud and Azure, and it’s cheaper than the model it replaces. The announcement says it does most work at the level of Claude Fable 5.1 while costing 40% less to run than Opus 5.

The 40% figure is for typical workloads. List prices are a smaller cut: $4 per million input tokens and $20 per million output, against $5 and $25 for Opus 5. Cache reads drop furthest, from $0.50 to $0.20 per million. If your agent loops re-read a big repo context on every turn, that’s the line that matters. A fast mode is also listed at $8 input and $40 output, which Anthropic describes as 2.5x faster.

Coding numbers

These are Anthropic’s own results, so treat them as claims until you’ve run your own tasks.

BenchmarkOpus 5.5Fable 5.1Opus 5
Terminal-Bench 4.0 (xhigh effort)66.4%55.8%52.3%
FrontierCode v1.154.4%50.3%48.0%
CursorBench 4.057.8%51.8%46.6%

Anthropic also compares against OpenAI models on cost per result, claiming it beats GPT-6 Astra on FrontierCode at roughly a fifth of the cost. Those comparisons are the vendor’s, and cost-per-result depends on how many tokens a model burns getting there. Reproduce them on your own repo before you rebuild a pipeline around them.

The post adds that Opus 5.5 is 30% faster than Opus 5 at default settings, and quotes a tester who ran a 680,000-line code migration in under a day. It also says GitHub found it solved more terminal tasks than Opus 5 in less than half the steps.

What changes when you use it

Thinking effort has five levels: Low, Medium (the default), High, Xhigh and Max. The Terminal-Bench figure above was measured at xhigh, so don’t expect it at the default.

Two restrictions to plan around. Opus 5.5 is no longer available with thinking switched off, so code that turned it off for latency will need a rethink. And Anthropic routes cybersecurity tasks to Claude Opus 4.8, so security tooling that sends exploit-adjacent prompts to the newest model may see different behavior.

Anthropic also says the model is less likely to take hard-to-reverse actions and more resistant to prompt injection than Opus 5. For anyone letting an agent run shell commands, that’s the claim to test hardest.

The caveat Anthropic includes

The most useful sentence in the announcement is a disclaimer: at these capability levels, benchmark margins have become a less reliable guide to real-world differences. It also says the model often suspects it’s being evaluated. Both are reasons to keep a small set of your own tasks, with known good outcomes, as a regression suite.

Copilot users

If you pick models in GitHub Copilot, GitHub’s October 2 changelog deprecates Claude Opus 4.7 and suggests Opus 5.5 as the replacement. Enterprise admins may need to enable the new model in Copilot policy settings.

Next step: swap claude-opus-5-5 into one non-critical pipeline at default effort and compare cost and pass rate against your current model for a week.