Moving to Sonnet 5.5? Four requests that now return a 400
Same price as Sonnet 5, faster, and better at agentic coding by Anthropic's numbers. But the API breaks in ways a model-ID swap won't catch.
Claude Sonnet 5.5 shipped on September 28 at the same price as Sonnet 5: $2 per million input tokens and $10 per million output. Anthropic says it runs 30% or more faster. If you call the API directly, though, changing claude-sonnet-5 to claude-sonnet-5-5 isn’t the whole migration. Several requests that worked yesterday now fail.
Claude Code users get the model without touching anything. Version 2.1.284 made it the default Sonnet, per the Claude Code changelog . This piece is for people with their own agent loops, evals or tooling built on the API.
The numbers, with the usual caveat
Anthropic’s announcement reports Terminal-Bench 4.0 at 70.6% for Sonnet 5.5 against 10.3% for Sonnet 5, and CursorBench 4.0 at 55.5% against 34.1%. FrontierCode 1.1 moves less, from 42.4% to 46.2%. These are vendor figures, and a jump from 10.3% to 70.6% says more about how badly Sonnet 5 did on that harness than about a new ceiling.
The same post is blunt about the limit: Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment. Sonnet 5.5 is pitched at well-scoped everyday tasks such as fixing bugs. Treat it as the cheap, fast tier, not a replacement for the model you escalate to.
What breaks
The what’s new page lists five breaking changes. Four are likely to hit an ordinary coding agent.
Thinking can’t be disabled
Sending thinking: {"type": "disabled"} now returns a 400. The lowest setting is between_tools, which skips up-front thinking but keeps short progress notes between tool calls.
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
thinking={"type": "between_tools"},
output_config={"effort": "medium"},
tools=tools,
messages=messages,
)
There are catches. between_tools only works at low, medium or high effort, so a config that runs xhigh or max has to use adaptive thinking instead. It also takes no other thinking fields, and you can’t change effort per message while it’s set.
Forced tool use is gone
tool_choice of {"type": "any"} or {"type": "tool", "name": "..."} returns a 400. Only auto and none are accepted. If you forced a tool to get schema-valid output, keep auto, add strict: true to the tool definition, and say in the prompt when the tool applies.
tools = [{
"name": "report_findings",
"description": "Report review findings. Call this once you have finished reading the diff.",
"strict": True,
"input_schema": {
"type": "object",
"properties": {"findings": {"type": "array", "items": {"type": "string"}}},
"required": ["findings"],
"additionalProperties": False,
},
}]
History has to be append-only
Thinking blocks are now tied to the model and the conversation. If you edit the system prompt, the tools or an earlier message and then replay a Sonnet 5.5 thinking block, the request can return a 400. Per the docs, that check is on by default for accounts created on or after August 31, 2026.
This matters for agents that rewrite their own history, for example by trimming old tool results to save tokens. Strip the thinking blocks from the edited turn onward, or change instructions with mid-conversation system messages instead of edits.
Computer use needs the new toolset
On the Claude API and Google Cloud, computer_20251124 is rejected. Use computer_toolset_20260801. Bedrock still accepts the old tool.
The one that fails silently
Text between tool calls now arrives in thinking blocks when it’s longer than a sentence or two. At the default display: "omitted", those blocks are empty. If your UI streams the agent’s running commentary, it goes quiet between tool calls with no error.
With between_tools, the text comes back. With adaptive thinking, set thinking.display explicitly. Check this before you ship, because nothing will tell you.
Re-run your effort sweep
Anthropic says effort levels are recalibrated, so a setting doesn’t produce the same amount of thinking as on Sonnet 5. For agentic coding its guidance is to start at medium for well-specified tasks and move to high for harder or longer ones. The API default is high.
The practical step: run your existing eval set at medium and high before changing any production default, and compare cost per solved task, not just pass rate.