Anthropic shipped Claude Sonnet 5.5 on 2026-09-28, the second model in the Claude 5.5 family. The headline number is not the price: it is 70.6% on Terminal-Bench 4.0 against Sonnet 5's 10.3%. Same sticker price as Sonnet 5, and Anthropic says it typically costs up to 30% less per task because it needs fewer tokens to finish the same work.
That combination — a mid-tier model within a few points of the previous frontier on agentic coding — is the thing to plan around. I would re-run the task-level routing in any agent pipeline built in the last few months.
The Benchmark Table
All figures from Anthropic's launch post. Dashes mean the vendor did not report publicly.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 (Main) | 46.2% Max | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | — |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Terminal-Bench and GPT-6 Sol were not reported together, so the OpenAI column above mixes GPT-6 Sol (FrontierCode, GDPval-AA, AA-Briefcase) with GPT-5.6 Sol (Terminal-Bench, CursorBench). Anthropic is explicit about that in its footnotes.
Read the table by distance, not by rank. GDPval-AA is 1844 versus Opus 5.5's 1846. AA-Briefcase is 1811 versus 1822. On the knowledge-work and computer-use evals Sonnet 5.5 is a rounding error behind the flagship.
Pricing and Effort Levels
| Price per 1M tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input tokens | $2 | $4 |
| Output tokens | $10 | $20 |
| Cache writes | $2.50 | $5 |
| Cache reads | $0.20 | $0.20 |
Half the flagship on input, output and cache writes, and the same on cache reads. Sonnet 5.5 also generates output 30%+ faster than Sonnet 5, which Anthropic calls its fastest Sonnet to date.
The more interesting lever is the effort setting, which Anthropic now documents as a first-class cost control. In Claude Code and the Claude apps the default is Medium; on the Claude Platform it defaults to High. On Terminal-Bench 4.0, Sonnet 5.5 at Medium effort "far exceeds Sonnet 5's best score for less than a tenth of the cost per task." On CursorBench, Sonnet 5.5 at Low effort beats Sonnet 5's best score for less than a tenth of the cost. On AA-Briefcase, Medium effort bests Sonnet 5's best for about one ninth of the cost.
My read: if you were paying for Opus-tier routing to get acceptable agent-completion rates, the burden of proof has moved. Start Sonnet 5.5 at Low or Medium and escalate per task, not per pipeline.
The Max-Effort Trap on FrontierCode
One detail worth internalizing before you set effort to Max everywhere. Sonnet 5.5 scores lower at Max effort than at Xhigh on FrontierCode: 46.2% at Max against 52.1% at Xhigh. Anthropic's explanation is that FrontierCode rewards mergeable changes and penalizes out-of-scope edits — and at Max effort the model more often invoked Claude Code's code-review skill, which splits review across many subagents, leading in two cases Cognition examined to a timeout or to edits beyond the task's scope.
So more reasoning budget bought a worse score on a benchmark that measures whether your change would survive review. That is a real tradeoff, not a footnote: cap effort at Xhigh for code-review-adjacent work.
What Changes in Your Code
The model ID is claude-sonnet-5-5. It is available on all platforms — AWS, Google Cloud, Microsoft Azure — with zero data retention, as with Opus 5.5 and Sonnet 5.
One migration item will bite people. If you run Sonnet with thinking turned off, you must switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Anthropic's migration guide has the details.
Two operational changes land alongside it. Sonnet 5.5 is the first Sonnet to launch with cybersecurity safeguards and fallbacks like Opus 5.5's, so higher-risk security tasks will visibly fall back to Sonnet 5 (routine bug-finding in your own code is unaffected). And it expands preserved thinking, so thinking can no longer be decoupled from the account that created it — Anthropic notes this matters if you move conversations between accounts or switch accounts mid-session in Claude Code.
What I Would Do This Week
Pull your agent logs and sort tasks by escalation rate. Anything you currently pin to Opus because smaller models stalled is a candidate. The customer quotes in the post point the same direction: CodeRabbit plans to move simple and moderate reviews over now, Slack reports roughly 14% fewer output tokens on unchanged prompts, Balyasny reports ~121k tokens per answer against Sonnet 5's ~497k across 2,441 finance tasks, and Atlassian says Rovo Agents run up to 30% faster.
Claude Haiku 5.5 joins the family in the coming weeks, which will push the cheap end of that ladder down again.
The catch is the usual one: these are vendor-run evaluations, and Opus 5.5 still wins on complex, open-ended work requiring sustained judgment — 54.4% on FrontierCode against 46.2%, and 64.4% on Chartography against 61.6%. The gap is now narrow enough that the default should be the cheaper model, with escalation as the exception rather than the rule.
Sources: