
Every few months a new AI model shows up with "the best ever" on the label, and most of the time the real-world change is small. Claude Sonnet 5.5, which Anthropic released on September 28, 2026, looks like one of the exceptions.
The headline: it's faster than Sonnet 5, it costs less per task, and on several benchmarks it scores the same as or better than Claude Opus 5.5, Anthropic's flagship model, at half the token price.
If you build with the Claude API, run coding agents, or just want to know which model to pick, here's what actually changed.
claude-sonnet-5-5For a long time the choice between models was simple: pay more for the smartest one, or pay less and accept worse results.
Sonnet 5.5 makes that choice harder. Anthropic describes it as "a faster, lower-cost complement to Claude Opus 5.5," and the numbers back that up. For many everyday workloads, including coding agents, support automation, and computer use, the gap between Sonnet and Opus is now small enough that the price difference probably matters more.
The per-task savings are worth a closer look. The token price didn't drop. What changed is that the model uses fewer tokens and fewer tool calls to finish the same job. In production, that's the number that ends up on your bill.
Here are the numbers Anthropic published, compared with Sonnet 5 and Opus 5.5:
| Benchmark | What it tests | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Agentic coding | 70.6% | 10.3% | 66.4% |
| FrontierCode 1.1 | Hard coding problems | 46.2% | 42.4% | 54.4% |
| CursorBench 4.0 | Coding in the IDE | 55.5% | 34.1% | 57.8% |
| OSWorld 2.1 | Computer use | 80.1% | 57.0% | 81.8% |
| Humanity's Last Exam | Expert-level reasoning | 64.5% | 54.9% | 67.7% |
| GDPval-AA v2.1 | Real-world knowledge work (Elo) | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1 | Business tasks (Elo) | 1811 | 1359 | 1822 |
| Chartography | Reading charts (no tools) | 61.6% | 15.6% | 64.4% |
1. It beats Opus 5.5 on agentic coding. The Terminal-Bench 4.0 score of 70.6% is the stat that will get the most attention, since it puts a mid-tier model ahead of the flagship on multi-step work in the terminal.
2. It's within a point or two of Opus almost everywhere else. On GDPval-AA the gap is just 2 Elo points (1844 vs. 1846). On OSWorld it's 1.7 percentage points. For most teams that difference won't show up in practice.
3. Opus still leads on the hardest problems. FrontierCode is where the gap is widest: 54.4% for Opus against 46.2% for Sonnet 5.5. If your work involves genuinely difficult, novel engineering problems, Opus is still the safer choice.
4. Vision took a big step forward. Chart reading went from 15.6% to 61.6%. If you've ever watched a model misread a bar chart, you'll understand why this matters.
A note on the numbers: the Sonnet 5 Terminal-Bench figure (10.3%) is what Anthropic's page reports. It's an unusually big jump, so for your own use case, run your own evals before you trust any single benchmark.
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| Input (per 1M tokens) | $2 | $4 |
| Output (per 1M tokens) | $10 | $20 |
| Cache writes (per 1M tokens) | $2.50 | $5 |
| Cache reads (per 1M tokens) | $0.20 | $0.20 |
Sonnet 5.5 costs half as much as Opus 5.5 for input, output, and cache writes. Add the per-task efficiency gains and the savings for high-volume workloads can be large.
Teams with strict data requirements can also get zero data retention.
The benchmarks look good, but results from real companies are more convincing. A few examples Anthropic shared:
A pattern shows up across these: fewer steps, fewer tokens, fewer retries. That's where the "up to 30% cheaper per task" claim comes from.
Sonnet 5.5 holds up better on long-running tasks and handles tool calls more efficiently, including batching them. For agent builders, that means fewer wasted round-trips.
You can set how hard the model thinks: Low, Medium, High, Xhigh, or Max. Use Low for quick classification and Max for tough debugging, and only pay for as much reasoning as the task needs.
Structured outputs are supported, which makes it much easier to get reliable JSON into your application.
Anthropic says the model is better at design and polishing UI, which will matter to anyone using it to build front-end components.
Sonnet 5.5 is "the first Sonnet model to beat Pokémon Red working only from screenshots." It sounds like a gimmick, but finishing a long game from pixels alone takes a lot of planning, memory, and visual understanding over many hours.
Anthropic shipped several safeguards with this release:
For most teams, switching is a one-line change to the model ID:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Refactor this function for readability..."}],
)
print(response.content[0].text)
One gotcha: if you had thinking disabled on Sonnet 5, Anthropic says you'll need to switch the thinking setting to between_tools when you migrate. Check the official migration guide in the Claude Platform docs before you push to production.
Switch now if you:
Stick with Opus 5.5 if you:
The practical approach: make Sonnet 5.5 your default and send only the hardest tasks to Opus. Many teams will see better results at a lower cost that way.
Model releases usually come with a trade-off: faster but less accurate, or smarter but pricier. Claude Sonnet 5.5 is unusual because it gets faster, cheaper per task, and more capable at the same time.
It doesn't replace Opus 5.5, and Anthropic doesn't claim it does. But for most day-to-day work, it's now hard to justify paying twice as much.
Try it on your own workload, measure tokens per task, and compare. The numbers may change which model you use by default.
When was Claude Sonnet 5.5 released? Anthropic released Claude Sonnet 5.5 on September 28, 2026.
How much does Claude Sonnet 5.5 cost? $2 per million input tokens and $10 per million output tokens, half the price of Claude Opus 5.5.
Is Claude Sonnet 5.5 better than Opus 5.5? On agentic coding (Terminal-Bench 4.0) it scores higher: 70.6% vs. 66.4%. On most other benchmarks Opus 5.5 leads, but usually by only a small margin.
How much faster is Claude Sonnet 5.5? It generates output more than 30% faster than Sonnet 5.
Where can I use Claude Sonnet 5.5? In the Claude apps, on the Claude Platform (API), and through Amazon Web Services, Google Cloud, and Microsoft Azure.
What is the model ID for Claude Sonnet 5.5?
claude-sonnet-5-5
Source: Anthropic – Claude Sonnet 5.5