On the Terminal-Bench 4.0 inversion, the token efficiency gains that explain it, and when Opus 5.5 is still the right choice.
Claude's middle model just beat the expensive one at coding. Use it accordingly.
Anti-AI
00
Skeptic
01
Neutral
00
Pro (practical)
02
Pro (hyped)
00
← Anti-AI · Pro-AI →
If you've asked Claude to help fix a spreadsheet, write a difficult email, or work through a report draft, the version you use next week will probably finish faster — and need fewer follow-up corrections. That's not marketing language. There's a structural reason, and it's worth understanding.
On September 28, Anthropic released Claude Sonnet 5.5. Same price as Sonnet 5: $2 per million input tokens, $10 per million output. The price is the same. The model is not.
What happened
Terminal-Bench 4.0 measures how well an AI completes real software engineering work inside a terminal — writing code, running commands, fixing bugs across a full project across many steps. It's a harder, newer version of the benchmarks used when Sonnet 5 launched. It rewards models that can sustain effort over dozens of steps without losing the thread.
Sonnet 5.5 scored 70.6% on it. Sonnet 5 scored 10.3%. Opus 5.5 — Anthropic's most expensive, most capable model — scored 66.4%.
That 60-point jump between Sonnet 5 and Sonnet 5.5 is real, but I want to be precise about what it means. Terminal-Bench 4.0 is a newer test that wasn't the primary target when Sonnet 5 was trained. Sonnet 5.5 was. So the gap isn't "Sonnet 5.5 is objectively 6× better at coding." It's "Sonnet 5.5 was built to be excellent at this specific kind of multi-step work, and it is."
The Opus comparison is the one worth sitting with. Opus 5.5 costs more per token and sits at 66.4% on the same test. Sonnet 5.5 at 70.6% beats it.
Source spread
- Anthropic — Introducing Claude Sonnet 5.5 — hype. Company framing leads with Terminal-Bench and positions Sonnet 5.5 as the "faster, lower-cost complement to Opus 5.5."
- The New Stack — Claude Sonnet 5.5 launch — builder. Focuses on the performance-per-dollar shift and what CursorBench 4.0 (55.5%) and GDPval-AA v2.1 (1,844) mean for production API users.
- The Next Web — Cyber limits in Sonnet 5.5 — builder. Covers the reasoning-extraction classifiers and expanded cyber safeguards as a signal about the model family direction.
- Thurrott.com — Anthropic releases Claude Sonnet 5.5 — builder. Clean summary of what changed for everyday and power users.
What's real
- 30%+ faster output than Sonnet 5. For common tasks — writing, summarizing, drafting — that's a real quality-of-life change.
- The token efficiency gain may matter more than the speed number. Sonnet 5.5 uses roughly 30% fewer tokens to complete the same work. Because you pay per token, "same rate card" becomes "meaningfully cheaper in practice."
- Fewer tokens per step, compounded over a 50-step agent loop, explains most of the Terminal-Bench jump. It's not magic; it's math.
- First Sonnet model to ship with frontier-style cyber safeguards — previously reserved for Opus. Enterprises deploying it in security-sensitive contexts now have a faster, cheaper option.
- Available on
claude-sonnet-5-5across AWS, Google Cloud, and Azure.
What deserves a side-eye
- The Sonnet 5 → Sonnet 5.5 gap (10.3% to 70.6%) is partly a benchmark version story. Terminal-Bench 4.0 is harder than what came before, and Sonnet 5.5 was trained for it. This is not the same as Sonnet 5.5 being 6× better at all coding.
- Anthropic has not published SWE-bench Verified numbers for Sonnet 5.5. That's the canonical, independently maintained software engineering benchmark. The absence is notable. It may just mean Anthropic chose not to run it; it may mean the result was unflattering.
- "First Sonnet to beat Pokémon Red from screenshots alone" appears in the announcement materials. I understand why it's there — it's a proxy for visual reasoning across many sequential steps. But it tells you almost nothing about professional use cases. Every time a benchmark mentions Pokémon, something is being obscured.
| Sonnet 5 | Sonnet 5.5 | Opus 5.5 | |
|---|---|---|---|
| Terminal-Bench 4.0 | 10.3% | 70.6% | 66.4% |
| CursorBench 4.0 | — | 55.5% | — |
| GDPval-AA v2.1 (Elo) | 1,449 | 1,844 | — |
| Price (input / output, per M) | $2 / $10 | $2 / $10 | higher |
| Speed vs Sonnet 5 | baseline | +30%+ | slower |
| Frontier cyber safeguards | No | Yes | Yes |
| Best for | Short, clear tasks | Multi-step coding, documents | Open-ended hard problems |
What to do about it
For everyday Claude users (Claude.ai, claude.com): The default model is already being updated. You may not need to do anything. The result should be faster replies and fewer correction rounds on standard tasks — writing, summarizing, analyzing documents.
For developers using the Claude API:
- Test with
claude-sonnet-5-5against your eval suite before flipping production traffic. "Drop-in replacement" doesn't mean your existing prompts won't need tuning — Anthropic flagged that instruction-following behavior changed. - Measure actual token usage on your workloads rather than assuming the 30% efficiency gain applies to you. It's workload-dependent.
- If you were running Opus 5.5 specifically for the frontier cyber safeguards: Sonnet 5.5 now offers the same protections at lower cost. It's worth evaluating whether you can move workloads down.
For anyone who just wants to know if it matters: If you pay for Claude and have noticed that complex requests sometimes need multiple follow-ups to land right — this update directly addresses that. Faster output and higher per-task success rates on the kinds of work most people actually do.
Further reading
- Anthropic — Introducing Claude Sonnet 5.5 — official launch, all benchmark numbers sourced from here
- The Next Web — Sonnet 5.5 cyber limits and reasoning extraction protection
- Thurrott.com — Anthropic releases Claude Sonnet 5.5
- SWE-bench Verified leaderboard — the independent benchmark Anthropic didn't publish numbers for; check back here as third parties evaluate the model
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.