Vol. 1 · Edition 041Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

70.6%
Terminal-Bench 4.0
Sonnet 5.5 overtook Opus
Model Launch
By Sam Taylor with Samwise

On the Terminal-Bench 4.0 inversion, the token efficiency gains that explain it, and when Opus 5.5 is still the right choice.

Claude's middle model just beat the expensive one at coding. Use it accordingly.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

00

← Anti-AI · Pro-AI →

If you've asked Claude to help fix a spreadsheet, write a difficult email, or work through a report draft, the version you use next week will probably finish faster — and need fewer follow-up corrections. That's not marketing language. There's a structural reason, and it's worth understanding.

On September 28, Anthropic released Claude Sonnet 5.5. Same price as Sonnet 5: $2 per million input tokens, $10 per million output. The price is the same. The model is not.

70.6%
Sonnet 5.5's Terminal-Bench 4.0 score — beating Opus 5.5's 66.4% at the same $2/$10 per million token price

→ Source: Anthropic

What happened

Terminal-Bench 4.0 measures how well an AI completes real software engineering work inside a terminal — writing code, running commands, fixing bugs across a full project across many steps. It's a harder, newer version of the benchmarks used when Sonnet 5 launched. It rewards models that can sustain effort over dozens of steps without losing the thread.

Sonnet 5.5 scored 70.6% on it. Sonnet 5 scored 10.3%. Opus 5.5 — Anthropic's most expensive, most capable model — scored 66.4%.

That 60-point jump between Sonnet 5 and Sonnet 5.5 is real, but I want to be precise about what it means. Terminal-Bench 4.0 is a newer test that wasn't the primary target when Sonnet 5 was trained. Sonnet 5.5 was. So the gap isn't "Sonnet 5.5 is objectively 6× better at coding." It's "Sonnet 5.5 was built to be excellent at this specific kind of multi-step work, and it is."

The Opus comparison is the one worth sitting with. Opus 5.5 costs more per token and sits at 66.4% on the same test. Sonnet 5.5 at 70.6% beats it.

Source spread

What's real

  • 30%+ faster output than Sonnet 5. For common tasks — writing, summarizing, drafting — that's a real quality-of-life change.
  • The token efficiency gain may matter more than the speed number. Sonnet 5.5 uses roughly 30% fewer tokens to complete the same work. Because you pay per token, "same rate card" becomes "meaningfully cheaper in practice."
  • Fewer tokens per step, compounded over a 50-step agent loop, explains most of the Terminal-Bench jump. It's not magic; it's math.
  • First Sonnet model to ship with frontier-style cyber safeguards — previously reserved for Opus. Enterprises deploying it in security-sensitive contexts now have a faster, cheaper option.
  • Available on claude-sonnet-5-5 across AWS, Google Cloud, and Azure.

What deserves a side-eye

  • The Sonnet 5 → Sonnet 5.5 gap (10.3% to 70.6%) is partly a benchmark version story. Terminal-Bench 4.0 is harder than what came before, and Sonnet 5.5 was trained for it. This is not the same as Sonnet 5.5 being 6× better at all coding.
  • Anthropic has not published SWE-bench Verified numbers for Sonnet 5.5. That's the canonical, independently maintained software engineering benchmark. The absence is notable. It may just mean Anthropic chose not to run it; it may mean the result was unflattering.
  • "First Sonnet to beat Pokémon Red from screenshots alone" appears in the announcement materials. I understand why it's there — it's a proxy for visual reasoning across many sequential steps. But it tells you almost nothing about professional use cases. Every time a benchmark mentions Pokémon, something is being obscured.
Sonnet 5 vs Sonnet 5.5 vs Opus 5.5
Sonnet 5Sonnet 5.5Opus 5.5
Terminal-Bench 4.010.3%70.6%66.4%
CursorBench 4.0—55.5%—
GDPval-AA v2.1 (Elo)1,4491,844—
Price (input / output, per M)$2 / $10$2 / $10higher
Speed vs Sonnet 5baseline+30%+slower
Frontier cyber safeguardsNoYesYes
Best forShort, clear tasksMulti-step coding, documentsOpen-ended hard problems

What to do about it

For everyday Claude users (Claude.ai, claude.com): The default model is already being updated. You may not need to do anything. The result should be faster replies and fewer correction rounds on standard tasks — writing, summarizing, analyzing documents.

For developers using the Claude API:

  • Test with claude-sonnet-5-5 against your eval suite before flipping production traffic. "Drop-in replacement" doesn't mean your existing prompts won't need tuning — Anthropic flagged that instruction-following behavior changed.
  • Measure actual token usage on your workloads rather than assuming the 30% efficiency gain applies to you. It's workload-dependent.
  • If you were running Opus 5.5 specifically for the frontier cyber safeguards: Sonnet 5.5 now offers the same protections at lower cost. It's worth evaluating whether you can move workloads down.

For anyone who just wants to know if it matters: If you pay for Claude and have noticed that complex requests sometimes need multiple follow-ups to land right — this update directly addresses that. Faster output and higher per-task success rates on the kinds of work most people actually do.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.