Vol. 1 · Edition 039Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

Opus 5.5 input/M

$4

GPT-6 Sol input/M

$2
Model Launch
By Sam Taylor with Samwise

On Terminal-Bench 4.0 vs. AutomationBench, the benchmark comparison problem, and whether Luna at $0.10/M is the actual story.

Opus 5.5 and GPT-6 Sol landed the same Tuesday. The choice is cleaner than it looks.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

03

Pro (hyped)

00

← Anti-AI · Pro-AI →

Anthropic and OpenAI shipped new models on the same afternoon this week. That's unusual. Both are meaningfully cheaper than what they replace. That's genuinely useful.

The natural question is: which one should you use? The honest answer is that the benchmarks are harder to compare than the pricing is, and the pricing gap is real. Let me work through both.

What launched

On September 22, 2026, Anthropic shipped Claude Opus 5.5. Same day, OpenAI shipped GPT-6 Sol and GPT-6 Luna. Both announcements within a few hours of each other.

The pricing picture first, because that's the part you can compare directly:

September 22 launches vs. the models they displace
ModelInput /MOutput /MContext windowKey change
Claude Opus 5.5$4$20200K40% cheaper than Opus 5, 30% faster output
Claude Opus 5 (prior)$5$25200KDisplaced by Opus 5.5
Claude Fable 5.1higherhigher200KFlagship; Opus 5.5 exceeds it on coding
GPT-6 Sol$2$101.05M (922K in)50% cheaper than GPT-5.6 Sol promo
GPT-6 Luna$0.10$0.501.05M (922K in)50% cheaper than GPT-5.6 Luna promo
$0.10
GPT-6 Luna input per million tokens — the cheapest frontier-tier API in the market as of September 22

→ Source: OpenAI

Source spread

Pros & cons

What's genuinely good here:

  • Opus 5.5 is the first model to exceed Claude Fable 5.1 on Terminal-Bench 4.0 (66.4% vs. 55.8%) and FrontierCode v1.1 (54.4% vs. 50.3%) — at 40 percent lower cost. That's a non-trivial leap.
  • The Fast mode option ($8/$40 per million) gives you up to 2.5x output speed when latency matters more than cost.
  • GPT-6 Luna at $0.10/$0.50 is, for the right workloads, absurdly cheap. If you're running high-volume extraction, classification, or structured output tasks where you're not bottlenecked by reasoning quality, Luna is worth testing immediately.
  • GPT-6 Sol's 1.05M token context (with 922K of usable input) is the longest context in a non-flagship model. If you're working with large document sets or long-running agent sessions, that window matters.
  • The pricing is confirmed permanent by OpenAI — not a promotional rate. That's relevant for budget planning.

What deserves a side-eye:

  • The benchmark comparisons are genuinely messy. Anthropic reported Terminal-Bench 4.0 and FrontierCode v1.1 for Opus 5.5 but OpenAI didn't report Terminal-Bench at all for Sol or Luna. OpenAI reported AutomationBench and DeepSWE v1.1. These benchmarks don't map cleanly onto each other. You cannot look at Opus 5.5's Terminal-Bench 4.0 score and GPT-6 Sol's AutomationBench score and draw a winner — they're measuring different things.
  • Anthropic's safety evaluations before launch (METR, Frontier Design, behavioral audits) are more thorough documentation than OpenAI provided for Sol/Luna. If your use case touches sensitive domains, that paper trail is relevant.
  • Luna's raw cheapness may encourage overuse in places where output quality actually matters. $0.10/M is a trap if you end up generating low-quality output at scale and paying to fix it.

Samwise's take

What builders need to know

For builders
  • Default Opus 5.5 for agentic workflows. Terminal-Bench 4.0 at 66.4% is the strongest coding-agent score in the market from a non-proprietary-closed model. If you're running Claude Code, complex multi-tool agents, or structured reasoning tasks, the quality gain is worth the $4/M vs $2/M premium over Sol.
  • Test Luna now on your cheapest tier. $0.10/M input at frontier-class output quality — that's a different cost regime than anything from three months ago. Batch extraction, structured output, classification, summarization at scale: run Luna and compare against what you were using.
  • Use GPT-6 Sol for long-context document work. The 922K usable input window is the longest in a mid-tier model. If you're building over large codebases, long meeting transcripts, or extended agent histories, Sol's context window changes what's possible.
  • Benchmarks don't cross-compare. Don't try to read Opus 5.5's Terminal-Bench number against Sol's AutomationBench number. They're different tasks. Run your own evals on tasks from your actual production distribution.
  • The Fast mode option on Opus 5.5 is underrated. $8/$40 per million for up to 2.5x output speed. If you're building anything user-facing where the model needs to appear responsive, the latency improvement at that price point is worth evaluating.
  • Migration note. If you're on Claude Opus 5, Opus 5.5 is a drop-in upgrade. Same context window (200K). API identifier: claude-opus-5-5. No breaking changes documented.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.