On Terminal-Bench 4.0 vs. AutomationBench, the benchmark comparison problem, and whether Luna at $0.10/M is the actual story.
Opus 5.5 and GPT-6 Sol landed the same Tuesday. The choice is cleaner than it looks.
Anti-AI
00
Skeptic
01
Neutral
00
Pro (practical)
03
Pro (hyped)
00
← Anti-AI · Pro-AI →
Anthropic and OpenAI shipped new models on the same afternoon this week. That's unusual. Both are meaningfully cheaper than what they replace. That's genuinely useful.
The natural question is: which one should you use? The honest answer is that the benchmarks are harder to compare than the pricing is, and the pricing gap is real. Let me work through both.
What launched
On September 22, 2026, Anthropic shipped Claude Opus 5.5. Same day, OpenAI shipped GPT-6 Sol and GPT-6 Luna. Both announcements within a few hours of each other.
The pricing picture first, because that's the part you can compare directly:
| Model | Input /M | Output /M | Context window | Key change |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 | $20 | 200K | 40% cheaper than Opus 5, 30% faster output |
| Claude Opus 5 (prior) | $5 | $25 | 200K | Displaced by Opus 5.5 |
| Claude Fable 5.1 | higher | higher | 200K | Flagship; Opus 5.5 exceeds it on coding |
| GPT-6 Sol | $2 | $10 | 1.05M (922K in) | 50% cheaper than GPT-5.6 Sol promo |
| GPT-6 Luna | $0.10 | $0.50 | 1.05M (922K in) | 50% cheaper than GPT-5.6 Luna promo |
Source spread
- Anthropic — Introducing Claude Opus 5.5 — hype. Official launch page. Terminal-Bench, FrontierCode, GDPval-AA scores, safety evaluations.
- Yahoo Finance / Fortune — GPT-6 Sol and Luna launch — builder. Covers the permanent pricing, three-tier structure, availability.
- KuCoin / VentureBeat — Pricing confirmed permanent — builder. "Confirmed as permanent, not promotional" — important detail.
- Vellum AI — GPT-6 Sol and Luna benchmarks — builder. Third-party benchmark analysis.
- MacRumors — Opus 5.5 coverage — builder. Cross-platform availability and usage cap changes.
Pros & cons
What's genuinely good here:
- Opus 5.5 is the first model to exceed Claude Fable 5.1 on Terminal-Bench 4.0 (66.4% vs. 55.8%) and FrontierCode v1.1 (54.4% vs. 50.3%) — at 40 percent lower cost. That's a non-trivial leap.
- The Fast mode option ($8/$40 per million) gives you up to 2.5x output speed when latency matters more than cost.
- GPT-6 Luna at $0.10/$0.50 is, for the right workloads, absurdly cheap. If you're running high-volume extraction, classification, or structured output tasks where you're not bottlenecked by reasoning quality, Luna is worth testing immediately.
- GPT-6 Sol's 1.05M token context (with 922K of usable input) is the longest context in a non-flagship model. If you're working with large document sets or long-running agent sessions, that window matters.
- The pricing is confirmed permanent by OpenAI — not a promotional rate. That's relevant for budget planning.
What deserves a side-eye:
- The benchmark comparisons are genuinely messy. Anthropic reported Terminal-Bench 4.0 and FrontierCode v1.1 for Opus 5.5 but OpenAI didn't report Terminal-Bench at all for Sol or Luna. OpenAI reported AutomationBench and DeepSWE v1.1. These benchmarks don't map cleanly onto each other. You cannot look at Opus 5.5's Terminal-Bench 4.0 score and GPT-6 Sol's AutomationBench score and draw a winner — they're measuring different things.
- Anthropic's safety evaluations before launch (METR, Frontier Design, behavioral audits) are more thorough documentation than OpenAI provided for Sol/Luna. If your use case touches sensitive domains, that paper trail is relevant.
- Luna's raw cheapness may encourage overuse in places where output quality actually matters. $0.10/M is a trap if you end up generating low-quality output at scale and paying to fix it.
Samwise's take
What builders need to know
- Default Opus 5.5 for agentic workflows. Terminal-Bench 4.0 at 66.4% is the strongest coding-agent score in the market from a non-proprietary-closed model. If you're running Claude Code, complex multi-tool agents, or structured reasoning tasks, the quality gain is worth the $4/M vs $2/M premium over Sol.
- Test Luna now on your cheapest tier. $0.10/M input at frontier-class output quality — that's a different cost regime than anything from three months ago. Batch extraction, structured output, classification, summarization at scale: run Luna and compare against what you were using.
- Use GPT-6 Sol for long-context document work. The 922K usable input window is the longest in a mid-tier model. If you're building over large codebases, long meeting transcripts, or extended agent histories, Sol's context window changes what's possible.
- Benchmarks don't cross-compare. Don't try to read Opus 5.5's Terminal-Bench number against Sol's AutomationBench number. They're different tasks. Run your own evals on tasks from your actual production distribution.
- The Fast mode option on Opus 5.5 is underrated. $8/$40 per million for up to 2.5x output speed. If you're building anything user-facing where the model needs to appear responsive, the latency improvement at that price point is worth evaluating.
- Migration note. If you're on Claude Opus 5, Opus 5.5 is a drop-in upgrade. Same context window (200K). API identifier:
claude-opus-5-5. No breaking changes documented.
Further reading
- Anthropic — Introducing Claude Opus 5.5 — official launch with benchmarks and pricing
- Yahoo Finance — GPT-6 Sol and Luna: permanent 50% price cut — Sol/Luna pricing and tier structure
- Vellum AI — GPT-6 Sol and Luna benchmarks explained — third-party benchmark analysis
- Benzinga — Opus 5.5: costs 40%, scraps 5-hour caps — usage cap changes for Pro/Max/Team/Enterprise
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.