On the DeepSWE lead versus the Terminal-Bench gap, what Google's 19-benchmark table actually shows, and who can use Argon right now.
Gemini 4 Argon's benchmark win is real. The coding narrative needs a footnote.
Anti-AI
00
Skeptic
01
Neutral
00
Pro (practical)
02
Pro (hyped)
01
← Anti-AI · Pro-AI →
If you use any of Google's AI tools — Gemini in Search, in Google Docs, in Workspace — something changed last Tuesday. Google announced Gemini 4 Argon on September 30, its first new frontier model in months. The announcement led with a headline number: 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.
The number is real. Two other coding benchmarks don't look as good for Google. That split is the story.
Source spread
- Google DeepMind — Gemini 4 Argon — hype. Official company announcement. 19-benchmark comparison table, pricing, access tiers. All numbers first-party; no independent methodology links cited.
- TechCrunch — Google releases Gemini 4 Argon — builder. Straightforward launch coverage; notes the Fairwind-first access restriction and the government voluntary review process.
- benchlm.ai — Gemini 4 Argon benchmark analysis — skeptic. Aggregates the full benchmark picture, including Terminal-Bench and FrontierSWE v2, which Google did not lead with.
What's real
The knowledge-work benchmarks are genuinely strong. On legal automation (Harvey's Legal Agent Benchmark: Argon 19.6% versus next model's 6.7%) and finance (Vals Finance Agent v2: Argon 65.4%), Argon leads by a margin that isn't close. These benchmarks test what happens when an AI has to read large amounts of documents and make decisions across them — the kind of thing lawyers and analysts do all day. The 1M output token limit matters here. Think of the output limit like a whiteboard: a much bigger whiteboard means the AI can return more work in one go before you have to start clearing it. One million tokens is a real capability jump for long documents and multi-step tasks.
CWE-bench v1 at 68%, tied for first with GPT-6 Astra: the one security benchmark where Google isn't alone in first. The Fairwind-first rollout — to vetted cyber defenders before anyone else — matches this positioning deliberately.
Pricing is competitive at $2/$10 per million input/output tokens intro rate. The 95% discount on cached input makes repeated queries against a fixed large document very cheap.
What deserves a side-eye
FrontierSWE v2: Argon 55.0%, Claude Fable 5.1 at 56.3%, Claude Opus 5.5 at 62.3%, GPT-6 Astra at 65.5%. Argon finishes last on this one. FrontierSWE tests real software engineering — fixing bugs and implementing features in existing codebases, not structured problem sets. Ten points behind the leader.
Terminal-Bench 4.0: Argon 57.4%, Claude Opus 5.5 at 66.4%. Terminal-Bench drops an agent into a real shell and asks it to complete tasks — closer to how most developers actually use these models. Google's own comparison table includes this row. Argon loses it by nine points.
The gap between DeepSWE (Argon wins by 3 points) and Terminal-Bench (Argon loses by 9 points) is the kind of thing that should make you pause before accepting the "best at coding" framing. Both are coding benchmarks. They test different things: structured problem-solving versus sustained autonomous operation in a real environment.
Also: Google's 19-benchmark table has no third-party verification at launch. These numbers come from Google. Standard for a model announcement, and worth keeping in mind.
| Benchmark | Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.2% | 74.1% |
| FrontierSWE v2 | 55.0% | 62.3% | 65.5% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | — |
| CWE-bench v1 | 68% | 67% | 68% |
What to do about it
For everyday Gemini users: Nothing changes today. Argon launched to vetted cybersecurity programs only — not yet to Google Search, Docs, or Workspace. When it reaches consumer products, you'll most likely notice it on long, complex tasks: summarizing a big document, drafting a detailed report, multi-part research questions. Keep an eye on Google's product announcements for the broader rollout.
For builders evaluating models:
- Don't switch based on DeepSWE alone. That benchmark and Terminal-Bench tell different stories about the same model.
- If your workload is long-document processing, legal or finance automation, or large-context knowledge retrieval: Argon is worth testing at the intro price.
- If your workload is terminal-heavy agent tasks: run your own eval before switching. The nine-point Terminal-Bench gap matters.
- Wait for third-party benchmark reproduction before making production decisions.
- Factor in the post-intro pricing ($4/$20/M) when doing production cost math.
Further reading
- Google DeepMind — Gemini 4 Argon — all first-party benchmarks, pricing, and access details
- TechCrunch — Google releases Gemini 4 Argon — launch news with access and government review context
- benchlm.ai — Gemini 4 Argon — aggregated benchmark comparisons including Terminal-Bench and FrontierSWE v2
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.