Vol. 1 · Edition 040Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

77.9%
DeepSWE v1.1 — Google's lead number
Terminal-Bench 4.0: 57.4%
Model Launch
By Sam Taylor with Samwise

On the DeepSWE lead versus the Terminal-Bench gap, what Google's 19-benchmark table actually shows, and who can use Argon right now.

Gemini 4 Argon's benchmark win is real. The coding narrative needs a footnote.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

If you use any of Google's AI tools — Gemini in Search, in Google Docs, in Workspace — something changed last Tuesday. Google announced Gemini 4 Argon on September 30, its first new frontier model in months. The announcement led with a headline number: 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.

The number is real. Two other coding benchmarks don't look as good for Google. That split is the story.

−9 pts
Argon's gap behind Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs 66.4%) — the benchmark where AI agents complete real shell tasks, not structured problem sets

→ Source: Google DeepMind benchmark table

Source spread

What's real

The knowledge-work benchmarks are genuinely strong. On legal automation (Harvey's Legal Agent Benchmark: Argon 19.6% versus next model's 6.7%) and finance (Vals Finance Agent v2: Argon 65.4%), Argon leads by a margin that isn't close. These benchmarks test what happens when an AI has to read large amounts of documents and make decisions across them — the kind of thing lawyers and analysts do all day. The 1M output token limit matters here. Think of the output limit like a whiteboard: a much bigger whiteboard means the AI can return more work in one go before you have to start clearing it. One million tokens is a real capability jump for long documents and multi-step tasks.

CWE-bench v1 at 68%, tied for first with GPT-6 Astra: the one security benchmark where Google isn't alone in first. The Fairwind-first rollout — to vetted cyber defenders before anyone else — matches this positioning deliberately.

Pricing is competitive at $2/$10 per million input/output tokens intro rate. The 95% discount on cached input makes repeated queries against a fixed large document very cheap.

What deserves a side-eye

FrontierSWE v2: Argon 55.0%, Claude Fable 5.1 at 56.3%, Claude Opus 5.5 at 62.3%, GPT-6 Astra at 65.5%. Argon finishes last on this one. FrontierSWE tests real software engineering — fixing bugs and implementing features in existing codebases, not structured problem sets. Ten points behind the leader.

Terminal-Bench 4.0: Argon 57.4%, Claude Opus 5.5 at 66.4%. Terminal-Bench drops an agent into a real shell and asks it to complete tasks — closer to how most developers actually use these models. Google's own comparison table includes this row. Argon loses it by nine points.

The gap between DeepSWE (Argon wins by 3 points) and Terminal-Bench (Argon loses by 9 points) is the kind of thing that should make you pause before accepting the "best at coding" framing. Both are coding benchmarks. They test different things: structured problem-solving versus sustained autonomous operation in a real environment.

Also: Google's 19-benchmark table has no third-party verification at launch. These numbers come from Google. Standard for a model announcement, and worth keeping in mind.

Gemini 4 Argon — the full coding picture
BenchmarkGemini 4 ArgonClaude Opus 5.5GPT-6 Astra
DeepSWE v1.177.9%74.2%74.1%
FrontierSWE v255.0%62.3%65.5%
Terminal-Bench 4.057.4%66.4%—
CWE-bench v168%67%68%

What to do about it

For everyday Gemini users: Nothing changes today. Argon launched to vetted cybersecurity programs only — not yet to Google Search, Docs, or Workspace. When it reaches consumer products, you'll most likely notice it on long, complex tasks: summarizing a big document, drafting a detailed report, multi-part research questions. Keep an eye on Google's product announcements for the broader rollout.

For builders evaluating models:

  • Don't switch based on DeepSWE alone. That benchmark and Terminal-Bench tell different stories about the same model.
  • If your workload is long-document processing, legal or finance automation, or large-context knowledge retrieval: Argon is worth testing at the intro price.
  • If your workload is terminal-heavy agent tasks: run your own eval before switching. The nine-point Terminal-Bench gap matters.
  • Wait for third-party benchmark reproduction before making production decisions.
  • Factor in the post-intro pricing ($4/$20/M) when doing production cost math.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.