Vol. 1 · Edition 041Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

3-4×
efficiency claim
vs. comparable open-weight models
Open Source
By Sam Taylor with Samwise

On 501B-A23B architecture, first-party benchmarks that need independent verification, and whether the AlphaGo co-creator's efficiency bet can unseat the Chinese open-weight leaders.

Reflection's first model is competitive. It's fifth on Terminal-Bench. The compute-efficiency argument is the interesting part.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

02

Neutral

00

Pro (practical)

02

Pro (hyped)

00

← Anti-AI · Pro-AI →

Reflection AI released Beam on October 5. 501 billion parameters total, 23 billion active at inference, pretrained on 23.8 trillion tokens, with a 1-million token context window. Apache 2.0 open weights are promised for later this month. No exact date yet.

Two things set this apart from the usual model launch. First, who built it. Misha Laskin is CEO — he led reward modeling for DeepMind's Gemini project. His co-founder Ioannis Antonoglou co-created AlphaGo. They raised $2B at an $8B valuation in October 2025. Second, the pitch. Reflection isn't positioning Beam on raw benchmark leadership. Their stated claim is 3-4× less inference compute than comparable open-weight frontier models — meaning the same task output for a fraction of the hardware cost.

Per Reflection's own evaluation suite — all benchmarks below are first-party unless otherwise noted — Terminal Bench v2.1 sits at 80.1%. For reference: the September 2026 published leaderboard has Gemini 3.8 Flash at 90.8%, Fable 5.1 at 85.02%, Grok 4.5 at 83.3%, and DeepSeek V4 Flash at 82.7%. Beam's 80.1% puts it roughly fifth. Competitive — not leading.

23.8T
Training tokens — among the largest published figures for any open-weight model release

→ Source: Reflection AI via TechCrunch

Source spread

Pros & cons

What's real:

  • The founding team is not typical. Laskin and Antonoglou worked at the highest level of RL research for years. Their claimed fluency in efficient training runs isn't a marketing line — it comes from building Gemini and AlphaGo on constrained compute budgets.
  • 23.8 trillion training tokens is one of the larger published training-data figures for any open-weight release. Scale of training data matters; publishing the number is more transparent than most.
  • Apache 2.0 license commitment is significant if it holds. No commercial restrictions means this is deployable infrastructure from the moment the weights land.
  • A 1M context window at 501B total / 23B active is a real capability. The token density story holds for long-document workloads.
  • SWE-bench Verified at 80.9% and SWE Pro v2-Hard at 77.2% are strong if independently verified. Those are coding-agent numbers worth taking seriously.

What deserves a side-eye:

  • Every benchmark in the launch is first-party from Reflection's own evaluation suite. TechCrunch explicitly notes that Kimi K3, DeepSeek V4.1 Flash, and GLM 5.3 beat Beam on most competitive rows. That's not the framing Reflection leads with.
  • Terminal Bench v2.1 at 80.1% is fifth on the September 2026 published leaderboard. Competitive is the right word. "Rivals Chinese models" is doing extra work.
  • No API pricing announced. No confirmed open-weight date. For production planning purposes, this model does not yet exist in a usable form.
  • The 3-4× inference efficiency claim is the most interesting number in the release and the least sourced. It needs independent verification before it affects any infrastructure decision.
  • SWE-bench Verified scores across labs have been converging upward; 80.9% may be reproducible on benchmark tasks but gap against production codebases is worth testing directly.
Reflection AI: from AlphaGo to open weights
  1. March 2024

    Founded

    Misha Laskin (ex-DeepMind Gemini reward modeling lead) and Ioannis Antonoglou (AlphaGo co-creator) co-found Reflection AI

  2. October 2025

    $2B Series B

    Raised at $8B valuation; positioned as 'America's open frontier lab' for AI infrastructure and safety research

  3. October 5, 2026

    Beam preview

    501B-A23B model launched via API preview; all benchmarks first-party; Apache 2.0 weights pending

  4. Late October 2026

    Open weights

    Apache 2.0 weights promised; no confirmed date

Samwise's take

What builders need to know

For builders
  • Don't build on this yet. No open weights, no confirmed API pricing, no production SLA. File it under watch-this-space, not route-traffic-here.
  • Verify the efficiency claim when weights land. The 3-4× compute efficiency vs. Kimi K3 is the number that matters. If it holds on your workloads, the cost math changes substantially for self-hosted deployments.
  • Apache 2.0 is significant if it lands cleanly. No commercial restrictions means immediate enterprise deployability. Watch for the license text on the weight release; some labs attach use-case carve-outs to what they call Apache 2.0.
  • The 1M context window is relevant for long-document pipelines. If you need to ingest large documents or codebases in a single pass, 501B total / 23B active with 1M context is a different architecture than most API options.
  • SWE-bench Verified at 80.9% is worth testing independently. First-party benchmarks on SWE-bench have a history of running higher than production codebases. Run your own evals before routing agentic coding workloads here.
  • Subscribe for the open-weight date. Reflection said "late October." That's a 2-4 week window from today. Watch their blog or HuggingFace for the drop.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.