Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

2.8T
parameters
Moonshot AI · Kimi K3 · July 2026
Open Source
By Sam Taylor with Samwise

On 2.8T MoE with Kimi Delta Attention, the coding-benchmark wins over GPT-5.6 Sol and Fable 5, the compute-disclosure gap, and what Moonshot's opacity means for your trust model.

Kimi K3 tops the coding leaderboard in three categories. The training story is the harder question.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

Moonshot AI released Kimi K3 on July 16. 2.8 trillion total parameters, mixture-of-experts with a proprietary attention architecture called Kimi Delta Attention, 1 million token context window. API access is live. Weights drop July 27 on HuggingFace under a Modified MIT license — which makes this the largest open-weight model ever announced, once it actually ships.

The coding benchmarks are worth taking seriously. On Program Bench, K3 scores 77.8 — edging GPT-5.6 Sol at 77.6 and Fable 5 at 76.8. On SWE Marathon, a benchmark designed around long-horizon coding sessions rather than short tasks, K3 scores 42.0 against Opus 4.8's 40.0, GPT-5.6 Sol's 39.0, and Fable 5's 35.0. These are independent benchmarks, not first-party claims.

The part without an answer: Moonshot hasn't disclosed training compute, hardware configuration, or structured safety evaluations. For a 2.8T model, those gaps are not minor.

Source spread

Pros & cons

What's real:

  • Program Bench 77.8 leads the field. Artificial Analysis' coding task suite is not a first-party benchmark. K3 topping it against Sol and Fable 5 is meaningful — particularly because the margin is razor-thin, which suggests these models have genuinely converged at the frontier on this task class.
  • SWE Marathon 42.0 leads the field. This one matters specifically for agentic coding workflows where you can't just retry every five minutes. If your use case looks like long-horizon autonomous coding sessions, K3's lead here is the number to replicate.
  • 2.5× scaling efficiency vs K2. Moonshot claims Kimi Delta Attention and Attention Residuals convert compute into capability more efficiently than K2's architecture — roughly 2.5× by their measure. If that holds, this is an actual architectural advance, not just "more compute." The technical paper isn't public yet.
  • 1M context at frontier MoE efficiency. 2.8T total parameters, 16 of 896 experts active per token. That keeps inference tractable at extreme scale.

What deserves a side-eye:

  • DeepSWE lags at 67.5. GPT-5.6 Sol scores 73.0; Fable 5 scores 70.0. DeepSWE is a broad real-world software engineering task set — arguably the one that looks most like daily engineering work. K3's coding advantage is real on some tasks and absent on others. Don't read "leads on Program Bench" as "leads on everything."
  • Training compute: undisclosed. Moonshot has not published FLOP count, hardware configuration, or energy consumption for the K3 training run. US export controls have blocked advanced Nvidia GPUs from China since 2022, tightening further in 2024 and 2025. The obvious question — how does a Beijing startup train a 2.8T frontier model — has no public answer. Either they solved a meaningful training-efficiency problem (in which case we'd have a paper), or there's hardware access that isn't being disclosed, or both are partly true.
  • Safety evals: not published. No structured red-team reports, no evaluation against known harm categories. That's the norm for Chinese labs right now. It's still a real risk for builders deploying into regulated or sensitive contexts.
  • "Open weight" means July 27, not today. The API is live. The weights aren't. $3/$15 per million while you wait is not cheap for a model you can't yet self-host.
  • Modified MIT: read the actual terms. The K2 family had commercial scale restrictions. Don't assume K3's "Modified MIT" is identical to MIT until you've read what they modified. The K3 license terms should be in the HuggingFace model card by July 27.
Kimi K3 vs frontier models — coding benchmarks (Artificial Analysis, July 2026)
BenchmarkKimi K3GPT-5.6 SolClaude Fable 5
DeepSWE67.573.0 ▲70.0
FrontierSWE81.271.386.6 ▲
Program Bench77.8 ▲77.676.8
SWE Marathon42.0 ▲39.035.0
Terminal-Bench 2.188.3

What builders need to know

For builders
  • DeepSWE is the watch benchmark. K3 trails GPT-5.6 Sol by 5.5 points there. If your workload looks more like broad real-world SWE tasks than long-horizon sessions, the leaderboard position is misleading.
  • API is live at platform.kimi.ai; weights drop July 27. Run your eval suite on the API now — you have 9 days of data before you need to decide whether local deployment is worth the wait.
  • Read the Modified MIT license terms before going to production. The K2 commercial scale restrictions surprised some teams. K3 terms should appear in the HuggingFace model card at weights launch.
  • Safety evals are absent. No published red-team results. Treat this as any undisclosed-safety model: your application layer carries the risk, full stop.
  • The architecture paper isn't public. KDA and AttnRes are Moonshot's internal names. The 2.5× efficiency claim is unverified by external researchers. Watch for the technical report — if the architecture advance is real, it has implications beyond K3.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.