Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

Luna at launch · Jul 9

$1.00

Luna now · Jul 30

$0.20
Industry
By Sam Taylor with Samwise

On Luna's drop from $1 to $0.20, Terra's 20% reduction, Sol staying untouched at $5, and what it means for the make-vs-buy calculation when the cheapest US frontier API matches Chinese model pricing.

OpenAI cut Luna 80% three weeks after launch. Read the Sol price to understand why.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

03

Pro (hyped)

00

← Anti-AI · Pro-AI →

OpenAI cut GPT-5.6 Luna 80% on July 30. The model went from $1 per million input tokens to $0.20, and output moved from $6 to $1.20. Terra, the mid-tier model in the GPT-5.6 family, dropped 20%: $2.50 to $2 input, $15 to $12 output. Sol, the flagship, didn't move. Still $5 input, $30 output.

The GPT-5.6 family launched July 9. Three weeks. 80%.

That speed deserves attention. OpenAI's explanation: efficiency gains they discovered while building GPT-5.6 — the model helped rewrite and optimize its own production inference code, which reduced serving costs, which they're passing on. That's a real explanation if it's accurate. It's also the kind of clean narrative that's hard to independently verify.

The alternative explanation is simpler: they priced Luna at $1, and the market didn't bite. DeepSeek V4 Pro sits at $0.44 per million input tokens on OpenRouter. DeepSeek V4 Flash is $0.09. Kimi K3 is available as open weights for self-hosting. At $1 input, Luna had no real value proposition over those alternatives except provenance and infrastructure reliability. At $0.20, the calculation looks different for a lot of workloads.

The two explanations aren't mutually exclusive. But the timing — 80% in three weeks — suggests the primary driver was demand signal, not a technical efficiency breakthrough. Genuine 80% efficiency gains don't emerge three weeks post-launch.

The Sol number

Sol didn't move. That's the structural tell.

Sol is $5 input / $30 output. It's the model that cleared the US government's pre-release frontier-model review before launch, that regulated enterprises are writing procurement contracts around, that defense and intelligence buyers use for sensitive workflows. There is no Chinese open-weight equivalent a defense contractor can use for those use cases. No DeepSeek alternative an enterprise legal team will approve for contract analysis. No open-source model that clears the data-residency requirements in regulated industries.

OpenAI has pricing power on Sol because the competitive set is genuinely restricted — by regulation, by buyer risk tolerance, by procurement rules that specifically exclude foreign-origin models. Sol doesn't need to respond to DeepSeek V4 Flash pricing. It never did.

Luna is a different product entirely. Luna competes with every cheap API. And the market for cheap APIs has gotten very cheap.

What $0.20 input actually means

At $0.20 per million input tokens, you're approaching parity with self-hosting open-weight models on cloud GPU infrastructure — for moderate utilization workloads without significant batch efficiency. Not all use cases, not all models. But builders who've been running smaller Llama or Qwen variants on spot instances primarily as a cost play should run the math again.

GPT-5.6 family pricing: before and after July 30
ModelInput — beforeInput — afterOutput — beforeOutput — after
GPT-5.6 Sol$5.00/M$5.00/M ←$30.00/M$30.00/M ←
GPT-5.6 Terra$2.50/M$2.00/M (−20%)$15.00/M$12.00/M (−20%)
GPT-5.6 Luna$1.00/M$0.20/M (−80%)$6.00/M$1.20/M (−80%)

The honest nuance: self-hosting costs depend heavily on workload shape, utilization rate, and ops overhead. High-batch, high-utilization inference on owned or leased hardware is still cheaper per token than any frontier API. Low-utilization bursty workloads frequently aren't. The $0.20 number shifts the crossover point — it doesn't eliminate self-hosting as a strategy.

Worth noting: Luna at $0.20 is still roughly 2× DeepSeek V4 Flash at $0.09. For workloads without data-residency or provenance requirements, the cheapest option isn't Luna.

Source spread

Pros & cons

What's real:

  • $0.20 input for a US frontier lab is genuinely significant. For the first time, a model with OpenAI's infrastructure guarantees, latency profile, and provenance documentation costs what some open-source alternatives cost to run.
  • Whether the cut came from efficiency gains or competitive pressure, the result is the same: builders pay less.
  • Sol not moving is actually the coherent part of this. The three-tier structure now makes sense: Luna competes on cost, Terra is the workhorse, Sol competes on access and provenance.
  • For builders building on Luna for high-volume, cost-sensitive workloads, 80% is a real budget change.

What deserves a side-eye:

  • 80% in three weeks implies the original $1 was either a test price or wrong. Both interpretations reduce confidence in OpenAI's pricing as a signal of value.
  • Luna at $0.20 still costs 2× DeepSeek V4 Flash at $0.09. Pure-cost optimization without data-residency constraints still points elsewhere.
  • Terra's 20% cut looks almost perfunctory next to Luna's 80%. If Terra is competing with Claude Sonnet 5 at $2/$10, a 20% reduction might not shift workloads.
  • The efficiency narrative is unfalsifiable in the short term. OpenAI won't publish inference cost structure; there's no way to verify whether efficiency gains explain 80% of the cut.

What builders need to know

  • Reprice Luna in your cost models immediately. If you have any active usage of GPT-5.6 Luna, your cost per token dropped 80% on July 30. Update your billing projections.
  • If cost was why you weren't using Luna, re-evaluate. At $0.20 input, Luna is now in the range where self-hosting open alternatives needs justification on non-cost grounds — latency, fine-tuning, data residency, output control.
  • Don't conflate Luna and Sol. If your reason for using GPT-5.6 is government clearance, regulated-industry procurement, or documentation requirements, you're in Sol — and Sol didn't change. The Luna cut doesn't apply to those workloads.
  • The cheapest US frontier API is no longer the cheapest option. DeepSeek V4 Flash at $0.09 input is still 2× cheaper than Luna. If you can use Flash for a workload (no data-residency or provenance requirement), the cost math still points there.
  • Terra at $2/$12 deserves a recheck. The 20% cut is modest, but if you've been in Terra for workload-size reasons, your cost has moved. Run the new numbers against Claude Sonnet 5 and Gemini 3.6 Flash before assuming where you sit competitively.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.