Microsoft built a model that doesn't talk. That's the point.
Jump to a section
See the story at a glance
Satya Nadella posted on October 9 about a new Microsoft model. The marketing hook — "fast decision-making AI" — sounds like every other AI launch. The actual product is different enough that I read it twice to make sure I wasn't misunderstanding.
Microsoft Decision-1 doesn't generate text. It takes a task description, a fixed list of options, and optional grading criteria, then returns a calibrated probability score for each option in one pass. No generation. No reasoning chain. Just: here are the choices, here's the best one, here's how confident I am.
This is the kind of model that slides under most builder radar because it doesn't do the thing that gets coverage.
What it actually is
Decision-1 is post-trained from Alibaba's Qwen3.5-9B, an open-weight 9-billion-parameter model. Microsoft didn't build this from scratch — they post-trained an existing open-weight base and optimized it for classification. That's either an embarrassing omission from the launch materials or a sensible engineering decision they're underselling. I think it's the latter.
The use cases Nadella called out: routing, classification, prioritization, verification, and workflow control. Microsoft's own Xbox Research team reportedly used it to categorize over 10,000 player reviews at GPT-6 Sol quality at 14× the speed. That sentence should land with anyone running any automated classification pipeline at scale.
Pricing: $0.042 per million input tokens, output free. The output-free part is load-bearing. Classification models don't generate many tokens — they return probabilities — so Microsoft can price by input only. For comparison, the cheapest tier of GPT-6 Luna is $0.20/M input and $0.80/M output. At any meaningful classification volume, the cost difference is not small.
| Decision-1 | GPT-6 Luna | |
|---|---|---|
| Input cost per M tokens | $0.042 | $0.20 |
| Output cost per M tokens | Free | $0.80 |
| P50 latency vs GPT-6 Sol | 35× faster | ~1× (baseline) |
| Avg benchmark accuracy (company-reported) | 83.5% across 36 benchmarks | Not published for routing tasks |
| Available now | Microsoft Foundry | OpenAI API, AWS Bedrock |
Source spread
- Satya Nadella on X, Oct 9 — hype. Launch post. Covers routing, classification, prioritization, verification, workflow control. References Xbox Research and incident response as internal testers.
- Windows Report — hype. Covers the 35× speed headline without engaging with methodology.
- Windows Forum megathread — skeptic. Flags that all benchmark figures come from Microsoft's own evaluation, not third-party reproduction. Explicitly notes the Qwen3.5-9B base model wasn't prominently disclosed in launch materials.
- FourWeekMBA — builder. Clean breakdown of pricing and Foundry availability.
Pros & cons
What's real:
- The speed claim is structurally plausible. Decision-1 doesn't generate tokens — it scores a fixed option list in one pass. GPT-6 Sol has to generate a response. These are different compute profiles, and a large latency gap between them isn't surprising.
- The pricing is genuinely cheap for what it does. $0.042/M with free output changes the cost math for classification-heavy workflows.
- Availability in Microsoft Foundry on day one. This isn't a preview or a waitlist. You can test it now.
What deserves a side-eye:
- The benchmarks are company-reported. 83.5% average accuracy across 36 suites with 147,137 questions sounds rigorous. It might be. But these are benchmarks Microsoft selected, run by Microsoft. No independent reproduction yet.
- The 35× latency comparison is against GPT-6 Sol specifically — the heaviest model in OpenAI's lineup. A classification model should be faster than the biggest frontier LLM. The comparison flatters the headline.
- The Qwen3.5-9B foundation. Building on Alibaba's open-weight model while separately benefiting from US export controls on Chinese AI hardware is a tension worth naming. It's not wrong — open weights are open weights — but it's there.
Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality.
What builders need to know
- Decision-1 is live in Microsoft Foundry today. OpenRouter access is coming, no date confirmed.
- Input format: task description + fixed option list + optional grading criteria. Output: probability scores per option. Not a chat model — don't use it like one.
- At $0.042/M input with free output, benchmark it against any classification pipeline currently running on an LLM. The cost delta is real at scale.
- All accuracy figures are Microsoft-reported. Run your own eval on a held-out dataset before trusting the 83.5% headline.
- Base model is Qwen3.5-9B (Alibaba). If compliance requirements flag Alibaba-derived model weights, check before deploying.
- The 35× speed claim is against GPT-6 Sol specifically. Verify latency on your actual workload before treating this as a universal truth.
Further reading
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.
Keep following the thread
A little more to explore.
Pipecat produced replies. Our local audio setup still took too long.
Nine offline audio runs completed. The eight warm runs reached first generated audio at a median of 2.9 seconds, before playback or telephony.
Our website-only AI agent accepted a fake cancellation policy
Sixty answers from one AnythingLLM configuration revealed repeatable failures, including a user message that overrode the real cancellation policy.
The browser automation worked. The token savings are still unproven.
A signed-in extension found an existing HeyGen output. That proves a useful task worked; it does not establish cheaper browser automation.