Microsoft-Decision-1 Is a Post-Trained Qwen3.5-9B at $0.042 per Million Tokens, With No Licence Given
A.I. / news
Microsoft-Decision-1 Is a Post-Trained Qwen3.5-9B at $0.042 per Million Tokens, With No Licence Given
Microsoft's launch post claims a win across 36 benchmarks and 35 times the speed of GPT-6 Sol, and every figure in it is the company's own.

Microsoft released Microsoft-Decision-1 on Oct. 9, a model that scores a fixed set of options with a probability for each, and priced it at $0.042 per million input tokens with output tokens free. The model is post-trained from Qwen3.5-9B, and Microsoft's launch post does not state a licence or a parameter count beyond that base name.
The post, by Achint Srivastava, vice president of software engineering in Microsoft's Office of the CTO, says the model is available on Microsoft Foundry and OpenRouter. It was published on the company's Command Line site.
What the model does
Decision-1 is built for single-pass decision scoring: routing, classification, prioritization, verification and workflow control. A request supplies a fixed set of options, in yes-or-no, multiple-choice or rating form, through a structured API call. The model returns a calibrated probability for each option.
Microsoft said it can also grade AI responses and agent actions against a rubric. The post says a 90 percent prediction should be right about nine times in ten on representative cases, which is a calibration claim, not a measured result in the post.
The company said it will rebase the model on other models, including Microsoft AI (MAI) and OpenAI models.
The benchmark claims are all Microsoft's
Microsoft said Decision-1 had the highest accuracy in a 36-benchmark comparison covering nearly 150,000 questions kept blind from training. It said the model was 35 times faster than GPT-6 Sol at median latency and 2.5 times faster than H2O-Lightning-4B v1.1, the fastest runner-up.
These are vendor-supplied figures. The post also says the model was tested against top models on the JevBench leaderboard across 36 additional public and private benchmarks.

JevBench is run by Benchmark Heaven, which says it is maintained independently of TypeSafe AI, the company behind the Jev model. Its board ranks open-weight models on a capability score that averages intelligence and calibration. The page The Terminal read on Oct. 10, part of it truncated, showed no Microsoft-Decision-1 row, so the model's rank there is unconfirmed.
- Quyet-1.0-Large81.7 points
- decisio v0.8.0 on gemma-4-31B-it79.6 points
- deck-31B77.6 points
- René-1 31B FP876 points
- H2O-Lightning-4B v1.175 points
Source: Benchmark Heaven JevBench leaderboard, accessed 2026-10-10
The board puts TypeSafe's own Jev 1.13.0 at 77.1, as an unranked API reference. Cloudflare's Clef was compared with Jev on four tests in our earlier coverage, so decision models now have a small public scoreboard to be measured against.
Internal results and what is missing
Microsoft listed four internal uses. All were reported by the company, and none has an outside check.
| Team | Task | Microsoft's reported result |
|---|---|---|
| Xbox Research | Labelled over 10,000 feedback items | Competitive with GPT-6 Sol, 14 times faster, 200 times cheaper |
| Copilot team | Quality control on chat and agent replies | Competitive with GPT5.6 Luna, 100 times faster |
| Microsoft Discovery | Adaptive replanning | 46 times more consistent than an LLM-based score |
| Incident response | Retrieval across logs, tickets, messages | Better and faster than an LLM |
The post gives a robustness figure: across eight input perturbations, the decision changed in 1.3 percent of cases on average, with no flips when option descriptions were paraphrased or options were reversed or shuffled. It also says the model was tested on 5,250 requests across 11 safety benchmarks.
The licence question
The base model, Qwen3.5-9B, carries an Apache-2.0 licence on its Hugging Face card, with a native context of 262,144 tokens. Microsoft's post does not say whether Decision-1's weights can be downloaded, what licence applies, or what context window it supports. Qwen's later release, Qwen3.8-27B, shipped under Apache 2.0 with a model card that gave training data one line.
Microsoft said it will update the model with new evaluations and data. It gave no date for the MAI or OpenAI rebases.
Sources
More in A.I.
- 01Anthropic's Claude Haiku 4.5 Filed a Fake Murder Tip in July, and Philadelphia Police Heard in OctoberAnthropic's own report says the model was building practice tasks on random web pages when it submitted the tip, 72 days before the company found it.
- 02OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 03Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M EmbedderThe Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.
- 04Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on OneThe Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.