LIVE
All stories ›
AI IN LIFENEWS
Tools & AppsBusiness & DealsAI ModelsSocietyResearchChips & ComputeSafety & SecurityRegulation & PolicyRoboticsReviews OpenAIAnthropicGoogle & DeepMindMetaAlibaba / QwenxAIByteDance
Home › AI Models › MODELS
MODELS

Microsoft Decision-1 routing model hits 83.5% at 85 ms

Post-trained from Qwen3.5-9B, the router runs on Foundry and OpenRouter at $0.042 per million input tokens, with output tokens free.

Microsoft Decision-1 routing model hits 83.5% at 85 ms
Symbolic image: a hand guides a patch cable into a port on the routing rack as two status LEDs flare up beside it.

In short

Microsoft's Decision-1 is a compact model that scores fixed options instead of writing prose, reaching 83.5 percent accuracy at 85 ms latency across the company's own 36-benchmark suite.

At a glance

  • Post-trained from Qwen3.5-9B for decision scoring; returns calibrated probabilities over fixed options as structured output.
  • 83.5 percent accuracy across 36 benchmarks and close to 150,000 questions — all figures from Microsoft's own testing.
  • 85 ms latency, which Microsoft calls 2.5x faster than runner-up H2O-Lightning-4B.
  • Price: $0.042 per million input tokens, output tokens free; available via Microsoft Foundry and OpenRouter.
  • Cloudflare's open Clef models are absent from the comparison, though they are also Qwen-based.

Microsoft's Decision-1 does not write prose. It picks among fixed options and returns calibrated probabilities as structured output, which is why the company reports its quality as one accuracy number: 83.5 percent, at 85 ms of latency, measured in house.

A scorer, not a chatbot

The model targets routing, classification, prioritization, verification and workflow control inside agent systems. It is post-trained from Qwen3.5-9B, Alibaba's open base model, and the selling point is stability: small variations in the input should not flip the decision. The Decoder notes that models of this class can also steer agents through complex environments.

Thirty-six benchmarks, one vendor

The 83.5 percent figure spans 36 benchmarks and close to 150,000 questions. Microsoft puts latency at 85 ms and calls that 2.5x faster than runner-up H2O-Lightning-4B, while its product listing claims 35x faster median latency than GPT-6 Sol. Every one of those numbers comes from Microsoft's own tests, not from an independent evaluation.

Pricing built for very short answers

Input costs $0.042 per million tokens and output tokens are free — a sensible structure for a model whose answer is a label and a probability. Distribution runs through Microsoft Foundry and OpenRouter. Microsoft also points to internal use: Xbox Research reportedly works 14x faster and 200x cheaper than with GPT-6 Sol, and Copilot teams have deployed it too.

A category born in mid-September

Jev opened this market in mid-September 2026, and OpenAI (with a Decisions API) and Cloudflare (with the open Clef models) followed quickly. Microsoft's own comparison places Decision-1 ahead of Jev 1.13.0 on both accuracy and speed. That comparison leaves out Cloudflare's Clef models, which are open source and built on Qwen as well.

What we could not verify

One of our three source URLs, MarkTechPost's launch write-up, answered our request with HTTP 403, so this article rests on The Decoder and Microsoft's own product listing. We have no independent measurement of accuracy, latency or calibration, and commenters on the launch thread raise calibration drift in production as an open concern. Whether Decision-1 will ship open weights is not stated in the sources we were able to read.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is Microsoft Decision-1?

A compact decision model post-trained from Qwen3.5-9B. It chooses among fixed options and returns calibrated probabilities for routing, classification, prioritization and verification rather than generating long text.

How much does Microsoft Decision-1 cost?

$0.042 per million input tokens, with output tokens billed at nothing. The model is available through Microsoft Foundry and OpenRouter.

How reliable is the 83.5 percent accuracy claim?

The 83.5 percent across 36 benchmarks and close to 150,000 questions is Microsoft's own measurement. No independent replication is available yet, and Cloudflare's Clef models were excluded from the comparison.

Sources

More reports