Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2B
A.I. / news
Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2B
AWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.

Amazon Web Services released Strands Decider 2B on Oct. 1, a roughly 2-billion-parameter model that picks between options and returns a confidence instead of writing text. The weights are free to download from Hugging Face under the Apache 2.0 licence.
The model comes from Strands Labs, AWS's experimental agent-development project, and is a direct answer to Jev, a decision model sold as an API by TypeSafe, according to VentureBeat and TechCrunch.

What Strands Decider 2B is and what it runs on
The model card lists the repository as StrandsAgents/strands-decider-2B-hobson-v19. The base is Qwen3.5-2B-Base from Alibaba, adapted with a LoRA adapter and a small readout head in place of the usual next-word output.
It answers three kinds of question: yes or no, multiple choice, and a rating on a scale. The card says each answer "carries a calibrated confidence" and lists model routing, tool selection, argument checking, triage and guardrails as intended uses.
Install is pip install strands-decider, with a command-line interface and an HTTP server. The Hugging Face API shows the repository was created on Sept. 30, 2026, is not gated and carries the Apache 2.0 licence tag.
The JevBench numbers and who ran them
The figures below come from the developers' model card and VentureBeat's report of them. They are vendor-supplied. No outside party is named as having reproduced them.
| Measure | Result | Conditions |
|---|---|---|
| JevBench accuracy | 0.723 (167 of 231 tasks) | Public benchmark, v19 |
| Brier score | 0.348 | Lower is better |
| Expected calibration error | 0.050 | Model card |
| Median latency | 106 ms | RTX 3090, includes HTTP round trip |
| 95th percentile latency | 296 ms | RTX 3090 |
| Median latency, laptop | about 150 ms | M3 MacBook |
Internal evaluations on the card range from 0.641 to 0.884 accuracy depending on the dataset. The card warns that questions are read less well than documents, that multi-step reasoning is weaker, and that calibration was fitted only on short classification tasks.
- Lowest internal set0.64 accuracy
- JevBench0.72 accuracy
- Highest internal set0.88 accuracy
Source: Strands Decider 2B model card on Hugging Face, accessed 2026-10-01
How it compares with TypeSafe's Jev
VentureBeat reports that Jev costs $0.042 per million input tokens, with latencies of 70 to 500 ms through TypeSafe's hosted service. Jev's weights are not released.
Strands Decider 2B ships with its training data, code and scripts, VentureBeat reported. The same article notes that AWS has not shown that self-hosting saves money once hardware and operations are counted.
TechCrunch reported that the project began when Marc Brooker, a distinguished engineer at AWS, saw Jev and built his own version. It briefly topped the JevBench ranking for its size, and Amazon engineers then cleaned it up for release.
Brooker told TechCrunch the appeal is a "perfect decider for a workflow step" with lower latency and potentially lower cost. TypeSafe CEO Diogo Almeida told the outlet that the current batch "seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful."
TechCrunch also reported that OpenAI announced a similar offering in the same week, and that dozens of decision models now exist.
What is unconfirmed
No independent run of JevBench v19 on this checkpoint has been published. None of the coverage read for this report describes what the training data contains, only that it is released.
The release builds on the same Qwen family that PrismML shrank in its Ternary Bonsai 2 27B, and it targets the agent plumbing that tools like NVIDIA's OpenShell sandbox. The Hugging Face repository showed 6 likes at the time of writing.
Sources
More in A.I.
- 01Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 02OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 03Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use ItGoogle priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.
- 04OpenAI and Synopsys Sign Chip-Design Model Deal With No Customers or Benchmarks NamedGPT-Synopsys will run Synopsys EDA tools on OpenAI-hosted infrastructure under a revenue-sharing agreement. The Sept. 30 announcement gives no dollar figure and no ship date.