Cloudflare Open-Sources Clef, a 27B Drop-In for TypeSafe's Jev
A.I. / news
Cloudflare Open-Sources Clef, a 27B Drop-In for TypeSafe's Jev
The Apache 2.0 decision models cut median latency from 524 ms to 39 ms on Cloudflare's own tests, and lose to Jev by 30 points on GPQA Diamond.

Cloudflare released two open-weight "decision models" on October 1, a 27-billion-parameter Clef and a 9-billion-parameter Clef-flash, both under Apache 2.0 and both built to be drop-in replacements for the Jev API that TypeSafe AI sells. Cloudflare's own latency table puts Clef-flash at a 38.8 millisecond median against 524.1 ms for Jev.
Every benchmark in the announcement comes from Cloudflare's run of TypeSafe's own Jev Decision Index, version 0.2.1, so the scores are vendor-supplied. No outside party has published a rerun in the sources reviewed.
What a decision model returns
A decision model does not write text. Per the Hugging Face model card, Clef is "a 27B multimodal model that turns a state and a schema of typed questions into decisions", and it returns a probability for each allowed option in one forward pass. A support-triage call might ask three typed questions and get back a probability for every answer to each.
Clef is post-trained from Qwen3.8-27B. Clef-flash sits on Qwen3.5-9B, according to Cloudflare's blog post. The card does not describe the post-training data. Cloudflare lists a 64k-token context window and says the API is "fully Jev-API compatible".
InfoQ reported that Clef takes images and video, while Jev handles only text classification today. It also put Jev's context window at 32k.
Where Clef wins and where it does not
On classification and tool-call tests Clef beats Jev in Cloudflare's run. On reasoning-heavy tests it loses by a wide margin.
| Benchmark (Cloudflare run) | Clef | Jev |
|---|---|---|
| BFCL case-exact accuracy | 98.5 | 95.8 |
| BANKING77 macro-F1 | 94.2 | 79.7 |
| CLINC150+OOS macro-F1 | 97.4 | 89.3 |
| When2Call accuracy | 72.4 | 81.0 |
| GPQA Diamond | 48.0 | 78.3 |
The model card carries the GPQA Diamond pair itself. Clef at 48.0 against Jev at 78.3 is a 30-point gap on graduate-level science questions, so a buyer wanting a model that also reasons should read the full table before switching.

Latency and price
Cloudflare's blog lists median and 95th-percentile decision times for each model.
- Clef-flash (9B)38.8 ms
- Clef (27B)209.3 ms
- Jev524.1 ms
Source: Cloudflare blog, Clef decision models, accessed 2026-10-10
The pricing is not on the blog page. Developers Digest lists Workers AI input rates of $0.24 per million tokens for Clef and $0.09 for Clef-flash, against $0.042 for Jev as checked on September 16. That makes Clef about 5.7 times Jev's input rate, so the trade is latency and image input against price.
Running it yourself
The card says it was tested on a single H200 with torch 2.11 and transformers 5.10.2, and links SGLang launch commands for H200, B200 and B300 hardware. Weights download without a gate.
Microsoft's Decision-1 is also built on a Qwen base. Google's EmbeddingGemma 2 shipped under the same Apache 2.0 licence. On Hacker News, as InfoQ quoted them, Jacek Złydach argued "It's not a 'new paradigm'" and that TypeSafe was first to market the approach.
Cloudflare said fine-tuning will start through its own forward-deployed engineers, with a self-service platform planned and no date given.
Sources
More in A.I.
- 01Nvidia Is Reportedly Weighing a Takeover of Reflection AI, Whose 501B Beam Model Has No Public Weights YetThe Financial Times reported early-stage talks on October 10. Reflection's own benchmark table has Beam behind Kimi K3 and Qwen 3.8 Max on every coding test where all three report.
- 02Qwen3.8-27B Is Apache 2.0. The 2.4T Max Weights Have No Published Licence TexteWeek reports a $50 million revenue trigger on the Max licence; RuntimeWire calls the terms unresolved, and the Hugging Face card gives only a name.
- 03OpenAI Posts Hundreds of AI-Written Maths Papers, Keeps the PromptsThe repository holds 719 manuscripts by README count, a 42% Lean share by OpenAI's measure and 22% by Decrypt's, and no model name.
- 04Microsoft-Decision-1 Is a Post-Trained Qwen3.5-9B at $0.042 per Million Tokens, With No Licence GivenMicrosoft's launch post claims a win across 36 benchmarks and 35 times the speed of GPT-6 Sol, and every figure in it is the company's own.