Cloudflare Releases Clef, a 27B Decision Model, Under Apache 2.0
A.I. / news
Cloudflare Releases Clef, a 27B Decision Model, Under Apache 2.0
The model returns probabilities for typed questions instead of text, answers in a median 209.3 milliseconds, and trails the closed Jev model by 30.3 points on GPQA Diamond in Cloudflare's own table.

Cloudflare released Clef, a 27-billion-parameter model that answers typed questions about a piece of text, JSON or an image with probabilities instead of prose, on Oct. 1. The weights are on Hugging Face under the Apache 2.0 licence with no access form.
Michelle Chen of Cloudflare announced the model in a post on the Cloudflare blog, alongside a smaller sibling, Clef-flash, with 9 billion parameters. Both run on Workers AI, Cloudflare's GPU service. The post does not list a price.
What Clef returns instead of text
Clef belongs to a category that its competitors call System One models. A caller sends a state, such as a support ticket, and a schema of questions. Each question is a choice among named options, a score on an ordered scale, or a true-or-false probability. The model card says the model returns "a probability for every allowed option of every question in a single forward pass," with no text generation to parse.
TypeSafe AI released a closed-weights model called Jev on Sept. 15 and markets it as the first of the category, according to a Hyperstack tutorial. Clef's card says its API is "fully compatible with Jev and SystemOne," so code written against Jev can point at Clef.
Cloudflare's post gives the build as a Qwen 27B backbone with rank-256 low-rank adapters. The card names the base as Qwen/Qwen3.8-27B and adds a small transformer head that scores all options jointly. Clef-flash starts from Qwen3.5-9B.
Latency: 209.3 ms against Jev's 524.1 ms
Cloudflare reports a median latency of 209.3 milliseconds for Clef, 38.8 for Clef-flash and 524.1 for Jev. The 95th-percentile figures are 238.6, 122.4 and 536.0. The context window is 64,000 tokens against 32,000 for Jev.
- Clef-flash38.8 ms
- Clef209.3 ms
- Jev524.1 ms
Source: Cloudflare Clef model card on Hugging Face, accessed 2026-10-03
Cloudflare also said its threat-intelligence team classified websites in 2.2 seconds with Clef, against 4.7 seconds for GPT-OSS-120B. That figure is Cloudflare's own and no independent party is named.
The Decision Index, run by Cloudflare
Every accuracy number comes from what the card calls "our internal run" of the Decision Index 0.2.1 suite, which is hosted on a Cloudflare Workers domain. These are vendor-supplied results. The Terminal did not rerun them.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL, case exact accuracy | 98.5 | 98.8 | 95.8 |
| BANKING77, macro-F1 | 94.2 | 90.9 | 79.7 |
| CLINC150+OOS, macro-F1 | 97.4 | 66.8 | 89.3 |
| GPQA Diamond, accuracy | 48.0 | 51.0 | 78.3 |
| MMLU-Pro, accuracy | 65.9 | 65.3 | 82.7 |
| BBH, accuracy | 73.7 | 68.9 | 92.9 |
The split is plain in the card's own table. Clef beats Jev on intent classification and tool-calling tests. Jev beats it on GPQA Diamond by 30.3 points, on MMLU-Pro by 16.8 points and on BBH by 19.2 points. The blog's claim that Clef is "currently the leader" holds for the selection of tests it reproduces, not for all 41 rows on the card.
The card also reports four business workflows from Typesafe Evals, an evaluation site named like Jev's developer. The Terminal did not establish who runs it. On invoice processing, Clef scored 64.7 percent on exact actions against Jev's 61.8. On the agent-trace task Jev led, 71.6 percent to Clef's 68.5.
Weights, hardware and what is missing
The card lists testing on torch 2.11 and transformers 5.10.2 on a single H200 GPU. It does not give a minimum memory figure. The model reads images and video as well as text, which Cloudflare's post says Jev, a text-only model, cannot.
The Hugging Face repository showed 2,620 downloads and 861 likes on Oct. 3. Alibaba's Qwen3.8-27B, the base model, has more than 6.8 million downloads, as covered in The Terminal's report on the Qwen3.8 release.
Cloudflare said it will offer a reinforcement-learning platform for fine-tuning decision models. For now it is hands-on, run by forward-deployed engineers, and a self-serve version has no date. Cloudflare also warned that general performance can fall when a model is tuned for one domain, a trade the table above already shows. For more on how vendor-run leaderboards are read, see the Gemini 4 Argon benchmark report.
Sources
More in A.I.
- 01GPT-6.1 Sol Is Priced at One-Fifth of Astra, and Its System Card Rates Cyber CriticalOpenAI's Sept. 29 addendum also shows the model misrepresenting its own coding work more often than GPT-6 Astra did.
- 02Gemini 4 Argon Goes to Cyber Defenders First, With Broad Access UndatedGoogle priced its new frontier model at $2 and $10 per million tokens for an introductory period, then $4 and $20, and has not said when most developers get it.
- 03Aleph Alpha's Kolibri Ships Under Apache 2.0, Compared Only With Spring ModelsThe 78B-parameter German-English model activates 3.46B per token, and its published benchmark table leaves out every open-weight release since the spring.
- 04Runway's Praxis-1 Robot Model Is Open-Weight on Paper, With Weights Still UnreleasedThe video-trained control model is being tested by Noble Machines, Standard Bots and Ultra, and Runway has not published a parameter count or a licence.