Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on One
A.I. / news
Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on One
The Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.

Cloudflare released Clef and Clef-flash on Oct. 1, two open-weight models that answer fixed-choice questions with a probability for each answer, and it says the larger one leads TypeSafe's Jev on three of four of Jev's own workflow evaluations. Cloudflare's announcement, written by Michelle Chen, lists Clef as the leader on the Jev Decision Index and says both models are compatible with Jev's API.
The weights are on Hugging Face under Apache 2.0, according to the Clef model card. Both models are also hosted on Cloudflare Workers AI.

What a decision model does
A decision model does not write text. Given a state, which can be text, JSON, images or video, and a set of typed questions, it returns one probability per allowed answer in a single forward pass. The model card describes Clef as a 27-billion-parameter model post-trained from Qwen3.8-27B. Clef-flash is built on Qwen3.5-9B, according to the announcement.
The announcement says training used rank-256 low-rank adapters and a routing head on a frozen backbone. The data was internal synthetic sets that vary field order, prompts and schema structure. The card does not give a release date or a training-data breakdown.
Clef accepts images and video. Jev handles text only, the announcement says, and Clef's context window is 64,000 tokens against Jev's 32,000.
The numbers, and who ran them
Every figure below comes from Cloudflare's own run of the Decision Index. The Register reported on Oct. 1 that the scores had not been reproduced on the official index at publication.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL case exact | 98.47 | 98.76 | 95.75 |
| BANKING77 macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS macro-F1 | 97.43 | 66.77 | 89.27 |
| When2Call accuracy | 72.37 | 65.58 | 80.97 |
The last row is where Jev wins. Clef-flash also falls to 66.77 on CLINC150+OOS, 22.5 points behind Jev, so the cheaper model is not a safe swap for every intent-classification job.
- Jev524.1 ms
- Clef209.3 ms
- Clef-flash38.8 ms
Source: Cloudflare announcement, Decision Index run by Cloudflare (vendor-supplied), accessed 2026-10-09
The Register also reported that Clef is slower than other open decision models in Cloudflare's tests, though faster than Jev.
Price and hardware
The announcement states no price. The Register put Clef at $0.24 per million tokens, nearly six times Jev's $0.042.
Michelle Chen, a product manager in Cloudflare's AI Platform group, told The Register that Clef-flash needs a GPU with at least 41 GB of memory and Clef needs 85 GB. Both figures assume one request at a time and a 64,000-token context. The model card's own test used a single H200.
What "open source" covers
Cloudflare calls the release open source. Chen confirmed to The Register that the training datasets are not public, which makes this an open-weight release.
Jev's maker, TypeSafe, has kept its architecture secret. A separate Hacker News item this week reported that the company raised $870 million at a $7.5 billion valuation. The Terminal has not read that round's primary announcement and does not report its terms here.
The Qwen3.8 base is the same checkpoint covered in the Qwen3.8-27B report, and Clef follows that checkpoint's Apache 2.0 licence, the model card says. Open licences are worth checking on every repository: the Terminal found AWS's Physical AI Toolchain shipped with none.
Cloudflare said it is also starting an RL fine-tuning service, run first by its own engineering team and planned to become self-serve.
Sources
More in A.I.
- 01OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 02Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M EmbedderThe Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.
- 03Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One LineAlibaba's 27B dense model scores 61.7 on SWE-bench Pro by its own table. The card does not say what it was trained on or what hardware it needs.
- 04ChatGPT's Intelligent UI Skips the Pro Thinking Level and Older Desktop Apps, and Publishes No Usage FiguresOpenAI put GPT-6 into the Chat tab on October 7 with generated buttons, forms and charts. The announcement gives one speed number and no accuracy figure.