Cloudflare's Clef Is a 27B Apache 2.0 Model That Never Writes Text, and Its Own Table Has the Baseline Winning Two of Six Rows
A.I. / news
Cloudflare's Clef Is a 27B Apache 2.0 Model That Never Writes Text, and Its Own Table Has the Baseline Winning Two of Six Rows
The decision model scores typed options in one forward pass, 209 ms at the median, but every number in the launch post is Cloudflare's own.

Cloudflare released Clef, a 27-billion-parameter model that picks among typed options and never writes text, on Oct. 1 under the Apache 2.0 licence. It also released a 9B sibling, Clef-flash, and announced a reinforcement-learning fine-tuning service.
The company set out the models in a post by Michelle Chen, whose job title the post does not give. Every benchmark and latency figure in it was run by Cloudflare, and the post cites no outside evaluation.
What a decision model does
Clef is post-trained from Qwen3.8-27B, according to its Hugging Face card. Clef-flash is built on Qwen3.5-9B. Given a state and a schema of typed questions, each returns a probability for every allowed answer in one prefill pass. The card says there is no free-form generation and no output parsing.
Cloudflare wrote that "a decision model makes classifications to help agents decide how to act." Its example is the company's own threat-intelligence team, which used Clef to classify domains in 2.2 seconds a site. The general model gpt-oss-120b took 4.7 seconds and returned only two classifications.
Both models run on Workers AI. The post gives no price for them.
Latency and the benchmark rows
- Jev (baseline)524.1 ms
- Clef209.3 ms
- Clef-flash38.8 ms
- Laya5.8 ms
Source: Cloudflare, Clef launch post, 2026-10-01, accessed 2026-10-08. Vendor-run.
Cloudflare said Laya is faster than both Clef models but scores lower on quality. At the 95th percentile the order tightens: Clef-flash 122.4 ms, Laya 222.5 ms, Clef 238.6 ms and Jev 536.0 ms.
| Benchmark (Cloudflare-run) | Clef | Clef-flash | Jev |
|---|---|---|---|
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS, macro-F1 | 97.43 | 66.77 | 89.27 |
| BFCL, exact case | 98.47 | 98.76 | 95.75 |
| When2Call, accuracy | 72.37 | 65.58 | 80.97 |
| BRIGHT, nDCG@10 | 45.91 | 39.26 | 47.52 |
| Home appliances, exact case | 82.95 | 97.73 | 52.27 |
The post reports 43 benchmarks in all and prints a selection. Jev leads on When2Call and BRIGHT. On PhishNChips, a Jev variant called DiffusionGemma Jev scored 85.35 against 79.60 for Clef.
Who Jev is, and the context-window mismatch
The post does not say which Jev checkpoint it tested. A model with that name is trending on the Hub: autotrust/JEV-27B-VL from AutoTrust, built on the same Qwen3.8-27B base under Apache 2.0, with about 1.53 million downloads last month. Whether it is Cloudflare's baseline is answered in neither source.

Three context figures do not line up. The post gives Clef a 64k-token window against 32k for Jev. The card says its encoder accepts 16,384 tokens by default. The Qwen3.8-27B base is rated at 262,144 tokens natively. The card says the model was tested on a single H200 and describes no training data; the post says it is internal synthetic data.
Fine-tuning plan and what is missing
The post describes a two-phase fine-tuning service. The first is hands-on, run by Cloudflare's forward-deployed engineers. The second is self-serve, built on AI Gateway, Workers AI, Containers and a new trainer. Cloudflare said it is early and warned that fine-tuning can trade general performance for accuracy in one domain.
The card shows 10,874 downloads last month for Clef and 17,587 for Clef-flash. Clef already appears as a finetune of Qwen3.8-27B, and on-device competitors such as EmbeddingGemma 2 take a different route to small models. No self-serve date has been announced for the fine-tuning platform.
Sources
More in A.I.
- 01OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 02Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M EmbedderThe Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.
- 03Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on OneThe Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.
- 04Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One LineAlibaba's 27B dense model scores 61.7 on SWE-bench Pro by its own table. The card does not say what it was trained on or what hardware it needs.