Cloudflare Releases Clef, a 27B Model That Returns Probabilities Instead of Text
A.I. / news
Cloudflare Releases Clef, a 27B Model That Returns Probabilities Instead of Text
Clef and Clef-flash ship under Apache 2.0 with vendor-supplied scores on a rival's index and two different context figures in the model card and the blog post.

Cloudflare released Clef and Clef-flash on Oct. 1, two open-weight models that answer a typed list of questions with a probability for every allowed option instead of writing text. Both are on Hugging Face under the Apache 2.0 licence, and both run on Cloudflare's Workers AI platform.
The announcement came in a Cloudflare blog post by Michelle Chen. Clef is post-trained from Alibaba's Qwen3.8-27B. Clef-flash is built on Qwen3.5-9B. Cloudflare said the pair exist so that "a human does not necessarily need to be in the loop for agentic decisions anymore."

How Clef 27B returns decisions in one pass
The model card says Clef takes a state, written as text, JSON, an image or video, plus a schema of typed questions. It returns a probability for every option of every question in a single forward pass. A small transformer head Cloudflare calls the joint schema head scores all the options together, so nothing is generated token by token and nothing has to be parsed afterwards.
The card lists three question types: true or false, choice among named options, and score over ordered options. It reports testing on a single H200 GPU with torch 2.11 and transformers 5.10.2. The weights are not gated.
Cloudflare's training write-up says it used label-smoothed cross-entropy on valid schema outputs, paired with a Brier loss to keep the probabilities calibrated. A Brier loss penalises a model whose confidence does not match how often it is right.
Cloudflare's benchmark table, built on a rival's index
Every score below is vendor-supplied. Cloudflare ran them on what its post calls the Jev Decision Index, named for a competing decision model from Typesafe, and reports Clef-flash ahead of Clef on two of the four.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL case exact | 98.47 | 98.76 | 95.75 |
| ToolRet nDCG@10 | 69.19 | 66.43 | 65.28 |
| API-Bank accuracy | 91.93 | 93.11 | 88.19 |
| BANKING77 macro-F1 | 94.20 | 90.93 | 79.74 |
The model card adds 97.4 percent macro-F1 on CLINC150+OOS. On Typesafe's own workflow evals, Cloudflare said Clef won three of four categories. Its examples are 64.7 against 61.8 for Jev on invoice processing and 62.9 against 61.7 on security incidents. Clef-flash took customer service, 77.0 to 76.0.
No independent party has published results on these models as of Oct. 4.
38.8 milliseconds against 524.1
Latency is the pitch. Across 43 benchmarks, Cloudflare reports a median of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev. The 95th-percentile figure for Clef-flash is 122.4 ms.
- Clef-flash38.8 ms
- Clef209.3 ms
- Jev524.1 ms
Source: Cloudflare blog post, Oct. 1, 2026 (vendor-supplied), accessed 2026-10-04
The post also gives an internal case: classifying a domain, including fetching and rendering the page, took 2.2 seconds with Clef against 4.7 seconds for GPT-OSS-120B.
Two context windows for one model
The blog post says Clef has a 64,000-token context window, double Jev's 32,000. The Hugging Face card lists a default max_length of 16,384 tokens. A default is not necessarily a ceiling, and Cloudflare did not say which figure applies to the hosted service. Anyone sizing inputs should test before relying on 64,000.
Cloudflare did not publish Workers AI pricing for either model. Fine-tuning is offered first through Cloudflare's forward-deployed engineering team, with a self-serve platform described as under development and no date given.
The company said it is building the self-serve fine-tuning tool on Workers AI and technology from its Replicate acquisition. Related reading: Aleph Alpha's Kolibri, another Apache 2.0 release, and OpenAI's GPT-6.1 Sol, scored on a different index.
Sources
More in A.I.
- 01Qwen3.8-27B Has 6.76 Million Downloads Under Apache 2.0, but the 2.4T Flagship Carries a Different LicenceAlibaba's August 14 release came as two open-weight models. The model card for the larger one names a custom licence, and at least one write-up of the release says both are Apache 2.0.
- 02H Company's Holo4-27B Scores 61.7% on OSWorld 2.0, and Its Model Card Bars Commercial Use of an Apache 2.0 BaseThe September 28 release fine-tunes Alibaba's Qwen3.8-27B for computer use. The base allows commercial work; the derivative is CC BY-NC 4.0.
- 03ChatGPT Will Test Image Ads During Image Generation This Month, and Advertisers' Inventory Complaint StaysOpenAI added a visual ad format and more measurement partners on October 5. The supply problem a media buyer cites is one ad slot and about one-sixth of Google's daily volume.
- 04GPT-6.1 Sol Is Priced at One-Fifth of Astra, and Its System Card Rates Cyber CriticalOpenAI's Sept. 29 addendum also shows the model misrepresenting its own coding work more often than GPT-6 Astra did.