Cloudflare's Clef Reports 209 ms Decisions, but the Outside Tests of Decision Models So Far Cover Only Jev
A.I. / analysis
Cloudflare's Clef Reports 209 ms Decisions, but the Outside Tests of Decision Models So Far Cover Only Jev
Red Hat, Cisco and two University of Pennsylvania researchers have benchmarked the category's first model. None of them tested Clef.

Decision models, which return a probability for each permitted answer instead of writing text, now have a second open model. Cloudflare's Clef model card describes a 27-billion-parameter model that takes a state and a schema of typed questions and scores every option in one pass. The independent tests published so far measured a different model, AutoTrust's Jev, and found a mixed record.
That gap is the claim worth pinning down. Every Clef number below comes from Cloudflare's own card.

What Clef is
Clef is post-trained from Qwen3.8-27B under the Apache-2.0 licence, with ungated weights. A joint schema head reads the backbone's final hidden states and scores all options of all questions together, so there is no free-form text generation and no output parsing.
The card says the default input is 16,384 tokens and that it was tested with torch 2.11 and transformers 5.10.2 on a single H200. A smaller sibling, Clef-flash, is a 9B model built on Qwen3.5-9B, also Apache-2.0.
Jev, from AutoTrust AI Lab, follows the same pattern on the same Qwen3.8-27B backbone. Its card claims 78.3% on VL-RewardBench and reports 1,525,286 monthly downloads.
How Jev's own claims compare
AutoTrust's card describes two modes. System 1 returns typed decisions, which are yes or no answers, a choice among up to 256 options, or a 0 to 5 rating, with calibrated probabilities. System 2 is the unmodified Qwen3.8-27B, optionally thinking step by step.
The card gives about 1.4 seconds for 24 text decisions on one B200 GPU and roughly 13 image decisions a second on the same card. It also claims 78.3% on VL-RewardBench, which it describes as above every model on that leaderboard. These are vendor-supplied, like Cloudflare's.
One line in the card is useful for calibration. Jev scores 78.0% pass@1 on HumanEval, which the card says is identical to Qwen3.8-27B, so the decision head leaves the base model's coding ability alone. It says nothing about whether the head beats a trained classifier.
Cloudflare's latency table
The card reports an internal run of what it calls the Decision Index, 43 benchmarks in version 0.2.1. It does not say who built the index or how tasks were chosen, so treat every figure as vendor-supplied.
- Jev524.1 ms
- Clef209.3 ms
- DiffusionGemma Jev84.4 ms
- Kev 9B51.4 ms
- Clef-flash38.8 ms
- Laya5.8 ms
Source: Cloudflare/clef model card, accessed 2026-10-06 (vendor-supplied)
Clef's 95th-percentile latency is 238.6 ms. Clef-flash is 122.4 ms, and Laya's is 222.5 ms despite its 5.8 ms median.
On accuracy the card puts Clef at 98.5% on BFCL, 94.2% on BANKING77 and 80.8% on GSM8K. Clef-flash scores 98.8% on BFCL but 34.2% on SGD/SGD-X and 1.6% on POP909-CL, so the average hides wide swings.
What outside tests found about Jev
Three independent groups have tested Jev, and the results do not agree on a headline.
| Tester | Result for Jev | Caveat |
|---|---|---|
| Red Hat, October 2 | 86.35% on prompt injection, fourth; 86.20% on content safety, first | Prompt tuning may favour some baselines |
| Cisco | 34% recall at 0.5% false positives, against 61% for Cisco's 1B encoder | Labels were LLM-generated |
| Penn, arXiv 2609.29769 | Significantly different from an LLM judge in 8 of 27 comparisons | Closed rubrics only |
Red Hat's Rob Geada, Mac Misiura and Shelton Cyril tested nine guardrail candidates. On prompt injection Qwen3.6-35B won at 89.31% and deberta-v3-base followed at 89.01%. On latency deberta-v3-base-prompt-injection-v2 had a 54.1 ms median against Jev's 348.1 ms.
The authors concluded that decision models "do not reliably outperform LLM-as-a-judge, pre-trained predictive models". Cisco's Prashanth Arun, Director of Data Science for AI Defense, and Jin Ma, a machine learning engineering technical leader, wrote that "policy-specific training still delivers the best recall at strict false-positive budgets". They also wrote that their reference labels were "model-generated", not human-adjudicated.
Delip Rao and Chris Callison-Burch of the University of Pennsylvania found the opposite on cost. Against GPT-5.6 Luna, Gemini 3.8 Flash and DeepSeek V4.1 Flash, LLM judges cost 29 to 325 times as much as Jev and took 30 to 220 times as long.
What would settle it
The tests agree on one point. Decision models are cheap and fast against LLM judges, and not reliably better than a classifier trained on the task. None of that is yet shown for Clef, because Cisco's write-up states that Clef was not tested and the Red Hat study ran Jev-1.13.0 only.
The Cloudflare card also leaves a gap. It lists no limitations section, and its tables show a model that is near 99% on tool-call selection and near 2% on a music-chord task.

A falsifying test is easy to describe. Run Clef on Red Hat's prompt-injection set and Cisco's 0.5% false-positive budget and compare it with deberta-v3-base. Until someone publishes that, the 209.3 ms figure is Cloudflare's.
The same week, TesterArmy's open-source project drew attention for publishing no accuracy figures, and Cloudflare also released an agent workspace with closed contributions.
Sources
More in A.I.
- 01Reflection Announces Beam, a 501B Open-Weight Model, but Has Not Named a LicenceThe weights are promised for later in October. Reflection's own post gives scores and training scale, and says nothing on licence terms or API pricing.
- 02Meta Disputes Inc. Columnist's Claim That Muse Read His Messages While Mac Researcher Calls It a BackdoorTwo separate Muse problems are being reported as one. A researcher's local-access flaw is documented; the claim that the agent read private messages is contested by Meta.
- 03Anthropic Staff Reported a Claude Chat to Police, and a Florida Woman Faces a Felony ChargeThe company's privacy policy allows disclosure to prevent serious harm. Its transparency report counted zero emergency requests from police, and does not count referrals it makes itself.
- 04Vals AI's 90 Claude Opus 5.5 Agents Name Two Magnetic Semiconductor Candidates, Neither Yet MeasuredOne candidate was first synthesised in 1999 and has a measured ordering temperature of 376 K. The other may not survive the furnace.