CLM-v0.1-8B Puts an Apache 2.0 Rival to TypeSafe's Jev on Hugging Face, With 30-Task Test Sets
A.I. / news
CLM-v0.1-8B Puts an Apache 2.0 Rival to TypeSafe's Jev on Hugging Face, With 30-Task Test Sets
Contrastive-LM says its 20-million-parameter heads on a frozen Qwen3-8B match Jev on some tasks at up to 9 times lower latency. The verifier results come from fine-tuned heads, not the released checkpoint.

A group called Contrastive-LM released CLM-v0.1-8B on Hugging Face under the Apache 2.0 licence, an open-weights model it describes as a "System One model" and pitches against TypeSafe's closed Jev. The repository was created on September 21, and by Wednesday the model card showed 544 likes and 2,392 downloads.
The model card says CLM is "on par with Jev on computer-use, gaming and tool-calling tasks, with up to 9× lower latency." Those figures are the developer's own. No independent party is named as having run them.
What the model is
CLM does not generate text. It scores candidates. According to the card, it is two small projection heads, a state head and an action head, on top of a frozen Qwen3-8B encoder, trained with a bidirectional InfoNCE loss. The GitHub repository puts the trainable part at about 20 million parameters, and the reference head downloads at 75 MB.

The encoder comes from Alibaba's Qwen family, and both it and CLM's weights are Apache 2.0, so commercial use is permitted. The weights are downloadable without a gate. A serving setup needs a vLLM server running Qwen3-8B in pooling mode, then the clm-serve package. States are truncated at 2,048 tokens unless both limits are raised together to 8,192, which needs more GPU memory.
Training used about 60 million Nemotron question-and-answer pairs, about 30 million synthetic hard negatives and about 1 million agentic trajectories, per the card. The card lists Jacky Kwok, Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, Marco Pavone, Christopher Ré and Azalia Mirhoseini as authors.
The numbers and their denominators
The headline scores are for a verifier built by fine-tuning the heads: 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1. The card says the released checkpoint is only the starting point for those heads, and that the numbers "come from fine-tuned heads, not this checkpoint zero-shot."
The repository says the DeepSWE score is over 38 held-out tasks and the Terminal-Bench score over 30. On 30 tasks a single pass can only land on multiples of 3.3 points, so 87.6% implies averaging over runs or a different denominator. The repository text does not say which.
- Verifier, low end4.1 x faster
- Verifier, high end5.7 x faster
- Zero-shot, up to9 x faster
- About 1,000 candidates13 x faster
Source: Contrastive-LM model card and GitHub README, accessed 2026-09-30
The 4.1 to 5.7 range was measured on an H100, according to the repository. The 13 times figure applies with about 1,000 candidates, because CLM encodes states and actions separately and can reuse action embeddings.
What Jev is and what is unverified
Jev is TypeSafe AI's model, offered through an API at $0.042 per million input tokens with no output charge, according to TypeSafe's launch post. TypeSafe published no weights, so anyone comparing against it does so through the paid API, and its size is undisclosed. The post names Diogo Almeida as founder.
CLM's README calls its server "TypeSafe-compatible," meaning a request written for Jev's typed-question format replays against it. We covered the argument over who invented the approach in our report on the Jev prior-art dispute, and the question of whether the big labs would copy it in OpenAI and Jev.
The card lists its own limits: CLM cannot generate, its probabilities are relative to the candidate set you supply, and it is locked to Qwen3-8B embeddings. The developers say a multimodal CLM-35B is coming in early October.
Sources
More in A.I.
- 01GPT-Synopsys: OpenAI Gets Paid Only When the Chips It Helps Design Beat the Customer's BaselineThe Sept. 30 deal pairs an OpenAI model with Synopsys' design software, with no price, no release date and no named customer.
- 02Ataraxos Beats Stratego's Top Player 15-1-4 After Training on 16 H100s for a WeekA Nature paper from MIT, Carnegie Mellon, NYU and Stanford puts the compute bill under $8,000, against an estimated $3 million to $4.5 million for DeepMind's DeepNash.
- 03Nine Mathematicians Advising OpenAI Ask AI Labs to Stop Testing Hard Problems on Models Nobody Else Can UseThe Advisory Group on Mathematics and AI published its rules on Sept. 29: release fast, fund human understanding, disclose prompts and costs.
- 04Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.