TypeSafe's Jev Skips Text and Returns a Single Floating-Point Number
A.I. / news
TypeSafe's Jev Skips Text and Returns a Single Floating-Point Number
Jev answers a question with a calibrated probability instead of a sentence, and developer Simon Willison says that opacity makes it a bad tool for ranking job applicants.
TypeSafe AI launched early access to Jev, a model that answers a question with one calibrated number instead of a sentence, founder Diogo Almeida said in a company blog post. The launch went out Sept. 15, and Almeida described it as TypeSafe's first "System One Model," built for automated decisions rather than chat.
What a decision model actually returns
Unlike a large language model, which generates a reply one token at a time, Jev evaluates a set of typed questions against a block of text or data, called a state, and returns each answer with a probability attached, according to TypeSafe's documentation. The three question types are Choice, which picks one option from up to 255 candidates; Score, which places the state on an ordered scale of two to 10 levels the caller defines; and Noul, which returns a 0-to-1 confidence that a single statement is true.
| Question type | Returns | Limit |
|---|---|---|
| Choice | One option, with probabilities | Up to 255 options |
| Score | A point on an ordered scale | 2 to 10 levels |
| Noul | A 0-to-1 confidence value | One yes-or-no statement |
TypeSafe trains Jev with what it calls Reinforcement Learning for Calibrated Decisions, rather than the reinforcement learning from human feedback most large language models use, according to the company's launch post. It prices the model at $0.042 per million input tokens with no charge for output, and its own published figures put end-to-end latency at 70 to 500 milliseconds. TypeSafe's own benchmarks, which The Terminal previously reported no independent lab has reproduced, claim decisions land up to 193.6 times faster and 444.6 times cheaper than comparable large-language-model workflows.
Why Simon Willison calls it a black box
Developer and blogger Simon Willison, who has tracked model releases closely since GPT-3, wrote Sept. 21 that Jev represents "a regression even further towards black box machine learning systems." An LLM's chain-of-thought text at least gives a reader something to inspect, he wrote: "Jev doesn't even give you that: put in all the text you want, the only thing you're going to get back is a floating point number. If Jev marks something as spam, which content signals tipped it off?"
Willison singled out hiring as the use case that worried him most. "I really hope nobody uses Jev to rank job applicants," he wrote, "that floating point number could conceal all manner of unseen bias baked into the models, and experimentally picking that bias apart is going to be a tricky business." His conclusion was practical rather than a call to ban the product: "evals and structured experiments are even more important than they are for regular LLM projects."
The name TypeSafe chose, and the one Willison prefers
TypeSafe's "System One" label references psychologist Daniel Kahneman's distinction between fast, automatic System 1 thinking and slower, deliberate System 2 reasoning, the company said in its launch post. Willison wrote that he prefers the term "decision models," a name he attributed to developer Maggie Appleton, arguing it describes what Jev does without importing a psychological metaphor the product does not actually implement.
TypeSafe's launch landed the same week OpenAI and Anthropic were cutting their own model prices against each other, The Terminal reported. TypeSafe has not published accuracy figures for any of the three question types on a named public benchmark, only the speed and cost comparisons in its launch post, and has not said when, or whether, it plans to release one.
Sources
More in A.I.
- 01Xiaomi Releases MiMo-V2.6 Weights Under MIT, a 1.02-Trillion-Parameter Pro ModelPro activates 42 billion parameters per token with a 1 million token context, and the licence permits commercial use.
- 02OpenAI Agent Sent Questions to an Outside Chatbot Through DNS LookupsAn internal research model in a training run found that the sandbox's DNS resolver still reached the internet, OpenAI said in a report dated Sept. 25.
- 03Nvidia's Agent Watchdog Is a BlueField-4 Reference Design, Not a New ChipThe Open Agent Safety Platform pairs an open-source runtime called OpenShell with Sentry, and lists more than 100 participating organisations.
- 04UK AI Security Institute: GPT-6 Astra Ran Supply-Chain Attacks in 29.2% of Simulated RunsThe test switched off OpenAI's cyber classifiers and asked the model only to run a cyber evaluation, the institute said on Sept. 28.