OpenAI's Decisions API Bills $0.10 Per Million Input Tokens and Nothing for Output
A.I. / news
OpenAI's Decisions API Bills $0.10 Per Million Input Tokens and Nothing for Output
The public beta runs only gpt-6-luna, returns probabilities instead of prose, and follows Strands Labs' open-source 2B model built for the same job.
OpenAI has put its Decisions API into public beta. It is a single endpoint, POST /v1/decisions, that answers yes-or-no, multiple-choice and rating questions with probabilities rather than text. Input costs $0.10 per million tokens and output is not charged.
Details come from OpenAI's Decisions guide, which calls the beta public and says general availability is expected. The guide reached 302 points on Hacker News.
What the endpoint does
A request has three parts: the model, the input evidence and a list of questions. Evidence can be text, images or both. Images must be inline base64 data URLs, because hosted image URLs are not supported.
Each question needs a unique name. Independent questions can share one request. A decision that depends on an earlier answer needs a separate, later request.
| Question type | Returns | Example in the guide |
|---|---|---|
| Predicate | A probability from 0 to 1 | Visible damage in a product photo |
| Choice | The chosen option, per-option probabilities, confidence | Routing a complaint to a department |
| Score | A probability-weighted rating, confidence, per-level probabilities | Severity on an ordered scale |
OpenAI says the API answers about 10 times faster than its Responses API. That is a vendor figure, and the guide as read gives no benchmark behind it.
Price and the one supported model
Only gpt-6-luna works on the endpoint today. OpenAI's model page describes Luna as its most efficient model for focused, high-volume tasks, with a 1,050,000-token context window.
| Route to gpt-6-luna | Input per 1M tokens | Output per 1M tokens |
|---|---|---|
| Decisions API | $0.10 | $0 |
| Standard model pricing | $0.10 | $0.50 |
Input pricing matches the standard rate, so the saving is the output line. The guide adds that regional processing carries a premium, and that Zero Data Retention and HIPAA compliance are available for eligible customers in US and European regions.
An open-source 2B model does the same job
On Oct. 1 the Strands team released Strands Decider 2B. Marc Brooker, Mike Chambers and Fabio Nonato de Paula wrote the post, which credits Strands Labs.
The model starts from Qwen3.5-2B. Its language head is replaced by a pointer head of just over 1 million parameters that scores the answers offered for each option, and the body is tuned with a rank-16 LoRA adapter. It picks from a list and does not write text.
The post says median latency is about 115 ms on an Nvidia RTX 3090 and about 153 ms on an M3 MacBook for small tasks. It ranks the model third of 33 in the 2B class on JevBench, and first of 30 once marginally larger models are excluded. Those are Strands' own results. Accuracy and calibration, by Brier score, are the metrics it names.
Hugging Face's StrandsAgents page lists two checkpoints, hobson-v19 and hobson-v21, with 61 and 15 downloads when read. The page did not show a licence, and the blog calls the release open source without naming one.
What neither source answers
OpenAI has not published accuracy or calibration figures for Decisions, so a buyer cannot yet compare it with Strands Decider or with the Clef and Jev models covered in Cloudflare's Clef reports 209 ms decisions. The endpoint also locks users to one model, which rules out a cheaper or larger choice.
The two products make the same bet. A fixed menu of answers with a probability attached is easier to wire into a workflow than a paragraph that has to be parsed. OpenAI's other recent release is covered in GPT-6.1 Sol's pricing.
OpenAI says general availability is expected. The guide gives no date.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.