The GenAI Field Guide · Trending
A new class of model that cannot write text: you give it state and a set of typed questions, and it returns probabilities over the answers you defined.
Two weeks after TypeSafe AI launched Jev on 15 September 2026, decision models became a category. OpenAI announced a Decisions API built on GPT-6 Luna at DevDay on 29 September, and on 1 October Cloudflare released the open-weight Clef models and AWS released the open-source Strands Decider 2B, joining Together AI's Tev1 from late September. If your software calls an LLM just to classify, route, score or gate something, this is the pattern to evaluate now, ideally with an open-weight model on your own labelled data.
A decision model takes some state (a ticket, a document, a tool result, sometimes an image) and one or more questions, each with a fixed set of allowed answers, and returns a probability for every allowed answer. It does not generate prose. Simon Willison summed up Jev as unstructured state in, typed probabilistic decisions out, and suggested the name decision models for what TypeSafe calls System One models, after Kahneman's fast, intuitive System 1. Most of the new entrants copy Jev's three question types: a Choice (pick one of several options), a Score (place the state on an ordered scale) and a Noul (a yes or no question that returns a probability between 0 and 1).
OpenAI's Decisions API is the highest-profile entrant and the least documented. OpenAI's DevDay material, as quoted on its developer forum, says the API is in limited preview and uses Luna to classify inputs, route requests, or choose an action from predefined answers. Context can be text or images. The Decoder reports a specialised version of GPT-6 Luna answering in about 150 ms against 1.6 seconds for a normal Luna call, a roughly tenfold speedup; that figure appears only in secondary coverage, and Firecrawl calls it a vendor claim under unstated conditions. As of 2 October there is no public endpoint, schema, price or rate card, and OpenAI's API changelog has no Decisions API entry. We could not open OpenAI's DevDay recap page directly (it returned 403).
The open side moved faster. Cloudflare's Clef (27B, post-trained on Qwen3.8-27B) and Clef-flash (9B, on Qwen3.5-9B) are Apache 2.0 on Hugging Face, hosted on Workers AI, and API compatible with Jev. AWS's Strands Decider 2B, from Strands Labs and built by Marc Brooker's team, is Apache 2.0 and runs on a laptop. Together AI's Tev1-4B-experimental is a LoRA fine-tune of Qwen3.5-4B released with its full data recipe. TechCrunch headlined Amazon's release as a Jev clone as decision models flood the web.
Do not confuse two products with the same name. OpenRouter already runs an endpoint called the Decisions API, at /api/alpha/decisions, which serves TypeSafe's Jev. It predates OpenAI's product and is unrelated to it.
The common trick is to keep a pretrained language model's body and change what comes out of it. Instead of decoding tokens one at a time, the model reads the state, the questions and every candidate answer in one prefill pass, then scores the candidates directly. Strands Decider removes Qwen3.5-2B's language-model head and adds a pointer head of just over a million parameters that scores the hidden state at each option's position against the hidden state at an answer marker; the torso is tuned with a rank-16 LoRA. Cloudflare describes Clef as non-autoregressive: a frozen base with rank-256 adapters, valid schema choices scored in parallel, and fields that cross-attend to each other and to the payload. Because the output is a distribution over your options, an invalid answer is impossible by construction, and a confidence number falls out of the distribution's shape.
Calibration is trained, not assumed. Cloudflare says it trains Clef with label-smoothed cross-entropy plus a Brier loss for calibration and uses RLCD (Reinforcement Learning for Calibrated Decisions, the name TypeSafe gave its own method) as a second stage. Together's Tev1 takes the simpler route: ordinary supervised LoRA on 37,840 examples (rank 8, one epoch), keeping Qwen's normal output head and constraining the answer to option letters. Its README says the token log probabilities it returns reflect preference and are not calibrated confidence. Together says on X that Tev1 cost about $17 to train.
Speed numbers are all vendor measured and should be read that way. Cloudflare reports, across 43 benchmark runs, median latency of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev, and says Clef scores highest on 7 of 10 decision benchmarks (for example Banking77 macro-F1 of 94.20 against 79.74 for Jev). AWS reports a median of about 115 ms per question for Strands Decider on an RTX 3090 and about 153 ms on an M3 laptop, 0.723 accuracy on the 231 public JevBench tasks, and third of 33 in its size class. Jev's maker quotes 70 to 500 ms. None of these have been reproduced independently in anything we found, and TypeSafe's CEO told TechCrunch the new batch looks more like ML people implementing a cool architecture than teams dedicated to making intelligence useful.
OpenAI's Decisions API is limited preview for selected customers, so most readers cannot call it yet. You can use a decision model today through Cloudflare Workers AI (a Cloudflare account and API token), TypeSafe's Jev directly or through OpenRouter, or by running Strands Decider 2B or Clef weights yourself. The example below is Cloudflare's own Workers AI call from its launch post.