The GenAI Field Guide · Trending
A hosted model that answers typed questions about your data with calibrated probabilities instead of generating text.
Jev is the first model from TypeSafe AI, a San Francisco lab that came out of stealth on 15 September 2026. It cannot chat or write: you hand it some state and a set of typed questions (pick one option, score on a scale, yes or no) and it returns typed answers, each with a probability distribution and a confidence value. If you run LLM calls today just to classify, route, score or gate things inside software, it is worth an afternoon of testing on your own data.
TypeSafe calls Jev a System One model, borrowing Daniel Kahneman's split between fast intuitive judgement and slow deliberate reasoning. The pitch is that a large share of what production LLM calls do is not writing at all. It is small judgements inside a program: which department does this ticket belong to, is this message urgent, does this passage answer the question, is this LLM output safe to show. Today those calls go through a text generator, get parsed back into a value, and sometimes fail to parse. Jev skips the text and returns the value.
It exposes three question types, which TypeSafe calls primitives. A Choice picks one option from a list of up to 255 and returns the choice, the probability of every option and a confidence. A Score rates the state against ordered levels (the docs and third-party guides describe 2 to 10 levels) and returns the level, its probabilities and a confidence. A Noul is a yes or no question that returns a single probability between 0 and 1. Several questions can go in one request; they are evaluated in parallel and in isolation against the same state, so one question does not leak into another's answer.
Under the marketing, the useful claim is narrow and checkable: because answers can only be drawn from the options you defined, Jev cannot return an invalid type or a value outside your schema. That is what TypeSafe means by zero hallucinations and a zero type-error rate. It does not mean the answer is right. A valid but wrong value is still possible, and the press coverage and TypeSafe's own docs both say so.
TypeSafe describes three pieces: a new model architecture, a parallel sampler that produces all outputs in one pass rather than token by token, and a training method it calls Reinforcement Learning for Calibrated Decisions (RLCD), aimed at probabilities that track how often the model is actually right. TypeSafe has not disclosed parameter counts or weights, and the model is only available as a hosted API. Marktechpost describes it as transformer based. None of the calibration claims have been independently measured in anything we could find.
Confidence is computed from the shape of the probability distribution, not reported by the model in words. For a Choice the docs give the formula (p_max minus 1/n) divided by (1 minus 1/n), which is 0 for an even spread across n options and 1 when all the mass sits on one option. Score confidence also accounts for ordering, so mass on an adjacent level costs less confidence than mass on a distant one. The docs suggest acting automatically above about 0.9, asking for confirmation between 0.5 and 0.9, and handing off to a person below 0.5, with thresholds set per action according to what a wrong answer would cost.
Everything goes through one endpoint, a POST to /v1/systemone on api.typesafe.ai, with a state, a model name and a map of named questions. The current model in the docs is jev-1.13.0, with aliases jev-latest and jev-preview. Input is text only; English is the most accurate language and others, including CJK, are handled less well. TypeSafe says Jev is not fine-tuned per customer and is not trained on customer requests; you steer it through the state, an instructions string and per-option criteria.
Speed and cost are the headline. TypeSafe's site claims Jev is 193.6 times faster and 444.6 times cheaper than LLMs on its workflow set (0.114 seconds against 8.566 seconds), and quotes 70 to 500 ms end to end. These are vendor-run numbers on workflows TypeSafe wrote, and the sources disagree on the baseline: Marktechpost reports the comparison against GPT-5.6 Terra, while TypeSafe's launch post describes the baseline as an average of GPT-6 Astra and Fable 5.1. InfoQ reports median user-reported figures of 7 times faster and 30 times cheaper, which are far smaller than the headline multiples.
You need a TypeSafe account and API key from console.typesafe.ai, or a Vercel account to call it through Vercel AI Gateway as typesafe-ai/jev. Access has moved around since launch (see availability), so check that signups are open before planning around it. The Python SDK needs Python 3.10 or later and reads the key from TYPESAFE_API_KEY.