Jev and the class of models that decide instead of chatting — why a typed probability in milliseconds is worth more, and costs less, than a frontier paragraph.
5 concepts 4 decision paths Diagrams
Share
The class, the interface, and the bill
Jev is TypeSafe AI’s first System One model, announced 15 September 2026. The class is theirs; the placement argument — cheap judgment in front of an expensive paragraph — is the read for anyone already paying a frontier bill for classification.
A new family
System One models decide. They do not chat.
"A sorting office, not a novelist. You hand it a parcel and a list of destinations; it stamps one, with a probability. It cannot write you a poem about the parcel."
System One is TypeSafe AI's name for models built to make fast, structured decisions that software can use directly — inspired by Kahneman's fast judgment, and named Jev after Jevons, because cheaper intelligence tends to get used more, not less. Jev, released 15 September 2026, is the first public instance. Like an LLM it reads natural language. Unlike an LLM it does not emit a string.
What it actually is
The model evaluates a state (text, JSON, or an array of text) and a set of questions whose answer space you define in advance. Training is Reinforcement Learning for Calibrated Decisions (RLCD): probabilities are optimised against outcomes so that higher confidence tracks higher accuracy. Sampling is parallel, not autoregressive. Schema match is guaranteed; a type error is not an empirical failure mode.
The useful test is not "is this cleverer than a frontier chat model?". It is "does this call need a paragraph, or a decision my code can branch on?". Most production traffic is the second. Sending that traffic to a frontier chat model is couriering a Post-it note.
A new family
System One models decide. They do not chat.
"A sorting office, not a novelist. You hand it a parcel and a list of destinations; it stamps one, with a probability. It cannot write you a poem about the parcel."
System One is TypeSafe AI's name for models built to make fast, structured decisions that software can use directly — inspired by Kahneman's fast judgment, and named Jev after Jevons, because cheaper intelligence tends to get used more, not less. Jev, released 15 September 2026, is the first public instance. Like an LLM it reads natural language. Unlike an LLM it does not emit a string.
What it actually is
The model evaluates a state (text, JSON, or an array of text) and a set of questions whose answer space you define in advance. Training is Reinforcement Learning for Calibrated Decisions (RLCD): probabilities are optimised against outcomes so that higher confidence tracks higher accuracy. Sampling is parallel, not autoregressive. Schema match is guaranteed; a type error is not an empirical failure mode.
The useful test is not "is this cleverer than a frontier chat model?". It is "does this call need a paragraph, or a decision my code can branch on?". Most production traffic is the second. Sending that traffic to a frontier chat model is couriering a Post-it note.
Function call
Unstructured state in, typed probabilities out
"A function call with a type signature. The surrounding code is the type system; the model is not allowed to return a sonnet."
You send state and questions. You get answers whose shape you declared: a choice from a closed set, a score on an ordered scale, or a probability that a statement is true. Every answer carries a confidence. Several questions on the same state run in one request, in parallel — extra questions barely change latency, and you pay only the tokens that describe them.
What it actually is
Three primitives. Choice returns a labelled option, a distribution over the set, and a confidence. Score returns a continuous value against ordered criteria, plus the underlying distribution. Noul (yes/no) returns the probability the statement holds. Cardinality on a single choice is capped (Jev currently at 255); above that you score independently, then choose. Images, audio and video are not accepted.
This is what makes the output composable. A chat model that "usually" returns JSON still needs parsing, validation, retries, and a human near the blast radius. A typed probability can sit in an if-statement, a queue priority, or a circuit breaker without any of that theatre.
Two orders cheaper
You stop paying for tokens you never needed
"An F1 car to deliver a stamped form. The form needed a franking machine."
Jev list-prices input at $0.042 per million tokens and does not meter output. End-to-end latency on TypeSafe's figures is 70–500ms, against 3–329 seconds for frontier chat models doing equivalent System One shaped work. On a 2,000-token decision, that is about $0.000084 a call — $84 for a million — against $25,000 if the same million is a reasoning frontier call that thinks for 2,000 tokens, or $6,500 if it only emits structured JSON.
What it actually is
The cost gap has three mechanical sources, not a promotional discount you should assume lasts forever. No generated tokens: output is free because there is almost nothing to generate. Parallel sampling: all answers arrive in one forward pass, so latency does not grow with the number of questions the way a chain-of-thought trace does. And you are not paying a reasoning model to narrate a classification it could have returned as a label. TypeSafe's workflow evals (193.6× faster, 444.6× cheaper versus the largest chat models, using those models as the reference) are the high end; the worked example above is the number you can check on a napkin.
The bill is the visible win. The quieter one is what the leftover budget buys: decisions inside a sub-second request path, classifiers you can run on every ticket rather than a sample, and an agent loop that does not call a frontier model to ask whether it should call a frontier model.
In front, not instead
A cheap judgment in front of an expensive paragraph
"A night-watchman in front of the expensive consultant. The watchman decides whether the consultant even needs to be woken."
Jev is not a drop-in replacement for an LLM. It cannot write the reply, the plan, or the code. It can classify, route, score, extract a decision, verify a trace, and refuse a tool call — the steps that currently burn frontier tokens because that is the only hammer on the bench. Put it on those steps; keep the LLM for the fraction that genuinely needs language.
What it actually is
The pattern that has landed first is a harness: a System One call on the latest user turn or tool result, then a branch in code. Model routing (cheap versus capable), auto-mode guardrails on irreversible tools, ticket triage, eval-as-a-judge, and map-reduce over large corpora are the same shape. Confidence is the gate: below a threshold you escalate to a person or a reasoning model, rather than hoping the chat model will say when it is guessing.
This is the FinOps control that was missing from the routing story. An estate without a cheap, typed decision layer has a cost floor set by its most expensive model, because every fork in the workflow is itself a frontier call. Lower the floor, then spend the frontier budget where a wrong paragraph actually costs you.
Closed answers
If you cannot name the answers, it is not this job
"Three stamps on the sorting desk: which pigeonhole, how far along a scale, and how sure it is a yes."
Choice, Score, and Noul are the whole surface. Mix them on one state. The constraint is the feature: if you cannot write the answer set down, you do not yet have a System One task — you have a writing task, and you still need a language model. Forcing a closed set onto an open problem just relocates the hallucination from the tokens to the schema.
What it actually is
Choice needs criteria for each option, not just labels. Score needs an ordered legend so the number means the same thing tomorrow. Noul is a probability, not a boolean dressed up — 0.51 and 0.99 are different actions in a workflow even if both round to true. Independent questions on one state is the reliability pattern: decompose, then let your code combine the probabilities. A single giant "what should we do?" question throws the judgment back into language.
The design work moves from prompt poetry to specifying the decision. That is a better use of a senior engineer's time, and it is the only way the result stays reviewable when the person who wrote the prompt has left.
The cost of one million decisions
Same workload, four model classes — 1 million structured decisions a month, 2,000-token state each.
Monthly bill, volume held fixed
A support ticket plus a short policy snippet, or one step of an agent loop
Frontier reasoning, every decision
2,000 thinking tokens out, $2.50 / $10 per million
$25K/mo
Frontier structured output
Same model, no reasoning trace, ~150 tokens out
$6.5K/mo
Cheap LLM classifier
Small-model API at $0.20 / $0.60, still generating tokens
$460/mo
System One / Jev
$0.042 per million in, output free, typed answers
$84/mo
Frontier reasoning at $2.50 / $10 per million with 2,000 thinking tokens out. Jev at $0.042 per million input, output free. TypeSafe’s 445× figure is versus the largest chat models on their own workflow evals — the high end, not this napkin.
End-to-end latency
The same decision, timed rather than priced
Frontier reasoning
Seconds to minutes. Fine for chat, lethal in a request path.
3–329s
Fast chat LLM
Usable for back-office, awkward inside UX.
300–800ms
System One / Jev
Inside the latency budget of a web request.
70–500ms
TypeSafe’s published range for Jev is 70–500ms. Frontier chat models doing equivalent System One shaped work land between 3 and 329 seconds. Fast enough for a person; a bottleneck inside a request path.
Decision framework
Frequently asked questions
What is a System One model?
A class of AI models built to make fast, structured decisions that software can use directly. They read natural language, like an LLM, but they return typed values and calibrated probabilities rather than generated text. Jev, released by TypeSafe AI on 15 September 2026, is the first public instance. The name draws on Kahneman’s fast judgment; Jev is named after Jevons, because cheaper intelligence tends to get used more, not less.
How is Jev different from a small language model doing classification?
A small LLM is still generating tokens. You prompt it for JSON, parse the result, retry when the schema fails, and hope it does not invent an extra field. Jev cannot leave the schema you declared — Choice, Score, or Noul — and it answers every question in a request in one parallel pass. Output tokens are unmetered because there is almost nothing to generate. The cost gap versus a cheap classifier is real but modest; the operational gap is type-safety, latency, and asking several questions in one call.
How much cheaper is it, really?
On a worked example of one million 2,000-token decisions a month, Jev list-prices at about $84. The same million as frontier reasoning with 2,000 thinking tokens out is about $25,000; as frontier structured output with no reasoning, about $6,500; as a cheap classifier API, about $460. TypeSafe’s own workflow evals claim 193.6× faster and 444.6× cheaper versus the largest chat models, using those models as the reference — they flag that as the high end. Check the napkin against your own state length before you put the ratio in a board paper.
Does this replace frontier models?
No. Jev cannot write the reply, the plan, or the code. It belongs on the steps that currently burn frontier tokens because that is the only hammer on the bench: classify, route, score, gate, verify. Keep the language model for the fraction that genuinely needs language, and use the System One confidence score as the gate between the two. An estate without a cheap, typed decision layer has a cost floor set by its most expensive model.
When should I not use a System One model?
When you cannot name the answers in advance — open-ended reasoning, novel situations, anything that resists a closed set. When the output has to be language a person reads. When the input is an image, audio or video: Jev currently accepts text, JSON, and arrays of text only. And when a wrong closed-set label is more expensive than a slow, hedged paragraph; in that case you still want a reasoning model, with a human on the irreversible step.
Where does TypeSafe’s claim come from?
From TypeSafe AI’s 15 September 2026 announcement of System One models and Jev, and from their published workflow evals. Prices, latency ranges, the three primitives, and the training method (Reinforcement Learning for Calibrated Decisions) are theirs. The placement argument and the worked cost example on this page are the practitioner read of those numbers, not a reprint of the launch post.