Decision Models (DMs)
Canonical version: Decision Models (DMs).
A decision model (DM) is an AI model whose output is a decision, not text. You give it an input and a bounded set of possible answers, and it gives you back numbers: which option, how likely yes is, what score. It's the category Jev belongs to, and I think it deserves its own name next to Large Language Models (LLMs) and Small Language Models (SLMs).
Why bother with a separate term? Because a LOT of what we ask LLMs to do today is actually deciding. Is this email urgent? Which tool should the agent call? Is this paragraph off-topic? Did the agent get stuck? We ask a text generator to write the answer as JSON, parse it, validate it, and retry when it breaks. A decision model answers that kind of question directly, and that changes the economics: when a decision costs a fraction of a cent and takes 200 ms, you can afford to ask it everywhere.
What makes a model a decision model
- The answer space is bounded and known upfront. Options, yes/no, or a numeric range. If you can't list the possible answers, you need an LLM
- The output is typed by construction. There's no text to parse, so malformed output can't happen
- The output is a probability. Ideally a calibrated one: when it says 80%, it's right about 80% of the time. That lets you set real thresholds ("send to a human below 0.7"), something you can't do with an LLM writing "I'm 90% sure"
- No generation loop. The input is read once and the answer comes out of a readout layer, not token by token. That's where the speed comes from
An old idea with a new twist
Decision models aren't new. We've had them for decades under other names:
- Classic classifiers: spam filters, sentiment analysis, intent routing (see Binary Classification and Data classification). Fast and cheap, but trained for ONE task with your own labeled data
- Zero-shot classifiers: NLI-based models (e.g., BART fine-tuned on MNLI) that score arbitrary labels without training. Closer to the idea, but limited in quality
- Rerankers and reward models: they score an input against a query or a preference. Decisions again, just packaged for search and for training LLMs
- LLMs used as deciders: prompt an LLM, force LLM Structured Outputs, or read the token probabilities of each option. It works, but you pay for a text generator to pick a label
What changed in 2025-2026 is generality. A general decision model answers arbitrary questions over arbitrary options at inference time, with no task-specific training, at a quality close to a frontier LLM. Laya went in that direction in March 2025. TypeSafe AI made it a product with Jev in September 2026, and calls the category System One Models, after Daniel Kahneman's fast, intuitive System 1. I prefer "decision model" as the generic term: it says what the thing outputs, and it doesn't depend on one company's branding.
DMs vs LLMs
| Decision model | LLM | |
|---|---|---|
| Output | A choice, a probability, a score | Text (possibly JSON) |
| Answer space | Bounded, defined by you | Open |
| Latency | Tens to hundreds of ms | Seconds |
| Cost | Fractions of a cent per decision | Much higher, output tokens included |
| Failure mode | A wrong but well-formed answer | Malformed output, hallucinated options, wrong answer |
| Explanation | None, just a number | Can explain (whether the explanation is faithful is another story) |
They're complementary. The pattern I find the most useful (from Nate B. Jones) is to think in three building blocks: code for rules and calculations, decision models for bounded judgment, LLMs for open reasoning, writing and exceptions. A decision model can even decide when an LLM is worth calling, which is Model routing done cheaply.
Current examples
- Jev by TypeSafe AI: closed, the first commercial general decision model
- Laya: open, multilingual, published in 2025
- Kev: open models on top of Qwen with a Jev-compatible API
- SemIf (formerly OpenJev): reads the probabilities of each answer from an LLM's output, runs locally
- Jevlike: a small trainable option scorer, closer to a classic classifier
- fast-jev-compaction: an example of what DMs enable, compacting a Claude Code session by scoring each tool call instead of summarizing
Where to use them
A decision model rewards a specific way of building: many narrow questions answered in parallel, combined in code (Atomic Question Decomposition), or fed as features to a small classical model trained on your labels (LLM-Generated Features for Classical ML).
- Classification at scale: tickets, emails, leads, posts, ads
- Guardrails and checks inside AI Agents: is this relevant, is the agent stuck, which tool, how much reasoning effort
- Asking many small questions about the same input (ten yes/no checks per paragraph for a prose linter, for instance)
- Triage before an LLM: let the DM filter, and only send the hard cases to the expensive model
Watch out for
- No reasoning to inspect. You get a number, not a why. Bias and blind spots are harder to see, so you need MORE evals, not fewer
- Calibration isn't universal. Calibrated on the vendor's data doesn't mean calibrated on yours. Measure it on your own distribution before trusting thresholds
- Wrong tool for open questions. If the right answer isn't in your list of options, the model will still confidently pick one. Pair the pick with an absolute yes/no check per option, or a "does an answer exist at all?" question (Relative vs Absolute AI Judgments)
- Vendor dependence. The best general DM today (Jev) is closed, and its terms forbid publishing benchmarks. The open alternatives are behind but closing the gap
My take: this is one of the most practical ideas in AI this year. Not because the tech is magic (it isn't), but because it makes judgment cheap enough to sprinkle everywhere. I expect many of the "small" LLM calls in my own automations to move to a decision model over time.
References
- Introducing System One Models & Jev (TypeSafe AI)
- Jev notes (Simon Willison)
- Nate B. Jones on Jev build patterns
- Jev's architecture unmasked (Archer Hume)
Related
- Decision Models (MoC)
- System One Models
- Jev
- TypeSafe AI
- Laya
- Kev
- SemIf
- Jevlike
- Large Language Models (LLMs)
- Small Language Models (SLMs)
- LLM Structured Outputs
- Model routing
- Binary Classification
- AI Agents
- System One Primitives
- Atomic Question Decomposition
- Relative vs Absolute AI Judgments
- Self-Consistency
- LLM-Generated Features for Classical ML
- Machine Native Intelligence
- AI Model Calibration
- Confidence-Gated Routing
- AI Model Cascades
- Jevons Paradox
- Zero-Shot Classification
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.