Index

A Decider Is Only as Good as Its State

In the week Jev launched, my feed was full of a ten-step guide to something called Jev, from a company called TypeSafe AI. The guide was breathless, so I went to the source. TypeSafe's founder, Diogo Almeida, published the launch post on 15 September, and the docs are open. What follows comes from those, not from the feed.

What Jev Is

Jev is what TypeSafe calls a System One model, after Kahneman's fast, intuitive mode of thinking. The company's own description is "a new class of frontier models built to make fast, structured decisions that software can use directly". It doesn't write. You hand it a state, which can be any text or JSON, plus a list of typed questions, and it answers each one.

A question takes one of a few shapes. A Choice picks one option from a list you supplied. A Score rates the state on an ordered scale you defined. A Noul returns the probability that a statement is true. Every answer arrives with a probability distribution and a single confidence number. TypeSafe's docs put the constraint plainly: "The model returns a probability distribution over your options or levels, never a value outside them." That's what the company means when it says Jev can't hallucinate: it can't invent an option, though it can still pick the wrong one.

The launch post claims the model reaches similar intelligence to leading LLMs on this class of task while running between forty and two hundred times faster, priced at $0.042 per million input tokens, in early access from day one. Those are TypeSafe's own numbers from launch week. The evaluation method, by the company's account, uses "the predictions of the largest, smartest, and most expensive external models as reference probabilities", so "as good as a frontier model" means agreement with those models, not with the world.

What's Right About It

Splitting the decision from the generation is the correct cut. An agent loop that asks a chat model to decide the next step gets a paragraph back, then has to parse a decision out of the prose. A model that only ever returns a typed value removes that whole class of failure.

Honest uncertainty is the second thing. TypeSafe's docs say: "If an intelligent system, whether human or machine, cannot express honest uncertainty, the system cannot be trusted." I have written a version of that sentence in this series more than once. A calibrated probability, where more confidence really does track more accuracy, is what a threshold needs to mean anything.

The third is how the docs tell you to use the threshold. "A confidence threshold is not one number." The bar for acting without confirmation "is higher for a destructive operation than for a read-only one", and below the bar the instruction is to "route to a human, request clarification, or fall back to a different system". That's Human-in-Command written as an inequality. The machine acts inside a declared boundary and hands the rest to a person. I'd rather see that in a vendor's docs than in a values statement.

The Shape of the Thing

The breathless guide did get one thing right. Every Jev workflow runs through the same four steps: state, questions, action, verify. It's the first step that matters here, because you supply it. You hand Jev the state, and you write the questions. Calibration only covers what happens after that: how closely the model's probabilities match the true answers, for the input you gave it. It can't vouch for the input. Feed it a thin or mistaken state and you get a well-calibrated answer about the wrong thing.

Ask "is this offer still valid?" about a web page that never says when it was published, and Jev gives you an honest probability about a document that didn't say. The number is calibrated, but what it measures is a guess. The docs even anticipate the gap in their own way, advising you to add "an other or none of the above option when the list might not cover every input". That's sound practice. It also tells you the model is only as complete as the state and the options you gave it.

This is the point I keep making in this series, and Jev sharpens it. A small model needs your metadata more than a big one does, because it can't read its way around a page that says nothing about itself. A decider that never reads around anything, that judges exactly what it's given and no more, needs the state to be honest even more.

Every MX Field Is a Question You Don't Have to Ask

MX, short for machine experience, is the practice of making a file explain itself to an agent that meets it in isolation, which is exactly how a decider works. The questions a decider asks about a document are mostly questions MX asks a publisher to answer up front. Who wrote this. When. Has it been changed since. Is it the current version. What may an agent do with it. Where does the evidence for its claims live. Each of those is a Noul, a yes-or-no with a probability attached, when a model has to work it out from the prose. Each is a plain comparison when the page declares the answer in a field a machine can read.

A signature with an expiry date answers "should I still trust this?" by comparing two dates. No model, no probability, no latency, and no bill. A declared content policy answers "may I quote this?" the same way. A declared fact is a decision at probability one, and it costs nothing to compute. The decider's job shrinks to the questions the page genuinely can't answer about itself, which is where a calibrated probability earns its keep.

The guide's closing rule was: code computes, LLMs create, Jev decides. Add a fourth: metadata declares. The fourth makes the other three safer and less expensive, and it shrinks the cheapest of them, the decision, further still. TypeSafe charges very little per token, and a question you never ask costs nothing at all.

A Threshold Is Governance Only if It Leaves a Record

Routing a low-confidence decision to a human is the right move. It's governance only if the routing leaves evidence. The state it judged, by content hash. The questions you asked. The probabilities that came back. What bar applied, and why that one for that class of action. Who confirmed, and when. A Jev answer already contains most of that in a typed, machine-readable form. Logging it takes a line of code. Not logging it turns the threshold into a private conversation between a model and a script that nobody can audit afterwards.

That's the same record MX asks a publisher's document to carry, read from the other side. The document declares its provenance, and the decider records its decision, so neither the human asked to confirm it nor the auditor who arrives later has to guess. The verify beat at the end of the loop, fresh state proving the result, is the same discipline again: measure, then say.

We keep that kind of record already, and you can open one. Every AI step in a file we produce gets written into an AI-provenance record that travels with it: which model or agent acted, what it did, who ran it, and the human check that signed the result off. In a PDF the record is embedded in the file's metadata, so it goes wherever the file goes.

We cover the decider too. Each decision goes into the same record as its own step: the state by hash, what was asked and the labels allowed, the probabilities, the bar and the kind of action it was set for, what happened next, and who confirmed it. The record also hashes the prompt and the label set, because adding one answer shifts every probability, and the thresholds you tuned on the old set stop meaning what they did. Nobody writes the verdict by hand. It's worked out from those facts, and our pre-push check recomputes it, so a record claiming a clean act over a below-bar call doesn't get through.

To see one, go to mx.allabout.network/tell. It loads our MX Explained PDF into the MX Inspector, which reads the file in your browser and walks the chain step by step: the build environment, the accountable people, and every recorded action. Drop in one of your files and it never leaves your machine. The Inspector also reports each trust signal separately (who published it, how current it is, whether the publisher still stands behind it, and what corroborates it) instead of folding them into one score.

When the Decider Becomes the Router

A second wave of posts about Jev followed the launch, and it moved on from describing the model to predicting what it does to the market. The argument I keep seeing runs like this. Most agent steps are decisions, not writing: which tool, is this done, is this safe, escalate or not, which bucket does this go in. Put a cheap, calibrated decider in front of every call and it becomes the router. It sends the easy work to a small model, strips stale context before anything reaches the expensive one, and escalates only when its own confidence says it should. Show the cost per task side by side and nobody keeps paying flagship prices for yes-or-no answers.

Most of that's right, and the waste is real. I've watched it in our own work: the strongest model gets used by default because switching is friction, not because the task needs it.

Where the argument overreaches is the conclusion that frontier spend shrinks. Cheaper decisions don't mean fewer of them. When something gets cheaper to use, people tend to use more of it. Economists call that the Jevons paradox, and it's hard to miss given the model's name. Cheaper routing makes agents cheaper to run, so more agents get built, they run longer, and the hard steps still go to the big model. The labs sell small models of their own and already route inside their products, so a lot of this moves money between one vendor's tiers rather than away from the vendor.

What does grow, and those posts skip it, is the number of automated choices nobody recorded. A router is a decider making thousands of quick calls per task: this one to the small model, that context dropped, this answer good enough not to escalate. Each is a decision about what a person eventually sees. When a routed answer turns out wrong, the questions are the same ones as before. What state did the router judge? What did it decide, and with what confidence? What bar applied, and who owns it? A router's calibration has the same limit as any decider's too. It's honest about the inputs it was trained on, and a router whose confidence drifts on unfamiliar traffic keeps routing with no alarm and a number that looks right.

The router is the decider problem at volume, and the same two answers apply. Declared metadata shrinks its work, because a page that states its date, owner, and permitted use has already answered the questions the router would otherwise pay to ask. The log makes it accountable: which model each task went to, why, what it cost, and what the router saw at the time. Our provenance record takes a router's calls the same way it takes any decider's, and rolls them up into the figures the Jevons paradox makes you need: how many decisions, the spend by model, how often it escalated, and how many calls it acted on below the bar.

What to Do

If you publish, declare. The author, the date, the currency, the permitted use, the evidence, in fields a machine reads without inference. Every one you declare is a question no decider has to ask a model about your page, which means one less guess.

If you build with a decision model, put the declared fields into the state before you ask the question, so it judges a document that says what it is. Set your thresholds by how reversible the action is, as TypeSafe's own docs tell you. Log the decision, every time, in a form a person can read back.

If you run a router, treat it as a decider too. Log where each task went and why, test its confidence on your own traffic before you trust it, and don't book the saving as a cut in total spend until you've measured the volume that follows.

I've since rebuilt a decider on a local model and measured it, with and without declared metadata. The numbers bear this out.

A declared interest. CogNovaMX sells the discipline this post recommends. The claims about Jev stand on TypeSafe's own published material, linked below, and I have no relationship with the company.

Sources

  • TypeSafe AI, 15 September 2026: Introducing System One Models and Jev, by Diogo Almeida.
  • TypeSafe AI docs: Confidence, how confidence is derived and when to escalate.
  • TypeSafe AI docs: Primitives, the Choice, Score, and Noul question types.
  • The ten-step guide that prompted this post circulated on LinkedIn, summarising a longer working note titled "Jev Engineering". It was the prompt, not a source for any claim here.

Where Your Own Content Stands

A decider judges the state it's given. Whether your pages give it an honest one is what an MX audit measures.