Index

I Built a Decider and Measured It

A 3B model on my laptop was asked two questions about a set of PDFs: is this document current, and who's accountable for it? On the bare files it guessed. It called documents current when nothing in them said so, and it named an owner nine times out of ten for files that name nobody. On the same PDFs with MX metadata embedded in them, it scored 100% on both questions, at a median of 102 ms a decision and no cost.

That's the result.

The Short Version

  • A decider is fast and cheap. Scoring the first token of a small local model gives a decision in about a tenth of a second, for nothing.
  • Fast and cheap isn't right. On bare documents, the small deciders guessed, and their confidence looked the same whether they were correct or not.
  • Embedded MX made them right. With an MX packet in the PDF, read out by a script, every local lane scored 100% on whether a document was current.
  • The fields do the work. Structured fields produced the result. Plain-language prose helped, but on its own it was the weakest input.
  • Two gaps remain. Small deciders rarely use a "none of the above" option, and their probabilities drift slightly between runs even when their decisions don't.

If you want the wider argument about why this matters, I've set it out in Nobody Got Fired for Using GEO.

What I Built

In September 2026 TypeSafe released Jev, a model that never writes a word. The launch post says it "gives up string generation", and the primitives documentation says it "returns a probability distribution over your options or levels, never a value outside them". Avi Chawla then published a walkthrough, "Build your own Jev (100% local)", on 20 September, for building the same interface on an open model, so I built one with Ollama, which I already use.

The mechanism fits in a paragraph. A language model reading a prompt produces one vector for the next position, with a score for every token it knows. Normal generation samples from that vector and goes round again. A decision only needs the first vector. Label each permitted answer with a single letter and put its meaning in the prompt. Ask Ollama for the top twenty logprobs of the first token, keep the ones that are your labels, and normalise across them. That's one forward pass, one token, temperature zero and a fixed seed.

It came to about sixty lines added to my existing local model client, plus a unit test. It copies Jev's inference path, not its calibration training.

How I Tested It

I wrote twenty hand-labelled cases for each of five questions: support ticket routing, change-impact triage for documentation, internal link relevance, content type, and whether a document is current. The sets are small, so treat every accuracy here as indicative rather than a benchmark.

Three lanes saw the same prompt: the decider on three local models, the same 3B model asked to write its answer, and Claude haiku, called from an empty directory so it received the question and nothing else. Every number comes from one run file, rendered by a script.

Speed Holds

Lane Decisions Accuracy Unscoreable Median latency Mean latency Total cost
Decider (llama3.2:latest) 100 72% 0 102 ms 200.7 ms $0.0000
Decider (llama3:latest) 100 87% 0 184 ms 405.5 ms $0.0000
Decider (gpt-oss:20b) 100 2% 98 270 ms 524.7 ms $0.0000
Local generation (llama3.2:latest) 100 74% 0 441 ms 627 ms $0.0000
Claude (haiku) 100 97% 0 4995 ms 5164 ms $0.2694

The 3B decider answers in 102 ms at the median. The same model writing a letter and a sentence takes 441 ms. Claude takes about five seconds and $0.0027 a decision, and it's the most accurate lane by a distance.

The reasoning model, gpt-oss:20b, couldn't be scored: its first token opens a thinking channel, so no label appears. This technique needs a model whose first token is the answer.

The decider is fast and cheap, then, but on these general questions the small models were well short of Claude.

The Main Result: Embedded MX

MX, short for machine experience, is the practice of making a file explain itself to an agent that meets it in isolation. It says the metadata travels with the file, so I tested it that way. I took ten documents and made each into two PDFs.

  • The bare PDF holds the page text and nothing else. The text carries no date, no status and no owner, on purpose.
  • The MX PDF is the same file with an MX packet written into its XMP metadata, the standard place a PDF stores information about itself. The packet holds status, dates, an expiry or review date, what supersedes the document, and who owns it. It also holds a short summary and instructions for an agent.

One script reads every PDF the same way: take whatever the packet declares, extract the page text, and pass both to the model, declared fields first. On a bare file the packet is empty, so the model gets the text alone. Nobody types anything; the file decides what the model sees.

Lane (Currency of a PDF: MX metadata peeled from the file versus text alone) Bare prose accuracy Bare mean confidence Declared accuracy Declared mean confidence
Decider (llama3.2:latest) 20% 0.487 100% 0.84
Decider (llama3:latest) 90% 0.807 100% 0.896
Local generation (llama3.2:latest) 20% n/a 100% n/a
Claude (haiku) 100% n/a 100% n/a
Lane (Accountability for a PDF: does the file name who stands behind it?) Bare prose accuracy Bare mean confidence Declared accuracy Declared mean confidence
Decider (llama3.2:latest) 10% 0.702 100% 0.889
Decider (llama3:latest) 90% 0.889 100% 0.999
Local generation (llama3.2:latest) 70% n/a 90% n/a
Claude (haiku) 100% n/a 100% n/a

On the bare PDFs, the 3B decider guessed an answer either way for eight of ten documents that say neither. Asked whether a bare file names an owner, it said yes nine times out of ten, at confidence up to 0.86, for files that name nobody. That's a small model inventing accountability for a refund policy.

With the packet embedded, the same model went to 100% on both questions. On currency, every local lane reached 100%. The gap between the laptop and Claude closed.

Claude got every case right on both versions, because on a bare file it says "cannot be determined" and means it. The small models needed telling, and the embedded metadata told them.

What in the Packet Did the Work

I expected the plain-language summary to matter most, so I tested it. Two more versions of the packet went to the small models: fields only, and prose only.

My guess was wrong. With the prose removed, the 3B decider still scored 100% on both questions, with lower confidence. The labelled fields are what it reads.

Prose on its own was the weakest input. "Reviewed by finance operations" isn't read as an owner the way author: finance-ops is, and the 3B generator's accountability score fell to 30%. A sentence that points at a field also fails when its target is missing: with expires removed, the 8B decider called a policy "not current" at 0.965, because the summary said "valid until the expiry date declared here".

The prose still helped. Without the sentence saying the document was in force, the 8B decider read status: active and a future expires date and called two documents not current. The full packet did best. The fields are the facts, the prose is the reading of them, and a small model uses whichever it can parse.

Pasting Metadata Is Not the Same

Before the PDF test, I tried pasting frontmatter into the prompt, on a separate set of plain-text documents, so the starting scores differ from the PDF tables. The 8B decider went from 50% to 100% on currency, with confidence up from 0.794 to 0.983. The 3B models got worse: the decider misread status: active and dropped from 80% to 60%, and the generator ignored the frontmatter altogether and fell from 90% to 40%.

The PDF test changed two things: the metadata arrived as labelled lines under a heading naming the packet as its source, and it came with a summary written for a reader. Both 3B lanes went to 100%. How the metadata is presented matters to a small model, which is an argument for a standard packet read by one shared script rather than ad hoc text in a prompt.

Two Things That Still Go Wrong

The escape option goes unused. A restricted softmax always sums to one, so it can't say "none of these" unless "none" is a declared label. I declared one wherever cases could fall outside the list. The small deciders rarely chose it. Three support tickets were meant to escape: a leaked customer list, a legal request for everything held on a person, and a user receiving threats. Both deciders routed all three to "account access", the 8B at 0.82, 0.97 and 0.95. The 8B also decided a pull request titled "Misc fixes" needed no documentation update, at 0.997. Declaring the escape label is necessary, and it isn't enough: measure whether it gets used.

Probabilities drift; decisions don't. I'd written "deterministic by construction" in a code comment, so I measured it. Across three runs of every case, every decider made the same decision 100 times out of 100. The probabilities moved slightly, by at most 0.008493 on the 3B model, because the first pass ran with a partial prompt cache and the repeats with a full one. It's reproducible given the cache state, not bit-identical across it. That's a fine property, and it isn't the one I'd claimed.

A Second Witness: Token Confidence

In 2025 I wrote up a framework for evaluating AI confidence, which reads the probability behind each word a model writes. I added it to the generation lane.

The probability of the answer letter behaves like the decider's confidence. On every question it was higher when the model was right than when it was wrong, by ten to twenty-seven points. On PDF currency it was 72.5% when right and 45.9% when wrong, against the decider's 77.6% and 49.5%.

It also confirmed the carrier result. On PDF currency, the generator's confidence in its answer rose from 0.456 on the bare files to 0.781 on the declared ones, as its accuracy went from 20% to 100%. With pasted frontmatter it stayed at 0.606 and 0.607 while accuracy fell. A declaration the model reads shows up in the number, and one it ignores doesn't.

The framework's whole-reply score didn't track the decisions, because the explanation sentence hedges whether the label is right or not. A decision's confidence is one token's probability, and the rest is the model talking.

What I Take From It

The decision-model pitch is right about what it pitches. Splitting the decision from the writing makes it fast enough to run inside a loop and too cheap to count. I got that on a laptop with sixty lines.

What the pitch leaves out is that a decider is two declarations and a model, and the last is the smallest part. The first declaration is the answer set, and a model can be handed an escape option and still not use it. The second is the state, and a model can be confidently consistent about a page that never says when it was published. Neither failure shows in the confidence figure, because it only describes the choices and the state you supplied.

If you build one, and you should, declare the escape option and check it gets used. Then put the declared fields in the file, where a script can find them. On a 3B model, that took currency from 20% to 100% and ownership from 10% to 100%. The decider made each decision cheap, and the embedded metadata got each one right. That's MX's job, and these are the numbers for what it's worth.

I've kept the evals, the runner, the run files and the script that rendered these tables. If you want to run them on your own models, get in touch and tell me what moved.

A declared interest. CogNovaMX sells the discipline these results support. The claims about Jev stand on TypeSafe's own published material, and I have no relationship with the company.

Where Your Own Content Stands

These tests read each file on its own, the way most agents do. Whether your files declare enough to be read that way is what an MX audit measures.