Index

Nobody Got Fired for Using GEO

In September 2026 I spent a week in Mallorca. On the Wednesday I gave a lightning talk on reverse centaurs at Boye & Co's CMS Experts day. Later that day, Nina Kozłowska asked what each reader needs when you write for humans and machines, and Deniz Ergun of Bloomreach wanted to know what happens if the page isn't the product. Then came CMS Camp itself. My thanks to Janus Boye for running the Experts day and to Volker Graubaum for running CMS Camp. I learnt a lot, but the lesson I keep coming back to came from the Friday masterclass rather than my own session. Architectural Trends in the CMS Space was run by William Borg Barthet, Professional Services Director for Bloomreach's CMS. It examined "the assumptions and architectural fashions that have shaped the industry", and it asked a question I haven't stopped thinking about: "How do we choose the right architecture for the problem, rather than adopting the architecture that happens to be fashionable?"

His subject was CMS architecture, from monoliths through headless to MACH. He also raised the possibility that CMS platforms will serve "not only websites and applications, but also machine and agent consumers", which is the audience MX serves. MX, short for machine experience, is to machines what UX is to users: it makes content something an agent can understand and act on without guessing. My takeaway reached beyond architecture: the marketing industry follows fashion too.

That isn't an insult. Fashion is how a crowded market handles risk. When no one can prove which option is right, you choose the one that can't be held against you. For a generation, that was the joke about IBM: nobody got fired for buying IBM. The purchase had to be defensible rather than the best.

The GEO Phase

In 2026 we're in the same phase, with a new name on the badge. Nobody gets fired for using GEO, generative engine optimisation: tuning content so that AI answer engines cite it.

I've written before that GEO is a tactic and MX is the specification, and that Google named GEO, then debunked it. I won't repeat those arguments. The point from Mallorca is different. GEO's popularity isn't evidence that it works. It's evidence that it's safe to buy. Everyone is following the trend, and in 2026 the market hasn't worked out what comes after it.

Jonathan Healey of IDHL put evidence behind this at the Experts day. He revisited eight predictions he'd made about where CMS was heading. One had landed at pace: his slide said "Every major CMS now ships with GEO and schema in mind". The next slide cited an Ahrefs controlled experiment in which 1,885 pages added JSON-LD, and "the citation lift was statistically indistinguishable from zero". Where the prediction fell short, his verdict was blunt: "Most CMSs model navigation. Almost none model meaning."

That gap is where MX sits. Markup added to chase citations is fashion. Metadata that declares what a document is, who stands behind it and whether it's still current is meaning, and it's the part I've been able to measure.

Fashion has a cost beyond the budget. In my talk I used Cory Doctorow's definition of a reverse centaur: a person forced by technology to work at an inhuman pace. The machine sets the pace and the human does the labour. A content team chasing whatever the answer engines reward this quarter is heading that way. It rewrites pages to suit a system it can't see, measures success by citations it can't control, and starts again when the rules change. Doctorow's governing question applies: the important thing about a technology "isn't what it does, but who it does it for and who it does it to". GEO, as usually sold, is done to the content desk.

What Is Left When the Bubble Bursts

I think we have to get through the bubble bursting before the next paradigm arrives. Doctorow predicted what survives. In his 2023 Locus essay, What Kind of Bubble is AI?, he wrote that "there will be little models - Hugging Face, Llama, etc - that run on commodity hardware", and that people would keep improving them, "possibly enough to satisfy most of those low-stakes, low-dollar applications".

His prediction is more cautious than mine, so I'll be clear about where we differ. He also wrote: "But these little models were spun out of big models". He doubted the big models would survive without bubble money. I expect them to survive in a smaller role. Many small, targeted models do the everyday work, each answering one narrow kind of question, and they hand the hard cases up to larger models. The small ones decide and the large ones write.

Models built this way already exist. TypeSafe released Jev in September 2026. It's a model that never writes text: it answers a typed question with a probability over a declared set of answers. I rebuilt the same interface on a local model and measured it. On a 3B model, a small AI with three billion parameters that runs on an ordinary laptop, a decision took a median of 102 ms and cost nothing. A hosted model, a large AI rented from a provider over the internet, took about five seconds per call.

What Solved It: A Jev Decider Plus Embedded MX

My working assumption going into those tests was that a targeted model hallucinates less because it's targeted. A model that can only answer "current", "stale" or "none" can't invent a paragraph.

That turned out to be half the answer. Targeting stops the model inventing text, but it doesn't stop it guessing. Handed PDFs with nothing declared, the 3B decider called documents current when it had no way to know, and invented an owner nine times out of ten. It couldn't write a false sentence, so it picked a wrong label instead, with confidence.

The other half was MX metadata embedded in the file. I made two versions of the same PDFs: one bare, and one with an MX packet in its XMP metadata, the standard place a PDF stores information about itself. One script peeled whatever each file declared and handed it to the decider. Nothing else changed.

With the packet embedded, the problems went away:

  • On whether a document was current, the 3B decider went from 20% to 100%.
  • On who owns a document, it went from 10% to 100%, and stopped inventing owners.
  • On currency, every lane I tested scored 100% on the declared files, including the same 3B model writing its answers and a hosted one.

Neither piece did it alone. The Jev-style decider made each decision fast and cheap: a median of 102 ms on a laptop, at no cost. The embedded metadata made each decision right. Together, a 3B model on a laptop matched a hosted service on these questions. The test set is small and hand-labelled, so treat the figures as indicative rather than a benchmark, but the direction was the same on every question I asked.

Targeting narrows the question. Only the document can supply the answer. A small model has no brute-force fallback: it can't read around the gap the way a frontier system with a huge context window sometimes can. That makes declared metadata worth more to small models than to large ones, and they're the ones I expect to do most of the work.

Prose on its own is the weakest input. "Reviewed by finance operations" isn't read as an owner the way a labelled field naming the owning team, author: finance-ops, is. With the labelled fields alone, the 3B decider still scored 100%. With the prose alone, the same 3B model writing its answers fell to 30% on ownership. A fact a machine must act on belongs in a field, not a sentence.

The Jevons Part

TypeSafe named Jev after William Stanley Jevons, whose paradox says that making a resource cheaper to use tends to increase total consumption rather than reduce it. I've traced the same effect in AI energy use. The launch post says so directly: the company expects machine intelligence to follow the path coal took after steam engines got more efficient. The company is betting on the paradox rather than hedging against it. Cheap decisions will multiply until the total goes up.

Follow that through to a document. Under GEO, a page might be read by a crawler and summarised by an answer engine. After the shift, the same page is read by a router deciding which model should handle it, a decider checking whether it's current, another asking who owns it, and a third deciding if it may be quoted. Each read is cheap, so there are many more of them. Every one of them reads what the document declares, and guesses where it declares nothing.

Calibration, a model being honest about how sure it is, doesn't save you here. A decider can be honest about its confidence in the state it was handed and still be wrong about the world, because the input left something out. That gap widens with volume. At a few decisions a day a human can catch the bad ones. At millions a day, the router's log is the only decision record anyone has, and only what the documents declared can make that log trustworthy. A 10% error rate is a curiosity at ten decisions a day. At a million a day, it's a hundred thousand wrong answers given in your name.

Put the two results together. Jev makes decisions cheap enough to multiply, and embedded MX makes each of them right. Without the metadata, the Jevons effect multiplies guesses. With it, the effect multiplies correct answers.

There's a cheaper step still. Many of those questions need no model at all once the document declares the answer. Checking whether an expiry date has passed is a line of code, not an inference, so every fact a document declares is one decision nobody pays for, in money or in energy.

That's the paradigm I think we're heading for: targeted small models coordinating with the larger ones, all reading the same documents at a volume nobody can supervise by hand. That's when the world will be ready for MX in every document.

What This Means for Your Business

Take a document every business has: a refund policy. An agent answering a customer reads it and has to settle two things. Is this the version in force, and who stands behind it? If the file doesn't say, the agent guesses. In my tests, a small model named an owner for files that named nobody nine times out of ten, and did it with confidence. In practice, that's a customer quoted an expired policy, or a partner told the wrong team approved a price, and nobody able to show afterwards why the machine said it.

The fix isn't a bigger model or a new platform. It's declaring a few facts in the documents machines already read about you. Start with the ones that carry money or liability: policies, prices, terms and product information. For each one, record three things as labelled fields: who owns it, whether it's current, and the date it stops being valid.

That work pays twice. Your own AI tools stop guessing about your own documents, which is the result the tests measured. And every declared fact becomes a dated record you can show an auditor, a partner or a regulator when they ask how a machine reached its answer.

Why We Promote MX Before the Market Asks

The MX Manifesto opens with "MX is to machines what UX is to users", and argues for determinism over inference: the same input gives the same result every time. The Jev tests show what that means in practice. A declared fact is a decision made with certainty. An undeclared one is a guess, made at speed and repeated at scale.

As I write, in September 2026, the market hasn't caught on to this. It's still in the GEO phase, buying the product that gets nobody fired. The sensible timing for MX is after the shift, when targeted models make the need obvious.

We're promoting it ahead of the market anyway, for two reasons.

The first is that standards compound. The manifesto calls this the architecture of participation. An organisation that declares its own documents, to make them more useful to its own machines, also produces something every other machine can read. That value builds over time, and it needs a head start.

The second is fashion itself. When the market moves on from GEO, it'll look for the next safe purchase. I'd like MX to be the known answer by then, with the evidence published and the records in place.

That answer won't be something you buy. MX is a practice, as UX is (I set out the case in Why Use MX), and teams build it into how they work. Declaring what each document is, who stands behind it and whether it's current becomes a habit of its authors, and no vendor can take that away.

Nobody got fired for buying IBM. I'd like the next version of that line to be about MX.

A declared interest. CogNovaMX sells the discipline this post recommends. The measured figures come from my own tests, published in full in the decider post, and I have no relationship with TypeSafe.