Thought Is Getting Cheaper Faster Than Transistors Did
Epoch AI has now put figures on the trend behind several of my recent posts. In The Plunging Price of Thought, published on 22 September 2026, Luke Emberson and David Roodman measure how fast the cost of a given level of AI performance is falling. Their answer is about 47% a quarter, or 13 times a year, across five benchmarks since 2023.
That's fast enough to make Moore's law look slow. We also have our own numbers, from a decider we built that already makes decisions at no marginal cost, and they point at the document as the thing that decides what all this cheap thinking is worth.
The Short Version
- The price of thought falls about 13 times a year. That halves the cost of a given result roughly every three months.
- Moore's law is about 1.4 times a year. In log terms the price of thought is falling six to seven times faster.
- Our local decider is already free per decision. It answers in 102 ms on a laptop. A hosted model takes about five seconds and $0.0027.
- Cheaper isn't more accurate. Accuracy dropped with the price. On whether a document was current and who owned it, embedded MX metadata closed the gap.
- Jevons multiplies the reads. When decisions are free, each document is read far more often, and each time the machine either finds a declaration or guesses.
Epoch's Numbers
Epoch took five benchmarks covering maths, hard sciences, and games of skill, and asked: what's the cheapest way, on each date, to reach a given score? OpenAI's o3 reached 75% on GPQA Diamond, a PhD-level science exam, for about 30 cents a question in January 2025. Just under 18 months later, GPT-5.6 Luna matched that score for $0.0004 a question. That's a 725-fold drop.
Averaged over the five benchmarks, the fall is 47% a quarter. It isn't even. A level of performance gets cheaper fastest right after it first appears, at 66% a quarter, and slows to 32% two years later. The authors suggest why: the maker of a new leading model can charge a premium for a while, until competitors catch up.
They list the limits themselves. The window is about three years. A benchmark score stands in for useful work, and a model can be trained towards a known test. The figure assumes a user who always switches to the cheapest capable model, which real users don't do. They call their own cross-technology comparisons "apples and oranges", because the thing being priced changes as it goes.
Against Moore's Law
Gordon Moore's 1975 observation was that the number of transistors on a chip doubles about every two years. As an annual rate that's about 1.4 times. Epoch uses a price series for compute instead, which fell 1.51 times a year between 1940 and 2001.
| Trend | Per year | Cost of a given result halves every | Over three years |
|---|---|---|---|
| Moore's law, two-year doubling | about 1.4 times | 24 months | about 2.8 times |
| Price of compute, 1940 to 2001 (Epoch) | 1.51 times | about 20 months | about 3.4 times |
| Price of thought since 2023 (Epoch) | about 13 times | about 3 months | about 2,200 times |
Measured in log points, which is the fair way to compare growth rates, 13 times a year is about six times the compute rate and about seven times Moore's two-year doubling. Epoch puts it at four times DNA sequencing, 18 times lithium batteries and 54 times electricity.
The comparison has a catch. Moore's law describes hardware. The price of thought depends on that hardware and adds everything else: better algorithms, smaller models distilled from larger ones, choosing how hard a model thinks, and margins squeezed by competition. Most of the gap between the two rates comes from software and competition, with faster chips as only one part.
Our Decider Is Already Below the Curve
Epoch measures what hosted models charge. We've built a Jev-style decider on a laptop, modelled on TypeSafe's Jev, and measured it across 100 hand-labelled decisions.
| Lane | Accuracy | Median time a decision | Cost a decision |
|---|---|---|---|
| Local decider, 3B model | 72% | 102 ms | nothing |
| Local decider, 8B model | 87% | 184 ms | nothing |
| Hosted model (Claude haiku) | 97% | about 5 seconds | about $0.0027 |
If the hosted price follows Epoch's average, that $0.0027 becomes about $0.0002 in a year. The local decider already costs nothing extra, on hardware we own, and about fifty times quicker. For small, bounded decisions (is this current, who owns it, may I quote it) the price of thought has already reached the floor.
The same table shows the part a price index can't. Epoch holds the score fixed and watches the price fall. We held the questions fixed, and accuracy dropped with the price: 97%, 87%, 72%.
The Jevons Bill
The Jevons paradox says that when a resource gets cheaper to use, total use tends to rise rather than fall. Moore's law is the best-known case in computing. Transistors fell in price for fifty years, and total spending on chips, and the power they draw, kept rising, because software used every gain. I've traced the same effect in AI energy use: Google's carbon footprint rose 48% since 2019 despite years of per-query efficiency gains.
TypeSafe named Jev after Jevons, and bet on the paradox. If the price of thought falls seven times faster than transistors did, and a local decision costs nothing, the number of machine decisions made about each document will grow more quickly than for any technology before it. A page once opened by a crawler and an answer engine will also be opened by a router and by small models asking if it's current, who owns it, and whether it may be quoted, many times a day.
Here's the bill. On PDFs that declare nothing, our 3B decider judged whether a document was current correctly 20% of the time, and named an owner nine times out of ten for files that name nobody. On the same PDFs with an MX packet embedded, it scored 100% on both. Only the document changed.
Whether those extra reads come back right depends on the document. MX, short for machine experience, is the practice of making a file explain itself to an agent that meets it in isolation. Where a document declares nothing, each cheap read is a guess, made at a volume nobody can check by hand. Where it carries MX metadata, the same reads return the declared answer. A decider is only as good as its state, and the document supplies it through what it declares.
There's a cheaper step still. Many of those questions need no model at all once the document declares the answer. Checking whether an expiry date has passed is a line of code, not an inference, so every fact a document declares is one decision nobody pays for, in money or in energy.
What I'm Not Claiming
Epoch measures price, not volume. That use will outgrow the fall in price is my inference from the Jevons pattern and the energy figures, not something their report measured. Our decider figures come from small, hand-labelled sets, so treat them as indicative; an independent re-run is still outstanding. Epoch's per-quarter rates and our per-decision costs measure different things, and I've set them side by side, not merged them.
Epoch will keep measuring the price. What you can change is what your pages declare when those decisions arrive: a date, a status, an owner, and what a machine may quote.
A declared interest. CogNovaMX sells the discipline this post recommends. The price figures are Epoch AI's, linked below, and I have no relationship with Epoch AI or TypeSafe.
Sources
- Epoch AI, 22 September 2026: The Plunging Price of Thought, by Luke Emberson and David Roodman.
- TypeSafe AI, 15 September 2026: Introducing System One Models and Jev, by Diogo Almeida.
Related
- I Built a Decider and Measured It: the local rebuild, and what embedded MX metadata did to its accuracy.
- A Decider Is Only as Good as Its State: what a decision model is, and why its calibration stops at the state it's handed.
- Nobody Got Fired for Using GEO: why cheap decisions multiply, and why MX is worth promoting before the market asks.
- Reduce Inference, Save Energy: the Jevons paradox applied to the energy AI inference uses.
- Humans Take Journeys, Agents Take One Fetch: why most agents read one file in isolation and decide, and why MX designs for both a drip feed and everything at once.
- Why Use MX: why a file needs metadata that travels with it, so the machine that meets it next doesn't have to guess.
Where Your Own Content Stands
Every machine that reads your pages will soon read them far more often. Whether they find a declaration or make a guess is what an MX audit measures.