Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Open-Weights Releases

How fast is AI inference getting cheaper? The 13x-per-year answer

Fixed-capability AI inference cost falls about 47% per quarter, roughly 13x per year, per Epoch AI. Here is why your per-task bill can still rise anyway.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single GPU server rack in a datacenter aisle under cool light, one amber status light standing out, standing in for the fixed-capability cost of AI inference falling every year the same hardware runs. Illustration: DeAI
A single GPU server rack in a datacenter aisle under cool light, one amber status light standing out, standing in for the fixed-capability cost of AI inference falling every year the same hardware runs. Illustration: DeAI

AI inference is getting cheaper at roughly 13x per year for any fixed level of capability, according to Epoch AI, which measures a 47% per-quarter decline in the cost of reaching a given benchmark score. The catch: your bill is per task, and new models spend far more tokens per task, so spending can rise even as prices collapse.

Key facts

  • 47% per quarter, about 13x per year — Epoch AI's estimate of the fixed-capability price decline across five benchmarks (math, science, and games of skill) from 2023 to 2026.
  • 725x in under 18 months — Epoch's example: scoring 75% on GPQA Diamond cost about $0.30 per question on OpenAI's o3 in January 2025 and about $0.0004 per question on GPT-5.6 Luna, which the authors compare to a car's sticker price falling from $50,000 to $69.
  • 5x to 10x per year, with about 3x from algorithms — a separate, peer-style analysis by Gundlach, Lynch, Mertens and Thompson (MIT FutureTech), whose abstract also finds the price of running frontier models rising 3x to 18x per year as models grow and reasoning lengthens.
  • 3x per task in eight months — Epoch's worked example: reaching roughly 27% accuracy on FrontierMath needed about 43 million output tokens from o4-mini in April 2025 but about 5 million from GPT-5.2 in December 2025, a cost drop the author puts near 3x over eight months.
  • 10x per year, 2021–2024 — a16z's earlier estimate that per-token prices for models clearing a fixed capability threshold fell tenfold annually, which Epoch cites as external validation of its longer-run extrapolation.

What the 13x number actually measures

The decline rate is not about any single model getting a discount. It is the price of buying a fixed score. Epoch AI assembled a dataset spanning three years and five benchmarks — GPQA Diamond, AIME (OTIS Mock), Chess Puzzles, FrontierMath Tiers 1-3, and Mystery Game Puzzles — and for each benchmark traced the cheapest way available, at each point in time, to hit a target score. The price of that frontier is what falls 47% per quarter. On math benchmarks the drop runs faster, at 50-52% per quarter (16-19x per year); on game-based puzzles it runs slower, at 39-43% per quarter.

Method matters for how much weight the number carries. Rather than running every model at every reasoning level, Epoch used a transcript-based approach developed by the federal Center for AI Standards and Innovation: take one benchmark run under an unrestricted budget, then predict performance under tighter token budgets from the transcript itself. Where the team priced open-weight models directly, it used the cost of rented hardware, and on the five models it could check against an API, the discrepancy stayed under 30%. Data and code are public on GitHub, so the analysis is checkable in a way most vendor pricing claims are not.

The comparison set gives the number its punch. Epoch puts the decline at four times faster than DNA sequencing, six times faster than compute, 18 times faster than lithium batteries, and 54 times faster than electricity's century-long fall up to 1973. No other general-purpose technology on record got this cheap this fast.

The three-way decomposition

Why is the curve this steep? The Gundlach et al. paper splits the decline into three sources: economic forces, hardware efficiency, and algorithmic efficiency. Hardware contributes a known rate — cost per FLOP falls with each GPU generation. Competition contributes another chunk: providers undercut each other on price for models at the same capability level.

To isolate the algorithmic component, the researchers looked only at open models, which strips out most competition effects, and then divided out hardware price declines. The estimate that survives: algorithmic efficiency improves around 3x per year. The same model or a smaller one does the same work for a third of the compute, every year, independent of chip roadmaps or price wars.

For builders, the 3x is the load-bearing number. Hardware savings and price wars are outside your control and partly cyclical, but algorithmic efficiency compounds into the models themselves. It is the reason a model you can run on a single GPU today scores like the cluster-scale frontier release of two years ago, and it is what makes the open-weight menu genuinely different from a price list of identical closed APIs.

Why your bill can rise while prices fall 13x

Here is the tension that makes this story confusing in practice. The 13x curve prices a fixed score. Nobody buys a fixed score. You buy completed tasks, and the amount of computation per task has been exploding.

The same research that documents 5-10x annual declines also finds the price of running frontier models rising between 3x and 18x per year, because models are bigger and reasoning models generate far more tokens per request. A query that once produced a 200-token answer now produces thousands of reasoning tokens plus tool calls. Multiply a rapidly falling price per token by a rapidly rising token count and the sign of the product is an empirical question, task by task.

Epoch's own worked example shows both curves at once: on FrontierMath, the token count needed for the same accuracy fell roughly 8x in eight months while per-token prices also fell, netting out to about 3x per task. That is a fast decline, and still far short of 13x, because the task's compute appetite grew along the way. A task that cost a cent in 2024 can genuinely cost a dollar in 2026 on a frontier reasoning model, and the 13x headline is not wrong.

The practical reading: whether your costs fall depends on whether your tasks are getting more compute-hungry faster than prices fall. Summarization, classification, and extraction get cheaper almost automatically. Open-ended agentic work is the case where per-task spending can rise for years. Teams that care should track cost per completed task, not cost per million tokens, and our monthly price snapshots exist to make that trackable.

Common misconceptions

"Cheaper inference means AI is getting less expensive to use." Not necessarily. The price of a unit of capability is collapsing; the number of units a typical task consumes is growing. If your workflow adopted reasoning models in the last year, your per-task cost may have gone up even as the per-token price fell. The honest metric is spend per completed task, and it needs measuring on your own traffic.

"The decline rate applies to the newest frontier model." It mostly does not. Epoch finds the steepest drops right after a capability level is first reached, when competition and distillation attack the premium fastest — 66% per quarter initially, settling near 32% per quarter (about 4.7x per year) two years later. Frontier launch pricing is precisely where the discount has not arrived yet. The pattern rewards a specific habit: anything that must run at volume, but does not need the frontier, belongs on a model one or two generations back, where the per-million-token prices already reflect most of the curve.

"This rate is guaranteed to continue." Both teams flag reasons it may not. The measured window is barely three years. Benchmarks over-reward distilled models, which are known to be more brittle than their teachers, so some of the measured decline may not survive contact with real workloads. There is likely a floor below which a model is simply too small for general agentic capability, no matter how good the distillation. And Epoch's own extrapolation before 2023 leans partly on the a16z dataset rather than direct measurement.

What the open-weight angle changes

A 13x annual decline in fixed-capability cost has a specific consequence for where models run. If the frontier's price advantage erodes within months of each release, the window in which a closed frontier API is the only way to buy a capability shrinks too. The capability gap between frontier and open weights closes on a schedule set partly by this curve, because distillation and open-weight retraining are the mechanism behind much of it.

That does not make open weights the right choice for every workload. It means the decision has to be re-litigated on a timescale of months, not years. A team that evaluated open models last quarter evaluated a different frontier; the price gap arithmetic moves underneath both options. The practical implication for builders is less about ideology and more about review cadence: re-test cheap candidates against your evals quarterly, keep routing simple, and use the OpenAI-compatible base-URL swap so a model change is a config change rather than a migration.

Current state (September 2026)

The two best public estimates bracket the decline rate: Epoch AI at about 13x per year on five benchmarks since 2023, Gundlach et al. at 5x to 10x per year on knowledge, reasoning, math, and software-engineering benchmarks, with roughly 3x per year attributed to algorithms alone. Both teams flag the same caveats — short windows, benchmark-heavy evidence, prices that vary enormously by domain and performance level. Epoch has since added a caution that the "price of thought" is a more ambiguous concept than the price of electricity, which complicates cross-technology comparisons. Watch the October price index snapshot for where per-token prices sit this quarter, and treat any single-provider claim about cost curves as marketing until it publishes its data.

Related reading

The September 2026 price index is the live sibling of this page — where the per-token numbers sit right now. Self-hosting versus API cost math covers the fixed-versus-variable tradeoff the 13x curve keeps repricing. The cheapest LLM APIs prices fourteen providers per million tokens, and what makes an OpenAI-compatible API the cheap way to act on any of it.

Questions

how fast is ai inference getting cheaper?
Epoch AI estimates the cost of reaching a given performance level on five benchmarks fell about 47% per quarter (roughly 13x per year) between 2023 and 2026. The Gundlach et al. paper estimates 5x to 10x per year on knowledge, reasoning, math, and coding benchmarks. Both are analyses, not measurements of any one provider's prices.
How much of the AI cost decline comes from algorithms rather than hardware?
The Gundlach et al. paper estimates algorithmic efficiency contributes roughly 3x per year. The researchers isolated open models to control for competition effects and divided out hardware price declines. The rest of the decline comes from cheaper hardware and market pressure.
If AI inference is getting 13x cheaper per year, why is my bill going up?
The 13x figure is per unit of benchmark performance, not per task. New models solve tasks with far more generated tokens, so a task that once cost $0.01 can cost $1.00 on a reasoning model even as the price per capability point collapses. Per-task spending can rise while per-performance pricing falls.
What is the cheapest way to buy AI capability?
On published benchmarks, the cheapest way to reach a fixed score is usually a heavily distilled open-weight model rather than the newest frontier release. Distilled models score disproportionately well on benchmarks for their size, so test cheaper candidates against your own evals before paying frontier rates.
Will AI inference prices keep falling 13x per year?
No one can promise that. The decline measured from 2023 onward is already slowing: Epoch AI finds first-year cost drops of about 75x per year for newly reached performance levels, settling near 4.7x per year two years later. Whether the curve continues depends on algorithmic progress, hardware cycles, and competition.

Sources

  1. The plunging price of thought (Epoch AI) — Epoch AI
  2. The Price of Progress: Price Performance and the Future of AI (arXiv:2511.23455) — arXiv / Gundlach, Lynch, Mertens, Thompson
  3. LLM inference prices have fallen rapidly but unequally across tasks (Epoch AI) — Epoch AI
  4. How persistent is the inference cost burden? (Epoch AI Gradient Updates) — Epoch AI
  5. LLMflation: LLM inference cost is falling fast (a16z) — Andreessen Horowitz
  6. Open-Model Inference Prices, September 2026 — DeAI News

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →