Counterpoint · the optimistic case · 2026

The other half
of the story.

The rest of this collection looks at the hard limits — soaring demand, finite power, carefully shared compute. All real. But it's only half the picture. The other half is genuinely exciting: AI is getting cheaper and more efficient faster than almost any technology in history. Both things are happening at once.

Cost · "LLMflation"
~10× cheaper / year

Price per token in free-fall

For a model of equivalent quality, inference cost has dropped roughly 10× every year for three years running — faster than compute in the PC era or bandwidth in the dotcom boom (a16z).

Energy · per prompt
33× less / year

Energy per query collapsing

Google reports a 33× year-over-year drop in energy for the median text prompt — from model, serving, and hardware efficiency stacked together.

How fast is "fast"?
1,000×
cheaper to run a GPT-3-quality model today than in 2021 ($/Mtok)
~2 mo
the cost to hit a fixed performance bar roughly halves this often (Epoch AI)
9–900×
per-year decline, depending on which capability milestone you measure
→200×
the median decline rate accelerated under competition (from ~50×/yr, Epoch AI)

The same intelligence that cost dollars a year or two ago now costs fractions of a cent. The frontier price stays high because the frontier keeps moving — but any fixed capability gets cheap, fast.

Why it keeps falling
Small models, big jobs

SLMs caught up

Today's 4–8B models match flagship quality from ~18 months ago on most everyday tasks. Right-sizing alone cuts energy and cost by 10×+ with no felt loss.

Algorithmic gains

More from the same FLOPs

Quantization, distillation, mixture-of-experts, speculative decoding, better attention — each shaves cost per token without touching the chip. These compound.

Hardware perf / watt

Each generation does more

Newer accelerators and tuned serving stacks push far more tokens per watt and per dollar — the same query lands on steadily more efficient silicon.

Competition

Open weights set a floor

Open models (Llama) and aggressive challengers (DeepSeek) drag the whole market's price-per-token down — and self-hosting turns marginal cost into just electricity.

The metric that's actually moving is intelligence per dollar.

Sticker prices on the newest flagship can rise even as the underlying trend plummets — because the product keeps getting better. Hold capability fixed and the line goes down and to the right, steeply. That's the optimistic case: efficiency is not a footnote to the demand story, it's the main event running alongside it.

The honest catch · keep this in mind

Efficiency gains get partly eaten by demand (Jevons' paradox).

Cheaper tokens don't shrink total energy use — they invite vastly more usage. A 33×-cheaper prompt times a 330×-bigger token volume still grows the grid bill. This is exactly why demand can stay vertical while supply is physical even as each query gets greener. And per Epoch AI, the declines are uneven — dramatic for solved tasks, slower at the moving frontier. Optimism about the rate, realism about the totals.

Sources — Andreessen Horowitz, Welcome to LLMflation (Nov 2024) — ~10× per-year inference cost decline for equivalent performance; GPT-3-quality ~1,000× cheaper than 2021. Epoch AI, LLM inference price trends (2025) — cost at fixed performance halving roughly every 2–3 months; 9×–900×/yr by milestone; acceleration to ~200×/yr under competition. Google, Measuring the Environmental Impact of Delivering AI at Google Scale (arXiv:2508.15734, Aug 2025) — 33× YoY energy reduction per median Gemini text prompt. TechTarget (Apr 2026) — small-vs-large model right-sizing.

Note — "10×/yr", "33×", and "1,000×" are headline figures from the cited analyses and vary by methodology, task, and time window. They describe a direction and rough magnitude, not a precise universal constant.