SLMs caught up
Today's 4–8B models match flagship quality from ~18 months ago on most everyday tasks. Right-sizing alone cuts energy and cost by 10×+ with no felt loss.
The rest of this collection looks at the hard limits — soaring demand, finite power, carefully shared compute. All real. But it's only half the picture. The other half is genuinely exciting: AI is getting cheaper and more efficient faster than almost any technology in history. Both things are happening at once.
For a model of equivalent quality, inference cost has dropped roughly 10× every year for three years running — faster than compute in the PC era or bandwidth in the dotcom boom (a16z).
Google reports a 33× year-over-year drop in energy for the median text prompt — from model, serving, and hardware efficiency stacked together.
The same intelligence that cost dollars a year or two ago now costs fractions of a cent. The frontier price stays high because the frontier keeps moving — but any fixed capability gets cheap, fast.
Today's 4–8B models match flagship quality from ~18 months ago on most everyday tasks. Right-sizing alone cuts energy and cost by 10×+ with no felt loss.
Quantization, distillation, mixture-of-experts, speculative decoding, better attention — each shaves cost per token without touching the chip. These compound.
Newer accelerators and tuned serving stacks push far more tokens per watt and per dollar — the same query lands on steadily more efficient silicon.
Open models (Llama) and aggressive challengers (DeepSeek) drag the whole market's price-per-token down — and self-hosting turns marginal cost into just electricity.
Sticker prices on the newest flagship can rise even as the underlying trend plummets — because the product keeps getting better. Hold capability fixed and the line goes down and to the right, steeply. That's the optimistic case: efficiency is not a footnote to the demand story, it's the main event running alongside it.
Cheaper tokens don't shrink total energy use — they invite vastly more usage. A 33×-cheaper prompt times a 330×-bigger token volume still grows the grid bill. This is exactly why demand can stay vertical while supply is physical even as each query gets greener. And per Epoch AI, the declines are uneven — dramatic for solved tasks, slower at the moving frontier. Optimism about the rate, realism about the totals.
Sources — Andreessen Horowitz, Welcome to LLMflation (Nov 2024) — ~10× per-year inference cost decline for equivalent performance; GPT-3-quality ~1,000× cheaper than 2021. Epoch AI, LLM inference price trends (2025) — cost at fixed performance halving roughly every 2–3 months; 9×–900×/yr by milestone; acceleration to ~200×/yr under competition. Google, Measuring the Environmental Impact of Delivering AI at Google Scale (arXiv:2508.15734, Aug 2025) — 33× YoY energy reduction per median Gemini text prompt. TechTarget (Apr 2026) — small-vs-large model right-sizing.
Note — "10×/yr", "33×", and "1,000×" are headline figures from the cited analyses and vary by methodology, task, and time window. They describe a direction and rough magnitude, not a precise universal constant.