Reference Collection · Claude & the AI economy

The Token Economy:
A Field Guide

Seven short reference sheets on how AI usage actually works — who burns the most, why the build-out is so massive, whether self-hosting is greener, what you're billed on, why it's all getting cheaper anyway, whether the whole thing's a bubble, and what it costs to build in the first place. Same engine underneath; very different views depending on where you stand.

Compiled June 2026 · figures current as of publication, subject to change
A
New to AI terms?
Glossary & Plain-English Guide
Every term in the collection — tokens, inference, kWh, grid intensity — explained simply, no jargon assumed.
Start here
The collection · 7 sheets
01
Scale · Per user vs. per company

Who Actually Burns the Tokens?

A person typing into a chatbot moves tokens at human speed; a company running APIs and agents moves them at machine speed, around the clock. Plotted on a log scale — each gridline 1,000× the last — the gap between an individual and a single enterprise runs to roughly a million-fold. The upside for you: that machine-scale demand funds the infrastructure and price drops behind your cheap personal plan.

~1M/yr individual → 38Q/yr platform ~1,000,000× gap log scale
Read the sheet
02
Physics · Demand vs. supply

Demand Is Vertical. Supply Is Physical.

Every token runs on a real chip drawing real power — and demand is racing so far ahead of supply that it's driving the biggest infrastructure build-out in a generation. The capex, the grid, and why power (not silicon) is now the thing being built. The upside: that scarcity is fueling record efficiency gains, and smart pricing keeps AI flowing to everyone.

demand ×330 vs. power ×2.4 ~1,000 TWh data centers power, not chips
Read the sheet
03
Field study · Self-hosted vs. cloud

Is Self-Hosted AI Actually Greener?

A measured look at one always-on home AI box — built from its own usage logs. The honest answer turns on how busy the box is: it gets greener the more you use it, and it already wins on embodied carbon, water, and privacy. The same query runs 2.6–43 Wh depending on how you count idle. Includes both attribution models, the break-even math, and the cheap wins that close the gap.

2.6–43 Wh / query greener the more you use it real telemetry
Read the sheet
04
Pricing · 8 providers compared

Consumer vs. Enterprise Token Usage

Two billing universes running across every major provider. One sells a flat-rate bucket of usage with rolling caps; the other meters every token you burn. Side-by-side consumer plans and API rates from the premium labs down to the cheap open-weight disruptors — plus how image inputs are counted and why generating an image is billed nothing like reading one.

~$20/mo apps vs. per-MTok APIs 8 providers caching · batch −50%
Read the sheet
05
Counterpoint · The optimistic case

The Other Half of the Story

The rest of this collection looks at the hard limits — soaring demand, finite power, carefully shared compute. This is the upbeat counterpart: AI is getting cheaper and more efficient faster than almost any technology in history. Inference costs fall ~10× a year, energy per prompt is collapsing, and small models keep eating big-model jobs. Plus the honest catch — why efficiency gains get partly eaten by demand.

~10×/yr cheaper inference 33× less energy/prompt Jevons caveat
Read the sheet
06
Economics · Bull vs. bear

Is AI a Bubble?

The most-asked question — and the one with the least honest answers, because it mixes up two different things: is the technology real? (almost certainly yes) and is the spending a financial bubble? (genuinely contested). The bull and bear cases weighed with real 2026 numbers — capex vs. revenue, GPU depreciation, circular financing — plus what the dot-com, Cisco, and railway busts actually teach us.

~$450B capex vs. real revenue both sides, weighed hedged verdict
Read the sheet
07
Build cost · Training the model

The Cost You Never See

Every other sheet is about the running meter — what each query costs. But before a model answers anything, someone spends hundreds of millions teaching it. Training is AI's other cost, with the opposite shape: enormous and one-time. The twist — spread across a model's life it's a fraction of a cent per query, yet inference still wins the lifetime total. Why "AI is wildly expensive" and "AI costs ~nothing" are both true.

$100–500M upfront, once ~0.02¢ / query amortized 60–90% lifetime = inference
Read the sheet
The footprint of this page itself

Building this collection with an AI assistant wasn't free either. Here's the rough footprint of the whole back-and-forth that produced it — estimated the same honest way as the sheets above, and small enough to be encouraging.

~1.1M
tokens processed end-to-end≈ 90K of it newly written
~250 Wh
energy (estimated)≈ a laptop running ~5 hours
~85 g
CO₂e @ ~330 g/kWh grid≈ driving ¼-mile in a gas car
~270 mL
water (data-center cooling)≈ one coffee mug

Estimate — A genuine back-of-envelope, not a meter reading. The token count is approximate; energy / carbon / water are derived from published per-token cloud-inference figures (plausible range ~150–650 Wh) and a ~330 gCO₂/kWh grid, using the same 1.08 mL-per-Wh ratio as the cited cloud benchmark. Real numbers swing with the model, data center, and grid. The hopeful part — prompt caching (sheet 04) and the efficiency trend (sheet 05) mean the same work keeps getting cheaper and greener — this is already a small fraction of what it would have cost a couple of years ago.

About — A small collection of standalone reference sheets on AI token economics. Plain HTML/CSS with one shared stylesheet for navigation — no build step, no frameworks. Sources are cited at the foot of each individual sheet (Anthropic docs, Google I/O 2026, IEA, Gartner, Menlo Ventures, and others).