
The Three Waves of AI Consumption

On February 6, 2026, agents on OpenRouter consumed more tokens than humans did, and they have not given the lead back. The post argues AI consumption is not a smooth curve anyone can extrapolate; it arrives in three waves, each order of magnitude larger than the last. The second wave has already broken, and the third is on the horizon.
The first wave is chat: a person asks, the model answers once, and it is done. The author estimates this at roughly 1 million tokens per active user per day. But chat is already the minority of enterprise output—by June 2026 it accounted for a little more than a third of enterprise use, with Codex representing the other 64%. The second wave is a single agent doing meaningfully larger work: vibe coding a family weekend sports scheduling app, or researching ten-year bond trends and their macroeconomic implications. That is roughly 100 to 200 million tokens per day, again an estimate, and the token-intensity gap between these waves is two orders of magnitude. The third wave stacks a meta-harness on top: one AI dispatching many agents in parallel, each spawning its own tool calls and sub-agents—another order of magnitude, pushing consumption into the billions.
Consumption does not grow because generation gets quicker; it grows because parallelization compounds. A concrete meta-harness example: a spreadsheet in which each cell is an AI answer rather than a number—a list of 200 Y Combinator startups with 30 columns covering founder backgrounds, company history, and value proposition. Each column represents tens of AI tool calls, and the run costs tens of millions of tokens before anyone notices. The shift is visible in traffic data: agent volume on OpenRouter rose from 0.51 trillion to 7.3 trillion tokens in six months, a 14x increase, while human volume managed 2.8x. Goldman Sachs projects consumer and enterprise agents will consume 120 quadrillion tokens per month by 2030, 24 times the 2026 level.
The waves do not replace each other; they stack. Chat keeps growing, agents multiply at least, and meta-harnesses will dwarf both. The author’s closing warning is direct: anyone sizing compute off a smooth extrapolation is planning for the wrong curve.


