Most Tokens Go to Non-SOTA Models: 77% Performance at 2.5% Price

State-of-the-art models are improving rapidly—two-thirds smarter than last November, with two new models released every three days. Yet 84% of tokens on OpenRouter are not state of the art. The six models carrying the supermajority of volume deliver about 77% of frontier performance at 2.5% of Claude Fable 5’s price.

Ramp‘s data shows buyers are price-elastic. Fable 5 at about $10 per million tokens captured 6% of Anthropic tokens and 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive. Each new state-of-the-art release should move less share than the one before it.

Enterprises will consolidate spend on one or two vendors, similar to the cloud era. Once a model clears a high-value job, the workload stays. Performance is already good enough at a meaningful discount. The gap keeps closing from below: the best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier.

Application deployment tells a different story. More startups default to smaller models, fine-tuned models, and open source, optimizing a different Pareto frontier: price over performance. If share stops shifting and good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, and that bar will rise over time.

Frontier models still win on software architecture and security design, where the best available model earns its price. But the open data suggests the frontier that matters for most tokens is the other one: price over performance.

Honestly, Who Buys SOTA?

View Original