Building abundant intelligence: OpenAI’s economic cycle for AI

OpenAI‘s latest post frames the value of AI infrastructure not by its size but by what it enables: more capable intelligence available to more people at lower cost. The company describes a self-reinforcing cycle: better intelligence drives broader adoption, which funds further investment in research and infrastructure, which in turn improves intelligence and efficiency.

Recent pricing changes illustrate the cycle. OpenAI reduced the price of GPT-5.6 Luna by 80 percent, to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra dropped 20 percent to $2 and $12. For GPT-5.6 Sol, a Fast mode delivers 2.5 times the speed at twice the price with no change in intelligence. The article argues that customers should evaluate models based on the cost of a successful outcome, not token price. A stronger model that completes work correctly on the first try can be more economical than a cheaper model requiring retries and human oversight.

Efficiency improvements come from making every unit of compute more productive, not just from adding capacity. GPT-5.6 Sol itself helped optimize the production serving software, reducing end-to-end serving costs by 20 percent, and improved speculative decoding to increase token-generation efficiency by more than 15 percent. System-level changes also matter: better routing, smarter context management, and stronger product design reduced steps and waste. In one example, improvements to retained reasoning and context management raised GPT-5.6 Sol’s score on the ARC-AGI-3 benchmark from 13.3 percent to 38.3 percent while using six times fewer output tokens, with no change to the model itself.

The full-stack approach—covering infrastructure, models, platform, and products—creates a feedback loop. Product use reveals where customers find value and friction, which shapes research. Research improvements strengthen products and lower serving costs. Demand across ChatGPT, ChatGPT Work, Codex, and the API guides capacity decisions. The scale is large: models reach over one billion active users and two million businesses. Users send 50 percent more messages daily after six months and use ChatGPT for twice as many kinds of work. Agentic work through Codex now accounts for 99.8 percent of weekly output tokens, and teams like Finance have adopted agentic tools as a primary work method.

Investment decisions are based on evidence: user growth, enterprise commitments, API consumption, utilization, revenue, and model capability progress. Technical and commercial milestones determine when projects advance. The goal is not maximum infrastructure but right capacity at the right time against credible demand. The post ends by stating that the true measure of progress is the amount of useful work intelligence makes possible, how efficiently it is delivered, and how widely its benefits are shared.

Building abundant intelligence

View Original