
OpenAI’s Full-Stack Compute Strategy and Jalapeño Chip Results

OpenAI frames its compute strategy as one integrated system rather than a collection of separate parts. The core idea is that progress compounds fastest when data centers, chips, models, the developer platform, products, and devices improve together: better software makes hardware more productive, hardware designed for OpenAI‘s workloads improves speed and efficiency, and more capable models unlock better products that generate more demand, usage, and learning, which feed back into the system.
The post shares the first measured performance results for Jalapeño, OpenAI‘s first custom inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. It also performed strongly on DeepSeek R1 and Kimi K2, indicating the gains extend across model families. Jalapeño gives OpenAI greater control over how models run and over serving economics. By developing the model, serving software, chip, memory, and network together, OpenAI can improve throughput, latency, energy efficiency, and cost as one system. It creates a credible first-party path alongside accelerators from other partners, and future generations are already underway.
OpenAI‘s portfolio now includes Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank. The goal is to stay on the Pareto frontier, seeking the strongest mix of capability, speed, reliability, efficiency, and cost for each workload. OpenAI actively manages this portfolio for both capability and economics, using premium systems where capability matters most and optimizing for efficiency where scale and cost matter more. Preserving credible choice across providers, hardware, and deployment models lets OpenAI direct demand toward the strongest performance per dollar and maintain pricing discipline. Direct control adds leverage where tighter integration improves the entire system, while partnerships help move faster where the ecosystem is stronger.
Data centers create another point of leverage. Project Camellia in Georgia shows how OpenAI designs facilities around customer workloads while creating jobs, supporting local businesses, covering project infrastructure and energy costs, conserving water through a closed-loop system, and subjecting commitments to an annual independent public audit.
The value of the system is measured by useful intelligence per dollar. On the Artificial Analysis Coding Agent Index, GPT-5.6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model. For customers, this means faster results, more dependable products, fewer retries, agents that complete longer workflows, and lower total cost for successful work. As useful intelligence becomes more capable and affordable, more work becomes economically practical, which the post connects to Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity.
The compounding advantage is that better technology creates better economics, better economics fund the next wave of progress, and every gain makes the whole system stronger. Growth funds continued investment in research, infrastructure, and safety.


