
GPT‑5.6 Sol Ultrafast preview: up to 14x speed with Cerebras

OpenAI has announced an early preview of Ultrafast, a new service tier for GPT‑5.6 Sol that runs up to 14× faster than Standard processing, first available via the OpenAI API. Powered by Cerebras, Ultrafast delivers up to 750 output tokens per second, aiming to eliminate the historical trade-off between model intelligence and real-time speed.
The post frames Ultrafast as enabling frontier intelligence in time-sensitive workflows where every second matters. Supported use cases include incident response (analyzing logs, code changes, and engineer reports during an active outage), financial research and security (assessing transactions and suspicious activity in changing market conditions), customer support and voice (resolving multi-step issues without interrupting the conversation), commerce (personalizing recommendations and resolving checkout issues in real-time), and live research (turning overnight experiments into interactive iterative sessions).
OpenAI is testing Ultrafast internally on incident response and research workflows. Internally, engineers use it to read logs, analyze traces, synthesize conversations, and validate fixes while an incident is still unfolding, with humans retaining judgment and deployment responsibility. In research, the team reports tightening the overnight batch experiment loop to support multiple iterations during a single workday.
The preview is limited to a select group of customers, and OpenAI plans to expand access as capacity grows. Customers interested in frontier intelligence at high speed can sign up for notifications. The announcement emphasizes that Ultrafast is built on the existing partnership with Cerebras for ultra-low-latency inference, and that learnings from the preview will guide future product decisions.

