Ultrafast: OpenAI speeds up GPT-5.6 Sol by up to 14x
The new API tier delivers up to 750 output tokens per second on Cerebras hardware — for now in limited preview for select customers.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Up to 14x speed for GPT-5.6 Sol in the new Ultrafast tier
- Up to 750 output tokens per second
- Runs on Cerebras wafer-scale infrastructure
- Currently a limited preview for select customers, expanding with capacity
- Pricing not yet published
OpenAI has introduced “Ultrafast”, a new service tier for its API that runs the frontier model GPT-5.6 Sol up to 14 times faster. The speed comes from wafer-scale hardware built by chipmaker Cerebras, designed for ultra-low-latency inference.
Concretely, OpenAI promises up to 750 output tokens per second — a figure previously associated with small, weaker models rather than the company's most intelligent one. That combination is the core of the announcement: full model quality at response times that make synchronous, interactive applications possible.
OpenAI's target scenarios include real-time incident response and system diagnostics, financial research and fraud detection, live customer support and voice interactions, and e-commerce product assistance and checkout support. According to OpenAI, early testers report that workflows previously blocked by the latency of intelligent models are becoming practical.
Ultrafast is initially available only as a limited preview for select customers, with access expanding as capacity grows. Pricing has not been disclosed. Even so, for builders of voice agents, support automation and real-time workflows — including Europe's mid-market — this is a directional call: the latency barrier to using frontier models in customer-facing products is coming down.
FAQ
What exactly is Ultrafast?
A new OpenAI API service tier that serves GPT-5.6 Sol at up to 14x speed for time-sensitive applications.
Who can use Ultrafast?
Only select customers in a limited preview for now; OpenAI plans to widen access as capacity increases.
What is the mode good for?
Real-time applications such as voice agents, live support, fraud detection and incident response that previously failed on large-model latency.


