LIVE
All stories ›
AI IN LIFENEWS
Tools & AppsBusiness & DealsAI ModelsResearchSocietyChips & ComputeSafety & SecurityRegulation & PolicyRobotics OpenAIAnthropicGoogle & DeepMindAlibaba / QwenxAIMetaByteDance
Home › OpenAI › Models
Models

OpenAI Launches Ultrafast: GPT-5.6 Sol Up to 14x Faster

OpenAI's new “Ultrafast” API tier runs GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from the $10B partnership.

OpenAI Launches Ultrafast: GPT-5.6 Sol Up to 14x Faster
Illustration · AI-generated (AI IN LIFE)

In short

OpenAI introduced a new speed tier for its API on August 14.

At a glance

  • Up to 750 output tokens per second in Ultrafast mode
  • Roughly 14x faster than standard operation
  • Built on the $10 billion Cerebras partnership
  • Three API tiers: Standard, Fast, and Ultrafast
  • Launching as a preview for select API customers

OpenAI introduced a new speed tier for its API on August 14. The “Ultrafast” mode lets flagship model GPT-5.6 Sol generate up to 750 output tokens per second — roughly 14 times faster than standard operation, according to the company. The acceleration runs on specialized Cerebras hardware, following the $10 billion partnership the two companies struck earlier in 2026.

The move establishes a three-tier system of Standard, Fast, and Ultrafast. Speed becomes its own pricing lever, much like in cloud computing: customers who need real-time responses pay for faster delivery, while batch and background workloads stay cheaper.

In parallel, OpenAI is cutting API prices across the GPT-5.6 lineup. Observers read this as a response to mounting price pressure from Chinese providers offering comparable performance at lower cost — even as DeepSeek simultaneously raises its own prices.

Ultrafast is initially available as a preview through the API and limited to select customers; companies can sign up via an interest form. OpenAI highlights real-time incident response, financial analysis, customer support, e-commerce, and interactive research workflows as target use cases.

For builders, the combination is what matters: falling base prices plus a paid turbo tier. Agentic applications that chain many model calls benefit disproportionately from higher token throughput — exactly where OpenAI is positioning itself against Google and Anthropic.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is OpenAI Ultrafast?

A new API tier that runs GPT-5.6 Sol on Cerebras hardware at up to 750 tokens per second — about 14 times faster than standard mode.

Who can use Ultrafast?

Select API customers during an initial preview, with capacity expanding gradually. Companies can register through an interest form on OpenAI's website.

Why is OpenAI cutting prices at the same time?

The company is responding to price pressure from cheaper rivals, especially from China, while turning speed into a paid premium feature.

Sources

More reports