BREAKING
+++ Anthropic reportedly in talks to buy Decart for $6 billion +++ OpenAI previews Ultrafast mode: GPT-5.6 Sol at up to 750 tokens per second +++ Anthropic experiment: AI agents sabotage each other +++ OpenAI run rate tops $40 billion – Dali Rajic named new sales chief +++ IBM and OpenAI launch enterprise partnership with dedicated practice +++ Microsoft merges Copilot apps and cuts unsuccessful features ++++++ Anthropic reportedly in talks to buy Decart for $6 billion +++ OpenAI previews Ultrafast mode: GPT-5.6 Sol at up to 750 tokens per second +++ Anthropic experiment: AI agents sabotage each other +++ OpenAI run rate tops $40 billion – Dali Rajic named new sales chief +++ IBM and OpenAI launch enterprise partnership with dedicated practice +++ Microsoft merges Copilot apps and cuts unsuccessful features +++
Updated 08:00
AI IN LIFE AI IN LIFENEWS
DAILY
Models

OpenAI Previews Ultrafast Mode: GPT-5.6 Sol at 14x Speed

A new preview mode pushes GPT-5.6 Sol to up to 750 tokens per second – powered by a partnership with chipmaker Cerebras.

OpenAI Previews Ultrafast Mode: GPT-5.6 Sol at 14x Speed

Illustration · AI-generated (AI IN LIFE)

At a glance

  • Ultrafast pushes GPT-5.6 Sol to up to 750 output tokens per second
  • Up to 14x the speed of standard operation
  • Hardware partner is Cerebras with its wafer-scale chips
  • Status: preview for a small customer group, no public pricing yet
  • Target uses: incident response, customer service, financial analysis, e-commerce

OpenAI has unveiled a new operating mode for its GPT-5.6 Sol model: “Ultrafast” delivers responses at up to 14 times the standard speed – up to 750 output tokens per second. The mode is in preview and initially available only to a small group of customers.

The speed jump is powered by a partnership with Cerebras. The chipmaker, known for its wafer-scale processors, confirmed the collaboration on August 13 in its own announcement. Cerebras systems specialize in extremely fast inference, competing directly with Nvidia-based setups.

OpenAI frames Ultrafast as an answer to a familiar dilemma: until now, getting real-time speed usually meant settling for a smaller or specialized model. “Ultrafast points to progress in a new direction: more useful work per second,” the company says – full model quality at real-time pace.

Target applications include incident response in IT operations, customer service, financial market analysis and e-commerce – scenarios where latency directly determines usefulness. For agentic workflows that chain many steps, every speed gain compounds.

Open questions remain: OpenAI has not announced pricing, and expansion beyond the small customer group will come only “as capacity grows.” How much Cerebras hardware actually sits behind the offering is not public.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is Ultrafast mode?

A new operating mode for GPT-5.6 Sol that runs the same model at up to 14x speed – up to 750 tokens per second.

Who can use it?

Currently only a small group of OpenAI customers in a preview. Broader availability is planned as capacity grows.

What is Cerebras' role?

Cerebras provides the specialized wafer-scale hardware behind Ultrafast and confirmed the partnership in its own announcement.