LIVE
All stories ›
AI IN LIFENEWS
Tools & AppsBusiness & DealsAI ModelsResearchSocietyChips & ComputeSafety & SecurityRegulation & PolicyRoboticsReviews OpenAIAnthropicGoogle & DeepMindAlibaba / QwenMetaxAIByteDance
Home › OpenAI › CHIPS
CHIPS

GPT-6 Astra Ultrafast: Up to 8x Faster on Blackwell

NVIDIA says the Blackwell-tuned Ultrafast mode generates tokens up to 8x faster than Astra Standard, and it is already in the OpenAI API.

GPT-6 Astra Ultrafast: Up to 8x Faster on Blackwell
Symbolic image: a tight crop on the half-open door of a liquid-cooled GPU inference rack as a gloved hand reaches for the latch.

In short

NVIDIA attributes the speed of GPT-6 Astra Ultrafast to inference optimizations written for the Blackwell architecture, which the company puts at up to 8x faster token generation than Astra Standard mode.

At a glance

  • Up to 8x faster token generation than Astra Standard mode — NVIDIA's own figure, with no published benchmark method.
  • Runs on NVIDIA Blackwell GPUs; the post names no specific board or system configuration.
  • Available in the OpenAI API and to eligible ChatGPT Work and Codex users.
  • OpenAI used its own models to produce and test inference kernels for NVIDIA GPUs, according to the post.
  • No pricing, no millisecond latency and no tokens-per-second numbers appear in the source.

NVIDIA's case for GPT-6 Astra Ultrafast rests on software rather than silicon: inference code tuned for the Blackwell architecture is what the company credits for the jump. The published figure is up to 8x faster token generation than the same model's Astra Standard mode. Access is already open through the OpenAI API, plus eligible ChatGPT Work and Codex accounts.

The one number in the post

Written by Dion Harris and posted on October 1, 2026, the piece commits to a single metric: the 8x multiplier on token generation, measured against Astra Standard rather than against a rival model. It does not say which Blackwell board was used, how long the prompts were, or how heavily the system was loaded.

The use cases NVIDIA names are code generation, tool use and interactive applications. That lines up with where the mode shipped first — Codex and the Work tier of ChatGPT.

Why coding agents feel decode speed

An agent loop is short and repetitive: emit code, call a tool, read the result, pick the next move. Every pass pays for a full decode, so output speed compounds across dozens of turns instead of hiding inside one long answer.

Whether an 8x decode gain turns into an 8x shorter session is not claimed anywhere in the post. Tool latency, network hops and test runs all sit off the GPU.

A model tuning its own serving stack

The more unusual claim is methodological. OpenAI says it put its own models to work refining the inference software that serves them on NVIDIA hardware. Philippe Tillet, who leads inference at OpenAI, credits NVIDIA's tooling and documentation for making those models good at producing “high-performance kernels” for Blackwell.

That is a feedback loop worth watching: the model writes the code that makes the model faster. The post offers no breakdown of how much of the optimization work was machine-generated.

What is not verified

This is a chip vendor writing about its largest customer's product, and it is the only account on the record. There is no pricing, no millisecond latency, no tokens-per-second figure and no reproducible benchmark description. No independent test of the 8x claim had surfaced when this article was written, so treat the number as a vendor estimate until you measure your own workload.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is GPT-6 Astra Ultrafast?

A faster serving mode of OpenAI's GPT-6 Astra that NVIDIA says runs on Blackwell GPUs and targets code generation, tool use and interactive applications.

How much faster is Astra Ultrafast?

NVIDIA cites up to 8x faster token generation versus Astra Standard mode. The post publishes no benchmark method, latency figures or tokens-per-second numbers.

Who can use GPT-6 Astra Ultrafast?

Developers through the OpenAI API, plus eligible ChatGPT Work and Codex users; no further tiers or rollout timeline are given.

Sources

More reports