Groq 3 LPX Ships: Nvidia Targets Agent Decode Latency
The accelerator posts 3,400 output tokens per second on Gemma 4 31B, and Nebius becomes the first AI cloud to run it. Vendor figures, no ship dates.
Illustration: a data hall at night, with long rows of dark server cabinets and bundles of fiber cabling running overhead.
Nvidia moved its Groq 3 LPX inference accelerator into full production on August 24, 2026, extending the Vera Rubin NVL72 rack platform with a tier built solely for fast token generation in agentic systems.
At a glance
- Nvidia says Groq 3 LPX entered full production on August 24, 2026, extending the Vera Rubin NVL72 rack system.
- Vendor benchmark: 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context window.
- Nvidia calls the setup up to four times more responsive than the nearest platform; no independent test exists yet.
- Up to 256 accelerators per rack deployment; Nebius is the first AI cloud running the chip.
- Spectrum-X Multiplane scales to 512,000 GPUs by vendor figures; CoreWeave has it in production.
Agents are slow in one specific place, and Nvidia has now shipped hardware aimed at that place. Groq 3 LPX, announced on August 24, 2026, is in full production as an add-on tier to the Vera Rubin NVL72 rack platform, and it does a single job: push output tokens out fast.
Why decode is the bottleneck
Inference splits into reading the prompt and writing the answer one token at a time. The second half sets the speed a user actually feels, because a tool call, a read of its result, and a follow-up request stack their delays on top of each other. Nvidia's argument is that a tier dedicated to that half beats throwing more general-purpose silicon at it.
SiliconAngle quotes Nebius Chief Technology Officer Danila Shtan to the same effect: generation, not prefill, is the phase that determines how responsive a system seems. Chief Executive Jensen Huang frames Vera Rubin there as a set of workload-tuned AI factory configurations built for agentic work.
What the vendor claims
The headline figure is 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context window. A rack deployment holds up to 256 accelerators. Nvidia puts the result at up to four times more responsive than the nearest competing platform.
All three are company numbers. No third-party measurement of throughput, tail latency, or power draw was available when the announcement went out, so treat them as claims rather than results.
The networking half of the story
Nvidia deliberately presents this as several layers moving together rather than one chip. The package covers:
- Spectrum-X Multiplane — an Ethernet fabric the company says scales to 512,000 GPUs and performs 1.6x better than standard Ethernet.
- Spectrum-6 — a switch ASIC rated at 102.4 Tb/s, paired with ConnectX-9 SuperNICs supplying 1,600 Gb/s per GPU.
- Failure behavior — in an eight-plane topology, losing one plane is said to retain 90 percent of bandwidth, with hardware recovery running 11 times faster than software load balancing.
- Scale-In — built on BlueField-4 and pitched by Nvidia as the fifth pillar of its AI networking stack.
- NVLink Fusion — pulls third-party silicon into Nvidia's rack architecture for semi-custom AI factories.
Multi-site NCCL collectives are credited with a 1.9x speedup.
Who is running it
Nvidia names Nebius as the first AI cloud with Groq 3 LPX, delivered to customers through the Nebius Token Factory. CoreWeave has Spectrum-X Multiplane in production. For SpaceX, Nvidia describes Vera CPUs going into data centers on the ground as well as satellites in orbit.
What the announcement leaves out
There are no delivery dates, no prices, and no unit volumes; beyond the named customers, availability is unstated. The blog post's own title lists a CPX variant alongside LPX, and the body never explains what it is.
The backstory rests on a single source: SiliconAngle values Nvidia's December 2025 licensing of Groq technology at $20 billion and reports that founder Jonathan Ross and President Sunny Madra joined the company. Nvidia's own post does not mention it, so that piece is not independently confirmed here.
FAQ
What is Nvidia Groq 3 LPX?
An accelerator dedicated to the output phase of inference that extends the Vera Rubin NVL72 rack system rather than replacing it. Nvidia aims it at agents, where per-token delay compounds across long chains of tool calls.
How fast is Groq 3 LPX?
Nvidia reports 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, and describes the configuration as up to four times more responsive than the nearest platform. Both are vendor figures with no independent benchmark.
When can you buy Groq 3 LPX and what does it cost?
Nvidia declares full production as of August 24, 2026 but gives no delivery dates, prices, or volumes. Nebius is named as the first AI cloud running the chip; wider availability is not addressed in either source.