Nvidia's upcoming Rubin platform: up to 67x efficiency in early measurements
SemiAnalysis benchmarked the Rubin NVL72 data-centre system against the current Blackwell generation — with a clear lead on agentic AI workloads.

Illustration · AI-generated (AI IN LIFE)
At a glance
- System tested: Nvidia Rubin NVL72 (successor to GB300 NVL72)
- Key metric: throughput per dollar on agentic AI workloads, up to 67x higher
- Also measured: 7x throughput per megawatt vs. official Blackwell figures, on pre-release software
- Technical basis: SGLang, TensorRT-LLM, vLLM, mixed-precision formats (MXFP4/MXFP8), NVLink linking 72 GPUs
- Source: independent measurements by SemiAnalysis, not Nvidia's own marketing figures
Market analysis firm SemiAnalysis has published early performance measurements of Nvidia's upcoming Rubin platform, based on pre-release software. At the centre is the Rubin NVL72 data-centre system, successor to the current Blackwell generation.
On so-called agentic AI workloads — tasks where AI models independently carry out multi-step actions — Rubin NVL72 achieves up to 67 times higher throughput per dollar invested than the current GB300 NVL72 generation, according to the measurements.
Notably, even on early, unfinished software, Rubin already reaches seven times the officially stated throughput per megawatt of the current Blackwell generation. SemiAnalysis estimates this lets Rubin earn more than twice the profit per gigawatt of data-centre capacity compared with the Blackwell platform.
The efficiency gains do not come from a single innovation but from several pieces working together: new model-serving software including SGLang, TensorRT-LLM, and vLLM, mixed-precision computing, and a faster NVLink connection joining 72 graphics processors inside a single rack.
For data-centre operators, the numbers — if confirmed in the final version — mean one thing above all: from the Rubin generation onward, the same investment budget should buy noticeably more usable AI computing power than it does today.
FAQ
What does "agentic AI workload" mean?
It refers to tasks where an AI model does not just answer a single query but independently carries out several steps in sequence. That places different demands on hardware than short, individual requests.
Are these official Nvidia figures?
No, these are independent measurements by analysis firm SemiAnalysis on pre-release software. Final, official figures may differ at launch.
When does Rubin NVL72 actually launch?
No precise launch date was given in these measurements; the tests explicitly ran on early pre-release software, suggesting a market launch still ahead of final optimization.


