Nvidia Vera Rubin Platform Ramps Production With 10x Efficiency Gain Over Grace Blackwell
Nvidia's Vera Rubin NVL72 platform is ramping production with partners CoreWeave, Google Cloud, Microsoft Azure, and Oracle. CoreWeave benchmarks show 10x more throughput per megawatt than Grace Blackwell.
Nvidia on July 21, 2026, announced that its Vera Rubin NVL72 platform is entering volume production, with racks already running at partners CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. The platform spans more than 350 factory sites across 30 countries, making it the largest rack-scale supply chain Nvidia has ever assembled.
CoreWeave's first benchmark on the DeepSeek-R1 model showed 10x more throughput per megawatt than the Grace Blackwell NVL72, a metric that directly addresses power constraints in AI factories. The result underscores the platform's focus on performance per watt and lowest token cost.
Vera CPU and Olympus Core Architecture
At the heart of Vera Rubin is the Vera CPU, Nvidia's first data center processor built with a custom core design. The Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency compared to competing chiplet designs, according to Nvidia. The CPU features 88 cores and 176 threads, with general release scheduled for the second half of 2026.
The platform integrates seven chips and five rack trays codesigned as a single system: Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX. Key specifications include:
- Sixth-generation NVLink scale-up networking with 2x throughput, 3x lower latency, and 10x higher packet rates versus off-the-shelf Ethernet.
- Spectrum-X Ethernet with 102.4T Spectrum-6 switches and 1.6T ConnectX-9 SuperNICs, enabling 1.6x higher RDMA bandwidth.
- NVIDIA Photonics with co-packaged optics for scale-out, delivering 5x lower power and 10x higher mean time between interruptions versus pluggable transceivers.
- A cable-free, fan-free, hose-free compute tray design that cuts assembly time from hours to one minute.
- 45-degree Celsius liquid cooling inlet temperature enabling chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually.
European Infrastructure and Partner Deployments
Vera Rubin is also powering a new multibillion-dollar partnership between Microsoft and Mistral to expand AI infrastructure in Europe. Mistral will draw on thousands of Vera Rubin GPUs to increase compute availability for customers, providing a shared platform for training, inference, and large-scale deployment. The partnership aims to deliver sovereign-ready AI across public cloud, cloud-connected, and fully disconnected private cloud environments.
Nvidia said the platform's full-stack architecture combines accelerated computing, networking, and software to support agentic workloads that can consume up to 15x more tokens than traditional AI applications. With production ramping now and general availability of the Vera CPU expected in H2 2026, the Vera Rubin platform is positioned to define the next generation of AI factory infrastructure.
Fact check
-
CoreWeave benchmark showed 10x more throughput per megawatt than Grace Blackwell NVL72 on DeepSeek-R1.
verified · source
-
Vera CPU has 88 cores and 176 threads, with general release in H2 2026.
reported · source
-
Vera Rubin platform spans 350+ factory sites in 30 countries.
verified · source
-
Microsoft and Mistral partnership is a multibillion-dollar agreement to expand AI infrastructure in Europe using Vera Rubin.
verified · source
-
Vera Rubin NVL72 compute tray has no cables, fans, or hoses, reducing assembly time from hours to one minute.
verified · source
Source reporting (10)
- NVIDIA Blog · NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
- Techmeme · Nvidia details its Vera CPU for data centers, its first CPU with a custom core design, featuring 88 cores and 176 threads, set for general release in H2 2026 (Jake Roach/Tom's Hardware)
- Tom's Hardware · Behind the scenes at Nvidia's Engineering SuperLab — Vera Rubin NVL72 running OpenAI workloads, 800VDC demonstrated, and more
- Tom's Hardware · Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack
- Tom's Hardware · Nvidia deep dives Vera CPU for AI data centers — SPEC CPU 2026 benchmarks revealed, Olympus architecture specifics, and more
- Tom's Hardware · Nvidia has shipped 'hundreds of thousands of Grace standalone servers’ — GPU firm pivots messaging as CPUs take center stage in agentic data centers
- NVIDIA Blog · Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
- ServeTheHome · Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks
- The Register · Nvidia shows off Vera Rubin platform for tokenmaxxing
- WIRED · Nvidia Wants to Own Every Chip Inside AI Data Centers
Join the conversation
You need to be registered and logged in to comment on blog articles.
Related Articles
Aging Grid and AI Demand Collide: Data Center Power Reliability Under Scrutiny After Virginia Near-Miss
Jul 26, 2026
US Data Centers Could Consume 194GW by 2035 as New Projects Emerge in Virginia, Ohio, and Australia
Jul 22, 2026
New York Becomes First US State to Impose One-Year Moratorium on Hyperscale Data Centers
Jul 14, 2026
0 Comments
No comments yet
Be the first to share your thoughts on this article.