NVIDIA Unleashes Vera Rubin NVL72 Rack Supercomputer Delivering 3.6 ExaFLOPS & Unveils 2-Gigawatt Clean AI Factory Alliance
SAN FRANCISCO, CA & SANTA CLARA, CA — September 16, 2026 — Addressing a standing-room-only crowd of institutional investors and systems architects at the Goldman Sachs Communacopia + Technology Summit, NVIDIA CEO Jensen Huang officially confirmed that the company's next-generation Vera Rubin NVL72 rack supercomputer has entered volume hyperscale delivery.
Packed with 72 Rubin GPUs featuring next-generation HBM4 memory, 36 custom 88-core Vera "Olympus" CPUs, and an astonishing 260 Terabytes-per-second NVLink 6 all-to-all fabric, a single liquid-cooled NVL72 cabinet delivers an unprecedented 3.6 ExaFLOPS of NVFP4 inference compute. Huang simultaneously announced a landmark 2-Gigawatt Clean AI Factory Alliance with Emerald AI, Google, and national energy partners to build dedicated Small Modular Reactor (SMR) nuclear and geothermal microgrids capable of powering the astronomical energy demands of frontier reasoning models.
1. The "Extreme Co-Design" Philosophy: Beyond the Chip
Huang emphasized that traditional chip-level scaling has met its physical and thermodynamic limits. In the era of autonomous reasoning agents—where models execute thousands of internal chain-of-thought tokens, multi-turn tool invocations, and memory rollouts—the bottleneck is no longer raw arithmetic logic, but memory bandwidth, interconnect latency, and sustained datacenter power delivery.
The Vera Rubin platform redefines the server rack as the atomic computing unit:
- 72 Rubin HBM4 GPUs: Fabricated on TSMC’s advanced 3nm process node with CoWoS-L packaging, each Rubin accelerator integrates ultra-wide HBM4 memory stacks delivering an aggregate 22 TB/s of memory bandwidth—essential for KV cache compaction during million-token agentic sessions.
- 36 Custom Vera "Olympus" CPUs: Built on ARM architecture with 88 specialized Olympus compute cores, Vera acts as a dedicated traffic director, orchestrating asynchronous data pre-fetching, multi-agent sandbox boundaries, and host OS overhead without bottlenecking GPU cores.
- NVLink 6 Spine: Delivering 3.6 TB/s bidirectional bandwidth per GPU, the unified NVLink 6 copper spine enables all 72 GPUs and 36 CPUs to address a single 30-Terabyte shared memory space with single-digit nanosecond latency.
- BlueField-4 DPU & ConnectX-9: Offloads zero-trust hardware security, automated checkpoint recovery, and InfiniBand/Spectrum-6 Ethernet fabric management at line rate (1.6 Tb/s per SuperNIC).
"We are no longer building computer chips; we are industrializing intelligence factories. When a reasoning model spends 15 minutes thinking before giving an answer, that is not compute—that is energy being converted into synthesized cognition. Vera Rubin was engineered from the electron up to make that cognition 10 times more energy-efficient."
Vera Rubin NVL72 Flagship Performance Metrics
2. The 2-Gigawatt Energy Barrier: SMRs & Clean Microgrids
While compute hardware has advanced rapidly, the global electrical grid has emerged as the single greatest constraint on AI advancement. Utility interconnection queues in Northern Virginia, Texas, and Western Europe now stretch to 2031, forcing hyperscalers to rethink data center power architecture.
To bypass municipal grid delays, NVIDIA’s newly forged alliance with Emerald AI, Google, and energy infrastructure developers commits 2 Gigawatts of dedicated off-grid clean power scheduled to come online between late 2026 and 2028:
- Modular Nuclear (SMR) Clusters: Co-locating 300MW Small Modular Reactors directly adjacent to AI factory mega-campuses in Wyoming and Ohio, providing uninterrupted 24/7 baseline power.
- Enhanced Geothermal Systems (EGS): Leveraging closed-loop geothermal drilling technology to power high-density cooling towers and thermal dissipation systems with zero freshwater consumption.
- Direct DC Bus Distribution: Vera Rubin racks interface directly with 800V DC power delivery buses, eliminating AC-to-DC conversion stages and recapturing 6% of lost transmission efficiency.
3. Hyperscale Deployment: AWS, Azure, Google Cloud & OCI
Every major cloud provider has confirmed initial multi-rack Vera Rubin cluster deployments slated for customer availability by Q4 2026:
- Microsoft Azure: Deploying Rubin NVL72 clusters in Wisconsin and Georgia to power OpenAI’s next-generation reasoning architectures and enterprise copilots.
- Amazon Web Services (AWS): Integrating Rubin racks into Project Rainier with bespoke liquid-cooling distribution units for frontier pharmaceutical and aerospace research.
- Google Cloud: Pairing Vera Rubin clusters alongside TPU v6e arrays within unified Vertex AI pods for hybrid multi-modal inference.
- Oracle Cloud Infrastructure (OCI): Deploying multi-rack clusters in its Abilene, Texas facility, supporting startup and open-weights developers needing ultra-high bandwidth RDMA interconnects.
4. What This Means for the Enterprise & AI Economics
The arrival of the Vera Rubin NVL72 fundamentally alters the unit economics of AI software. By lowering inference costs per token by up to 85% compared to Hopper-era clusters, complex reasoning workloads that once cost dollars per query can now be delivered at penny scale.
For enterprise teams deploying autonomous agents—from automated customer support agents to software engineers that write, test, and deploy software autonomously—the transition to Vera Rubin marks the moment where continuous background agent cognition becomes economically viable across global business operations.