Nvidia Vera Rubin NVL72 Boosts Agentic AI Throughput by 30x

Nvidia Vera Rubin NVL72 Boosts Agentic AI Throughput by 30x

Nvidia’s Vera Rubin NVL72 rack-scale systems combine CPUs, GPUs, networking, and software to support high-throughput agentic AI workloads. Image: Nvidia

Nvidia says Vera Rubin NVL72 delivers up to 30x more agentic AI throughput, but enterprises should weigh workload needs, power, cooling, and cost.

Written By
Eric Mboizi
Eric Mboizi
Sep 21, 2026

Adding more GPUs does not automatically make AI workloads scale efficiently. Achieving higher throughput also requires CPUs, GPUs, memory, networking, and software designed to work together.

At the AI Infrastructure Summit 2026, Nvidia said its Vera Rubin NVL72 platform delivered up to 30 times the AI factory throughput of the GB300 NVL72 on agentic workloads. Ian Buck, Nvidia’s vice president of hyperscale and high-performance computing, presented the results and detailed the rack-scale architecture behind them.

Buck described Vera Rubin as a “fungible” full-stack platform, meaning its computing resources can support open and proprietary models across training, inference, and other accelerated workloads.

He also said agentic AI processing can require up to 100 times more compute than conventional chatbot interactions. Nvidia co-designed Vera Rubin’s CPUs, GPUs, networking, and key-value cache management to address those demands.

For IT leaders, the results show how purpose-built infrastructure could improve agentic AI performance. However, the gains do not mean every workload will run 30 times faster or that enterprises should immediately replace existing Blackwell systems.

NVIDIA Vera for Agents

Nvidia announced at the end of May 2026 that its Vera Rubin platform had entered full production. The company says the rack-scale system can deliver 10 times more performance per watt than the previous-generation Grace Blackwell platform.

At the 2026 AI Infrastructure event, Nvidia shared benchmark results for the platform using SemiAnalysis AgentX, an open-source agentic AI benchmark.

Tests involving DeepSeek-R1 and Qwen3-VL showed up to 3.7 times the performance of the Blackwell Ultra-based GB300 NVL72, according to Nvidia. This model-level result is separate from Nvidia’s broader claim of up to 30 times greater AI factory throughput across agentic workloads.

One of the notable design features of Vera Rubin is its support for efficient model routing in the increasingly popular Mixture-of-Experts (MoE) model architecture. In MoE, a token may need to be routed to an expert running on another GPU. NVL72’s high-bandwidth fabric is designed to make this communication much less costly.

What this means for enterprises

Advertisement

The comparison may suggest that Vera Rubin is simply “better” than Blackwell, but the results require context. Vera Rubin succeeds the Blackwell Ultra-based GB300 NVL72 and improves the full rack-scale system, including its CPUs, GPUs, memory, networking, and software. Benchmark gains will vary by model, workload, and configuration.

Enterprises should not rush to replace existing infrastructure based on these benchmarks alone. IT leaders should first assess workload growth, power and cooling capacity, software compatibility, deployment timing, and the cost of adopting a new rack-scale system.

Nvidia said software optimization can also improve existing hardware performance. According to the company’s preliminary results, GB300 NVL72 delivered a 1.6x improvement with its 6.1 software release compared with version 6.0.

Existing Blackwell infrastructure can therefore continue to deliver value, particularly as software updates improve performance. Enterprises should consider Vera Rubin when projected agentic AI demand justifies its cost, power, cooling, and migration requirements—not simply because one benchmark produced a higher number.

Read more: Nvidia’s GTC 2026 announcements show how Vera Rubin, agentic AI, and expanding AI factories are reshaping enterprise infrastructure plans.

Eric Mboizi

Eric Mboizi is a technology news writer covering software development, emerging technologies, and the evolving digital landscape for TechRepublic and eWeek. He holds a bachelor’s degree in software engineering from Makerere University and has more than five years of experience creating technical content for developers and technology professionals. In addition to his work as a journalist, Eric is an Ethereum developer with more than four years of experience in blockchain technology. His hands-on development background gives him a practical perspective on software engineering, decentralized technologies, and the real-world implications of new technology trends.