NVIDIA’s Vera Rubin delivers 30x more AI work per megawatt

NVIDIA’s Vera Rubin NVL72 preview results achieve up to 30 times higher AI-factory throughput per megawatt than GB300 NVL72, according to testing with the SemiAnalysis AgentX benchmark. This leap in performance arrives as the scale of agentic AI rapidly expands; OpenRouter’s State of AI report found single agentic requests now consume 15 times the tokens of ordinary chat.

AgentX specifically evaluates AI infrastructure for agentic-coding inference by replaying production-style sessions, measuring efficiency in a way that captures long-context, KV-cache reuse, and MoE execution. The benchmark highlights how Blackwell GB300 NVL72 extends an order-of-magnitude throughput-per-megawatt advantage to these dynamic agentic workloads.

SemiAnalysis AgentX Benchmark Measures Agentic-Coding Inference

SemiAnalysis AgentX assesses AI infrastructure by replaying production-style sessions, focusing on agentic-coding inference and revealing substantial performance differences between hardware platforms. The benchmark specifically measures how efficiently accelerators handle request patterns from real coding agents, accounting for long-context prefill, KV-cache reuse, interactive decode, tool-call gaps, and distributed mixture of experts execution under realistic concurrency. AgentX differs from traditional benchmarks by simulating variable inference sequences, unlike the fixed prompt-and-response patterns of earlier tests like the InferenceX static 8K/1K sequence length results for DeepSeek-R1-0528.

Tokens per megawatt is prioritized as the most critical metric for AI factories, alongside four user-experience values: E2E normalized interactivity, standard interactivity, E2E latency, and time to first token. AgentX testing methodology preserves the original session timing of replayed agent traffic, reproducing KV cache-capacity pressure that a realistic benchmark must capture, and varies concurrency to map the trade-off between throughput and interactivity.

The benchmark reports sustained throughput per provisioned megawatt, alongside each stack’s ability to reuse repeated context, a key factor in agentic workloads.

Vera Rubin NVL72 Achieves 30x Throughput per Megawatt

This performance leap impacts the economics of running large-scale agentic systems, allowing for more interactive capacity within a fixed power and infrastructure budget. NVIDIA reports these results were measured using the AgentX workload and are pending review by SemiAnalysis, highlighting a commitment to external validation of performance claims. OpenRouter’s analysis of 100 trillion tokens reveals single agentic requests consume 15 times more tokens than ordinary chat interactions, underscoring the need for optimized infrastructure.

The NVLink scale-up fabric, connecting the 72 GPUs within GB300 NVL72, facilitates the high-bandwidth communication and KV-cache movement necessary to serve large models as a coordinated rack-scale system. The benefits of GB300 NVL72 become more pronounced as model scale increases, with throughput per megawatt for the Kimi K3 2.8T model reaching roughly 80 times that of H200 NVL8 at comparable interactivity.

This allows GB300 NVL72 to extend the interactivity frontier to roughly 215 tokens per second per user, exceeding the capabilities of H200 NVL8. The broader Vera Rubin platform builds on this approach, extending it across the entire workflow, and demonstrating what is possible when a rack-scale system is tuned for the specific demands of agentic inference.

Blackwell GB300 NVL72 Outperforms H200 in Agentic Workloads

This substantial increase in performance arrives as agentic AI workloads become increasingly complex, demanding more efficient infrastructure to handle longer, stateful interactions. Agentic sessions differ from conventional AI interactions by chaining model calls, incorporating tool use, and continuously expanding context, creating a variable and demanding workload for AI hardware. GB300 NVL72 delivers up to 15 times higher AI-factory throughput per megawatt than the H200 NVL8 when running DeepSeek V4 Pro 1.6T on the AgentX workload, sustaining more responsive agentic inference within a fixed power budget.

For data center operators, this means a fixed power and infrastructure budget can support significantly more interactive agentic capacity, or maintain existing capacity at a substantially reduced operating expense. NVIDIA reports that “Vera Rubin delivers 30 times higher throughput per megawatt at 160 interactivity.” The Vera Rubin platform extends this performance optimization beyond individual components, encompassing the entire workflow to maximize efficiency.

AgentX Metrics Define AI-Factory Performance & Interactivity

The benchmark’s design prioritizes capturing the nuances of agentic workloads, moving beyond fixed sequence length scenarios that have become less representative of true serving performance. The benchmark’s methodology, which incorporates KV cache-capacity pressure through preserved session timing, ensures a realistic evaluation of performance under demanding conditions. For the Kimi K3 2.8T model, throughput per megawatt reaches roughly 80 times that of H200 NVL8.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Rusty Flint

Rusty Flint

Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things Quantum Physics especially. Rusty likes the more esoteric side of Quantum Computing and the Quantum world. Everything from Quantum Entanglement to Quantum Physics. Rusty thinks that we are in the 1950s quantum equivalent of the classical computing world. While other quantum journalists focus on IBM's latest chip or which startup just raised $50 million, Rusty's over here writing 3,000-word deep dives on whether quantum entanglement might explain why you sometimes think about someone right before they text you. (Spoiler: it doesn't, but the exploration is fascinating)

Latest Posts by Rusty Flint: