NVIDIA’s new chip delivers leading performance in MLPerf

NVIDIA’s Vera Rubin NVL72 system achieved up to 3.7 times better throughput than the GB300 NVL72 in its first MLPerf Inference v6.1 submission, signaling a substantial gain in AI inference performance, the company says. The results, released by NVIDIA, demonstrate the importance of system performance and efficient scaling for AI infrastructure economics. “Higher system performance means more tokens generated, resulting in higher revenue,” according to the company. This debut highlights NVIDIA’s accelerated innovation and how continuous software optimizations are improving performance beyond benchmark results.

Vera Rubin NVL72 Achieves 3.7x Throughput in MLPerf Inference v6.1

NVIDIA’s Vera Rubin NVL72 system achieved up to 3.7x better throughput than NVIDIA GB300 NVL72 in its first MLPerf Inference v6.1 submission, specifically on the Qwen3-VL benchmark across offline, server, and interactive scenarios. This performance gain was realized using vLLM alongside the open-source NVIDIA Dynamo inference framework, demonstrating a significant advancement in AI inference capabilities, according to the company. The results, published by MLCommons, highlight the importance of full-stack codesign, integrating hardware and software, in maximizing AI performance and efficiency.

A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, meaning throughput increased almost linearly with added hardware. This near-proportional growth supports cost-effective scaling by minimizing the resources needed to serve a growing number of users. NVIDIA’s platform is designed to maintain high utilization rates across various workloads, from training to reasoning, and the MLPerf results validate this approach.

The company’s focus on platform fungibility, the ability to run any model on the same infrastructure, contributes to maximizing return on investment. The Vera Rubin system’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while the use of NVFP4 precision reduces memory requirements for model weights and caches. On the DeepSeek-R1 benchmark, the Vera Rubin NVL72 achieved up to 2.5 times higher throughput than the GB300 NVL72, utilizing the NVIDIA TensorRT-LLM library.

This improvement translates to a substantial increase in the number of tokens each Vera Rubin NVL72 rack can deliver, allowing for more users to be served and increased revenue generation per rack, all while lowering the cost per token, the firm reports. NVIDIA launched NVQLink in January 2026, an open architecture for integrating quantum processors with GPU supercomputers.

NVIDIA reported a quantum calibration model working across six qubit modalities in July 2026. The cuQuantum SDK v25.11, for example, delivers up to a 4,000x speedup in quantum simulation using new Pauli propagation and stabilizer primitives on GB200 NVL72 systems.

TensorRT-LLM and Dynamo Drive 1.6x Performance Gains in v6.1

This near-perfect scalability demonstrates that the system’s architecture, interconnect, and software components are harmonized to maximize performance gains as resources grow, avoiding diminishing returns common in large-scale deployments. If adding GPUs yielded only marginal improvements, the infrastructure costs would quickly outweigh the performance benefits, a challenge NVIDIA appears to have overcome.

Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0, highlighting the importance of continuous refinement beyond hardware advancements, the company states. These gains stemmed from techniques including lower KV cache precision, additional kernel fusion, and improved kernels, alongside disaggregated serving utilizing vLLM and NVIDIA Dynamo.

These optimizations extended beyond the official v6.1 submission deadline, with preliminary, unverified results on GPT-OSS-120B and DLRMv3 indicating further performance improvements are still being realized. “Performance, scaling efficiency and software velocity are important considerations for determining long-term inference economics,” according to NVIDIA, emphasizing a comprehensive approach to AI infrastructure development. The Vera Rubin NVL72 platform’s debut in the MLPerf Inference v6.1 suite showcased substantial gains on demanding benchmarks, delivering up to 3. These early results, NVIDIA asserts, demonstrate an accelerated pace of innovation and a trajectory of continued improvement through ongoing software development.

Beyond large-scale deployments, NVIDIA also submitted results from the Jetson AGX Thor, employing NVIDIA TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3. 6-27B, extending AI inference capabilities to edge devices. The company’s commitment to software development is evident in the continuous delivery of performance and feature improvements, positioning it as a key player in the evolving landscape of AI infrastructure.

Vera Rubin NVL72 Delivers 30x Advantage on SemiAnalysis AgentX Benchmark

The NVIDIA Vera Rubin NVL72 system achieved up to 30 times better performance than the GB300 NVL72 on the SemiAnalysis AgentX benchmark during preview testing, demonstrating a substantial leap in AI agent capabilities. This result, revealed alongside MLPerf Inference v6.1 results, highlights a focus on metrics that capture the evolving demands of AI, moving beyond simple throughput to assess reasoning, planning, and action. Nebius independently validated these gains, also submitting preview results with the Vera Rubin NVL72 and achieving strong performance, indicating broad applicability of the platform’s advancements.

NVIDIA attributes this performance boost to a full-stack approach, integrating hardware and software optimizations designed to maximize efficiency across all stages of AI inference, by the company’s account. The adoption of NVFP4 precision reduces the memory footprint of model weights, attention mechanisms, and the key-value cache, allowing for increased throughput without significant loss of output quality.

This careful balance between performance and resource utilization is particularly important for cost-effective scaling of AI services. Disaggregated serving, a technique separating prefill and decode operations, played a key role in the Vera Rubin NVL72’s success, alongside large-scale expert parallelism utilized within mixture-of-experts layers powering models like DeepSeek-R1 and Qwen3-VL.

NVIDIA, headquartered in Santa Clara, United States, is a public company listed on Nasdaq as NVDA. Established in 1993, it has become a significant player in advanced computing, extending beyond traditional graphics processing to include quantum computing services and infrastructure through its CUDA Quantum platform and DGX Quantum systems. NVIDIA maintains partnerships with companies including Quantinuum, QuEra Computing, and IonQ, collaborating on the development of hybrid quantum-classical workflows. These collaborations are supported by software tools like cudaq-algorithms, accelerated algorithm libraries launched in May 2026, and the open NVQLink architecture, introduced in January 2026, designed for low-latency integration of quantum processors with GPU supercomputers.

Recent activity demonstrates NVIDIA’s focus on optimising performance across both classical and quantum systems. In September 2026, the company reported a quantum calibration model functioning across six qubit modalities and charted a new path for enterprise quantum computing. Simultaneously, NVIDIA has continued to refine its classical computing infrastructure, as evidenced by reports of innovations in AI power efficiency and error resilience, including NVLink 6. The ten patent families and four publications registered in the last twelve months reflect ongoing research and development efforts. This sustained innovation builds on a 2023 integration of the Classiq quantum software platform with NVIDIA’s CUDA-Q.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Ivy Delaney

Ivy Delaney

Ivy Delaney has been working with neural networks and machine learning since the mid-nineties, back when a couple of hidden layers and a long afternoon of training counted as ambitious. She has watched the field go from academic curiosity to the thing quietly running underneath everything, and she brings that long view to quantum computing. For Quantum Zeitgeist she covers the ground where the two fields meet. That means quantum machine learning and the variational algorithms it leans on, and it also means the less glamorous but more interesting story of classical machine learning already doing real work inside quantum machines, decoding error-correcting codes, calibrating noisy hardware and learning the error models that simulators depend on. She writes about the hardware those algorithms have to run on too, and about the post-quantum cryptography scramble that the same hardware has set off. Her stories typically start with the paper, whether that is peer-reviewed work, conference proceedings or an arXiv preprint, with the source linked so you can hold a claim up against the research it came from. She is unimpressed by benchmarks that will not say what they beat, and by demonstrations that only work in the press release.

Latest Posts by Ivy Delaney: