NVIDIA’s GB300 NVL72 Achieves 1,648 TFLOPs with DeepSeek-V3

NVIDIA’s GB300 NVL72 has achieved 1,648 TFLOPs per GPU during pre-training of the DeepSeek-V3 671B model, setting a new performance record and demonstrating gains in AI training capability. Software innovations, including NVIDIA Megatron Core, TorchTitan, and JAX, have delivered three to ten times performance improvements over previous generations, with per-GPU throughput holding above 97% when scaling from 256 to 1,024 GPUs. This demonstrates the effectiveness of NVIDIA’s continuous hardware-software co-design and open-source contributions. The performance relies on advances across the entire AI platform, from silicon to networking to software, pushing training performance forward as frontier models increasingly use mixture of experts (MoE) architectures. Unlike dense models, MoE models activate only a subset of parameters per token, reaching frontier scale at a lower per-token cost; however, this introduces communication challenges as experts reside on different GPUs.

MoE Architecture Drives Demand for Communication Efficiency

The shift toward mixture of experts (MoE) models is reshaping bottlenecks in large-scale artificial intelligence training, with communication efficiency now the primary challenge. This efficiency introduces communication overhead, as tokens must be dispatched to and results gathered from experts residing on different GPUs. Every MoE layer necessitates an all-to-all communication pattern in both forward and backward passes, creating a critical path where delays quickly compound. NVIDIA explains that “Because it occurs at every layer in every training step, small delays compound until the all-to-all can no longer be hidden behind compute, and adding GPUs no longer increases throughput.” This demands a tightly coupled scale-up domain where GPUs can communicate with each other via a non-blocking fabric offering high bandwidth and low latency, alongside a robust scale-out network to link multiple domains.

The challenge is two-tiered; success is measured in delivered FLOPs, not peak FLOPs. NVIDIA’s GB300 NVL72 addresses this dual challenge through extreme co-design, integrating silicon, interconnect, networking, and software into a unified platform. Central to this is fifth-generation NVLink, providing each of the 72 Blackwell Ultra GPUs with 1.8 TB/s of bandwidth and 130 TB/s of non-blocking, all-to-all bandwidth within the rack. This memory-semantic fabric allows GPUs to directly read and write each other’s high-bandwidth memory, bypassing software overhead and minimizing latency. Beyond the rack, NVIDIA ConnectX-8 SuperNICs at 800 Gbps per GPU, coupled with either Quantum-X800 InfiniBand or Spectrum-X Ethernet, maintain predictable performance during scale-out. Software innovations further amplify these hardware gains.

NVIDIA GB300 NVL72 Enables Record-Breaking DeepSeek-V3 Pre-Training

The current landscape of artificial intelligence model training is increasingly defined by the architecture of mixture of experts (MoE), a shift demanding new approaches to computational scaling. Unlike traditional dense models where every parameter is activated by each input, MoE models selectively activate subsets of parameters, allowing for larger models with manageable per-token costs. Successfully training models like DeepSeek-V3 671B, with 671 billion parameters but only approximately 37 billion activated per token, requires overcoming this communication challenge. This performance is not solely attributable to faster silicon; it’s a result of a holistic, co-designed platform integrating compute, interconnect, networking, and software. The system’s core is built around 72 NVIDIA Blackwell Ultra GPUs linked by fifth-generation NVLink, providing each GPU with 1.8 TB/s of bandwidth. NVIDIA explains that “A shortfall in any one caps the whole,” highlighting the importance of a balanced system.

This scale-out networking is crucial for models exceeding the capacity of a single rack, ensuring that added GPUs translate into increased throughput rather than communication overhead. Software innovations, including NVIDIA Megatron Core, TorchTitan, and JAX, further amplify these gains. The continuous optimization of these frameworks, with nearly ten times performance on DeepSeek-V3 671B at 256 GPU scale, all from software optimizations, demonstrates that performance gains extend beyond initial silicon capabilities. This sustained performance is not merely a matter of faster hardware, but of a system designed to deliver “delivered FLOPs, not peak FLOPs,” according to NVIDIA. The GB300 NVL72’s architecture prioritizes maintaining high throughput throughout the entire training process, ensuring that communication delays do not negate the benefits of increased compute power. The record-setting performance achieved with DeepSeek-V3 671B is therefore not just a benchmark, but a foundation for even greater advancements in AI model training.

Software Innovations Yield 3x-10x Performance Gains with GB300

The demand for increasingly large artificial intelligence models is driving innovation not just in hardware, but also in the software that unlocks its potential. Recent results with the NVIDIA GB300 NVL72 demonstrate this synergy, achieving substantial performance gains, between three and ten times, over previous generations through optimized software stacks. A key element of this advancement is the efficient scaling of mixture of experts (MoE) architectures. This efficiency introduces a communication bottleneck; experts reside on different GPUs, requiring all-to-all communication for each layer and training step. The company reports that fifth-generation NVLink gives each GPU 1.8 TB/s of bandwidth. NVIDIA’s contributions to open-source frameworks like Megatron Core, TorchTitan, and JAX are central to these gains. This represents approximately a three-fold increase in delivered throughput per GPU in a single generation. Further optimization through TorchTitan yielded approximately six times higher delivered performance on the same infrastructure.

JAX saw even more dramatic improvements, with NVIDIA JAX optimizations lifting performance by nearly ten times on DeepSeek-V3 671B at 256 GPU scale, all from software optimizations. The latest software version reaches an exceptional performance throughput of 1,025 TFLOPS/GPU, and software optimizations continue to evolve from there. Notably, this performance isn’t limited to smaller scales. Scaling DeepSeek-V3 671B pre-training from 256 to 1,024 GPUs, per-GPU throughput holding above 97%, demonstrates the efficiency of the system. This means nearly all the added infrastructure translates into increased system-level tokens per second throughput. As NVIDIA explains, “The open frameworks tell the same story. Performance continues to improve on the same platform as the software evolves.” These results are not the ceiling, but rather a foundation for even higher performance in the future.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of The Quant

The Quant

The Quant possesses over two decades of experience in start-up ventures and financial arenas, brings a unique and insightful perspective to the quantum computing sector. This extensive background combines the agility and innovation typical of start-up environments with the rigor and analytical depth required in finance. Such a blend of skills is particularly valuable in understanding and navigating the complex, rapidly evolving landscape of quantum computing and quantum technology marketplaces. The quantum technology marketplace is burgeoning, with immense growth potential. This expansion is not just limited to the technology itself but extends to a wide array of applications in different industries, including finance, healthcare, logistics, and more.

Latest Posts by The Quant: