NVIDIA Vera Rubin boosts AI power efficiency with Groq 3 LPX

NVIDIA is boosting the efficiency of its Vera Rubin AI platform with the addition of the Groq 3 LPX low-latency accelerator, and software innovations that allow operators to recover stranded power and provision up to 40% more GPUs within the same site-power envelope. The NVIDIA DSX MaxLPS software shifts power between racks as workloads demand shifts, allowing an operator to recover stranded power and provision up to 40% more GPUs within the same site-power envelope and deliver 35% higher token throughput.

These factory and rack-level power management techniques, alongside the deterministic execution model of the Groq 3 LPX, are designed to maximize performance per watt, the ultimate measure of an AI platform’s value, as AI workloads increasingly strain power budgets. The platform’s innovations allow prediction of electrical current draw, enabling technologies like Preemptive Power and Clock Period Synthesis to reduce electrical safety margins and dedicate more power to the AI workload itself.

NVIDIA Vera Rubin Achieves Efficiency with DSX MaxLPS Power Management

This capability addresses a critical constraint for AI factories, power, by maximizing output within a fixed budget, making performance per watt the primary metric for platform value. The Vera Rubin platform, centered around the NVL72, aims to deliver this efficiency across a range of AI compute needs, from high-volume throughput to the smaller batches required for interactive applications and both open and closed models. Beyond software-driven power allocation, rack-level innovations further refine Vera Rubin’s efficiency.

Inside each NVL72 rack, capacitors coupled with Intelligent Power Smoothing software absorb the rapid power fluctuations inherent in both training and inference workloads. This allows operators to plan around sustained power demand rather than peak spikes, optimizing resource allocation and reducing waste. The result is a more deployable compute capacity for each megawatt of power consumed, which is a significant advantage as AI model complexity and scale continue to increase.

The highest tiers of AI interactivity present unique challenges, demanding exceptionally low latency. This addition isn’t simply about adding more compute; it’s about fundamentally changing how that compute is delivered. The LPU compiler generates a precise schedule dictating when data moves to specific compute units and when operations execute, all down to the clock cycle, extending this control across all 256 LPU chips within the rack.

“After the LPU compiler has generated the workload execution schedule, it can predict the electrical current draw for each cycle in that schedule,” explains the NVIDIA team. This predictability unlocks two complementary technologies: Preemptive Power (PEP) and Clock Period Synthesis (CPS).

PEP prepares the power-delivery system for anticipated changes in demand before they occur, while CPS shapes the abruptness of those rises and falls. Together, these techniques reduce the electrical safety margin continually provided to the chips, allowing a greater proportion of available power to be directed to the AI workload itself. Groq 3 LPX determinism, therefore, isn’t just about speed; it’s about fundamentally optimizing power utilization, the company says.

The system delivering current to the chips must be designed to handle the demands of AI workloads, which involve high transistor switching rates and, consequently, rapid fluctuations in current demand on a nanosecond scale. Decoupling capacitors, positioned close to the chips, provide an immediate current source, but their effectiveness depends on anticipating and mitigating voltage drops. PEP commands the power delivery network (PDN) to adjust the supplied voltage based on the compiler’s predicted current spikes.

By initiating this change early, the decoupling capacitors have less ground to make up, minimizing voltage drops when current ramps up. NVIDIA estimates that these technologies will lead to a high single-digit percentage decrease in the baseline voltage required by the electrical system, translating into even larger percentage decreases in overall power consumption, all without impacting workload performance. “A smaller voltage guardband lowers the supplied voltage, and the power saved scales with the square of that reduction on every cycle,” the team notes, highlighting the significant impact of even small voltage adjustments.

Compared to a similarly specified, nondeterministic system, Groq 3 LPX deterministic execution can potentially reduce the power required to run the same workload by a low-double-digit percentage. In a power-constrained AI factory, this reduction frees up valuable power budget for increased token production.

Groq 3 LPX Deterministic Execution Enables Cycle-Level Power Prediction

The NVIDIA Groq 3 LPX rack employs a deterministic execution model that forecasts electrical current draw cycle by cycle, a capability enabling proactive power management and significant efficiency gains. Individual LPUs utilize specialized hardware, MXM for matrix multiplication, VXM for vector operations, and SXM for transposing and reshaping, synchronized to the clock cycle, further enhancing predictability and speed. This pre-scheduled approach resolves resource conflicts, such as simultaneous writes to the same memory bank, before execution even begins, a stark contrast to dynamically scheduled accelerators.

The deterministic nature of Groq 3 LPX is therefore not merely a performance enhancement but a fundamental enabler of energy efficiency, according to the company. CPS, in particular, uses the rack’s plesiosynchronous clock system, maintaining synchronization between chips and compute elements, to fine-tune current spikes by lengthening individual clock cycles, effectively lowering the rate of current increase.

Voltage Droop Mitigation via Groq 3 LPX Preemptive Power & CPS

These innovations aren’t simply about minimizing wasted energy; they fundamentally alter how power is managed within AI infrastructure, allowing for denser deployments. Preemptive Power operates by commanding the power delivery network, the physical electronics supplying the LPUs, to anticipate voltage changes before demand spikes. Clock Period Synthesis further refines this control by shaping the rate at which demand rises and falls, smoothing the current waveform and reducing the peak di/dt, or rate of change of current.

This collaborative approach directly addresses the challenge of voltage droop, a phenomenon where rapid power draw discharges capacitors and temporarily lowers voltage, potentially causing errors if the voltage falls below a minimum operating threshold. Maintaining a stable voltage is critical because of the power-hungry nature of AI workloads, particularly matrix and vector multiplications. These operations involve intense transistor switching, creating a high and rapidly fluctuating current demand on a nanosecond timescale.

A larger di/dt necessitates larger capacitors, increasing system size and cost, while a slower rate of change allows for smaller, more efficient components. “Decreasing this voltage guardband means a greater proportion of scarce power can be spent on the workload,” explains the research team. The Groq 3 LPX system, through its deterministic execution model, effectively shrinks the necessary voltage guardband, freeing up power for actual computation.

The benefits of this approach extend beyond simply reducing wasted energy; it directly impacts the number of GPUs that can be deployed. This is achieved without compromising performance, as the reduced voltage guardband doesn’t impact the workload itself. This improvement is particularly important for demanding operating points where power is a limiting factor.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Ivy Delaney

Ivy Delaney

Ivy Delaney has been working with neural networks and machine learning since the mid-nineties, back when a couple of hidden layers and a long afternoon of training counted as ambitious. She has watched the field go from academic curiosity to the thing quietly running underneath everything, and she brings that long view to quantum computing. For Quantum Zeitgeist she covers the ground where the two fields meet. That means quantum machine learning and the variational algorithms it leans on, and it also means the less glamorous but more interesting story of classical machine learning already doing real work inside quantum machines, decoding error-correcting codes, calibrating noisy hardware and learning the error models that simulators depend on. She writes about the hardware those algorithms have to run on too, and about the post-quantum cryptography scramble that the same hardware has set off. Her stories typically start with the paper, whether that is peer-reviewed work, conference proceedings or an arXiv preprint, with the source linked so you can hold a claim up against the research it came from. She is unimpressed by benchmarks that will not say what they beat, and by demonstrations that only work in the press release.

Latest Posts by Ivy Delaney: