NVIDIA’s DSX MaxLPS fits 40% more GPUs in the same power budget

NVIDIA is challenging traditional AI factory design with DSX MaxLPS, a system demonstrated to deploy up to 40% more GPUs within the same power budget, the company says. The approach moves beyond static power provisioning by dynamically reallocating available energy across resources, a capability validated in testing with Kimi K2.5 workloads on NVIDIA GB200 NVL72 systems. This evaluation, conducted at Nscale’s data center in Keflavík, Iceland, used entirely renewable energy, and showed how DSX MaxLPS combines chip, system, and software technologies to maximize AI output within existing infrastructure constraints.

NVIDIA DSX MaxLPS Enables 40% More GPUs Within Fixed Power Budgets

NVIDIA’s DSX MaxLPS enabled a 40% increase in deployed GPUs within a fixed power budget during testing, a result achieved by dynamically reallocating energy across resources. The evaluation, conducted jointly with Nscale, demonstrated the ability to operate 192 GPUs, compared to a baseline of 140, while remaining within a provisioned power limit of 264.4 kW. This approach moves beyond traditional static power provisioning, which often leaves capacity unused to accommodate peak demand scenarios.

The core of DSX MaxLPS is a control loop combining chip, system, thermal and software technologies, all operating within established land, power and shell constraints. Dynamic Power Software is the control layer, organizing nodes into managed groups with aggregate power budgets, but does not increase the overall site power supply. Operators define these limits, and the system then intelligently distributes available energy where it’s most needed, maximizing output without exceeding the pre-approved electrical boundaries.

Measurements from the Nscale deployment revealed a 49.2% increase in aggregate throughput, reaching 1,618,443 tokens per second, when using DSX MaxLPS compared to the static baseline of 1,084,503 tokens per second. Per-instance throughput for both high-throughput and low-latency workloads remained consistent, indicating that the increased capacity didn’t come at the expense of existing performance levels. The baseline configuration used 35 four-GPU nodes, while the DSX MaxLPS configuration expanded to 48 four-GPU nodes, adding a third high-throughput instance while maintaining the same low-latency instance.

This scaling demonstrates the system’s ability to accommodate increased demand without requiring infrastructure upgrades, according to NVIDIA. The evaluation used NVIDIA GB200 NVL72 systems, running Kimi K2.5 with NVIDIA Dynamo and NVIDIA TensorRT LLM. The workload mix intentionally combined high-throughput and low-latency inference instances, creating distinct power and service profiles within the managed group.

The resulting data showed a 12.3 percentage point increase in power-budget use, rising from 62.9% in the static baseline to 75.2% with DSX MaxLPS. This improved use translated directly into a 49.2% increase in throughput per provisioned watt, climbing from 4.10 tokens/s/W to 6.12 tokens/s/W. NVIDIA emphasizes that the 264.4 kW provisioned power budget remained constant throughout the evaluation, meaning the gains in throughput per watt are a direct result of more efficient power allocation.

“DSX MaxLPS gives operators a framework for turning workload variability into managed capacity,” the company states in accompanying documentation. The engineering effort, they explain, centers on defining operational boundaries, measuring representative behavior, tuning policies against service objectives, and proving compliance under both normal and adverse conditions. This validation process is essential for ensuring reliable and predictable performance in a production environment.

NVIDIA notes that electrical distribution, cooling systems, network fabric, floor space, and rack positions must all be sized for the validated lifecycle target, even if fewer racks are initially populated. This forward-looking approach ensures scalability and avoids costly retrofits as demand increases. The company highlights that DSX MaxLPS will also combine dynamic power management with performance-per-watt techniques and infrastructure designed for 45°C liquid-cooling inlet operation in future NVIDIA Vera Rubin NVL72 AI factories.

However, NVIDIA clarifies that any Vera Rubin capacity projections should be considered separate from the GB200 NVL72 evaluation results. The evaluation methodology itself provides a repeatable validation sequence for operators seeking to implement DSX MaxLPS in their own AI factories. By following the documented steps, organizations can establish production operating limits tailored to their specific workloads and infrastructure.

This standardized approach facilitates wider adoption and ensures consistent performance across different deployments. “Review NVIDIA DSX MaxLPS and NVIDIA Dynamic Power Software documentation, then use this validation sequence to establish production operating limits for your own AI factory,” the company advises, signaling a commitment to transparency and collaboration within the AI community. The findings highlight a shift toward more intelligent and adaptive power management strategies in AI infrastructure, promising significant gains in efficiency and scalability.

Static Power Provisioning Limits AI Factory GPU Utilization

The evaluation of NVIDIA’s DSX MaxLPS revealed a 40% increase in GPU density. This gain in GPU density was not simply a matter of adding hardware. The system demonstrably increased aggregate throughput by 49.2%, from 1.085 million to 1.618 million tokens per second, demonstrating a more efficient use of existing infrastructure. This approach contrasts with static provisioning, where reserved power often remains unused, limiting overall capacity and hindering scalability.

The system meticulously collects power telemetry from GPUs, nodes, racks and groups, identifying available headroom and responding to emerging power events at intervals sufficient for effective management. Operators define rules establishing node and group limits, allocation priorities and reserve requirements, allowing for customized control over the AI factory’s power distribution.

When resources draw less than their allocated power, the software intelligently adjusts GPU power limits, enabling other resources to use the available capacity, a process that requires reliable telemetry to function effectively, the company says. Site-level telemetry was used during the evaluation to verify the accuracy of rack-level power measurements, highlighting the importance of accurate data for dependable fleet-level decisions.

The evaluation’s workload mix was carefully constructed to mimic the diverse demands of a production AI environment, combining two 52-GPU high-throughput instances with one 36-GPU low-latency instance, then adding a third high-throughput instance with DSX MaxLPS. This configuration allowed for a direct comparison of performance and power consumption between the static baseline and the dynamic allocation system.

The results showed that while mean GPU power increased by 35.9%, from 97.0 kW to 131.8 kW, total measured power rose by 19.7%, from 166.2 kW to 198.9 kW, demonstrating that the increased throughput was not simply achieved by consuming more energy. Power-budget use also improved by 12.3 percentage points, rising from 62.9% to 75.2%, indicating a more efficient use of the available power infrastructure.

However, the evaluation also highlighted the inherent trade-offs involved in dynamic power allocation. While DSX MaxLPS effectively uses previously unused headroom, adding capacity inevitably increases average power use, requiring careful consideration of overall energy consumption. Workload mix influences the available headroom, and operators must define production acceptance criteria as part of the validation process.

Before deploying DSX MaxLPS, operators should confirm that increased throughput does not compromise service quality or compliance with power limits, establishing a representative baseline under static provisioning to accurately measure performance and power consumption. To facilitate wider adoption, NVIDIA recommends a staged validation process with explicit boundaries and acceptance criteria. This process begins with defining the managed boundary, mapping the utility, distribution, rack, node, and GPU topology, and setting the resource-group budget, reserve requirements and escalation behavior.

Operators must confirm which measurement represents the enforceable limit, ensuring that the system operates within established power constraints. The findings emphasise that dynamic power management is not a panacea, but a sophisticated tool that requires careful planning, implementation and validation.

The ability to observe and control engineering trade-offs at the fleet level offers significant advantages, but it also demands a thorough understanding of the underlying system and the specific characteristics of the workload. NVIDIA emphasizes that the validation sequence detailed in their documentation is essential for establishing production operating limits tailored to individual AI factories, enabling organizations to maximize throughput and efficiency while remaining within their power constraints.

DSX MaxLPS Control Loop: Topology, Telemetry, and Dynamic Allocation

DSX MaxLPS uses a layered approach to infrastructure management, beginning with a detailed mapping of resources and their organization into managed groups with defined power budgets. This topology and resource grouping forms the foundation for dynamic power allocation, allowing the system to intelligently redistribute energy based on real-time demand and availability. Operators define these groups, establishing node limits, overall group limits, allocation priorities and pre-defined responses to both routine maintenance and unexpected emergency events.

The system’s core functionality relies on continuous telemetry collection, monitoring GPU, node, rack and group power consumption at intervals designed to capture both available headroom and the onset of potential power events. When individual resources operate below their allocated capacity, NVIDIA’s Dynamic Power Software adjusts power limits for participating GPUs, directing unused energy to resources that can benefit from it.

This redistribution is not a net increase in power draw, but rather a more efficient use of existing resources, ensuring that the total consumption remains within the operator-defined budget. The evaluation demonstrated a 40% increase in normalized aggregate throughput, climbing from 1.618 million tokens per second.

This performance gain was achieved within the same 264.4 kW provisioned power budget, resulting in a substantial improvement in power efficiency, increasing from 4.10 to 6.12 tokens per second per watt. This ability to maximize output from a fixed power supply is particularly relevant given the increasing energy demands of large-scale AI deployments and the growing emphasis on sustainable computing practices.

NVIDIA, founded in 1993 and headquartered in Santa Clara, has expanded beyond traditional graphics processing to become a key player in AI infrastructure, including hybrid quantum-classical computing with platforms like CUDA-Q and cuQuantum. However, dynamic power allocation introduces engineering trade-offs that must be carefully managed. The available headroom is directly influenced by the workload mix, meaning that the system’s effectiveness can vary depending on the types of tasks being performed.

DSX MaxLPS makes these trade-offs visible and controllable at a fleet level, but does not eliminate them entirely. Operators must then confirm which measurement represents the enforceable power limit, establishing a baseline performance level under static provisioning before introducing dynamic allocation.

This baseline is established by running the intended AI factory workload mix under static provisioning, measuring performance and power consumption over a period sufficient to capture workload variations and ensure repeatability. Before adding capacity, operators should confirm telemetry accuracy, control response and budget compliance. Incremental increases in the managed population allow for comparison of aggregate and per-instance performance at each step, with ongoing verification of service behavior under peak demand, operating transitions, telemetry failures and simulated power reductions.

A configuration is only approved when it meets throughput and latency objectives, remains within the defined budget, maintains the required reserve capacity, and exhibits predictable behavior during fault conditions and transitions. The evaluation demonstrated that DSX MaxLPS can effectively reclaim previously unused capacity within a power-constrained AI factory, offering a pathway to increased efficiency and output.

Kimi K2.5 Workloads Demonstrate 49.2% Throughput Gain with DSX MaxLPS

Kimi K2.5 workloads achieved a 49.2% throughput gain when deployed with NVIDIA’s DSX MaxLPS, a result demonstrated in a joint evaluation with Nscale at its data center in Keflavík, Iceland. The increase, measured in tokens per second, signifies a substantial boost in processing capacity without expanding the overall power budget, a critical consideration for rapidly scaling AI infrastructure.

The testing environment was deliberately configured with four racks, distributing workloads within a single rack to isolate performance variables and ensure accurate comparisons between static provisioning and the dynamic allocation offered by DSX MaxLPS. The evaluation revealed that DSX MaxLPS enabled the deployment of 40% more GPUs, increasing the fleet from 140 to 192, within the same 264.4 kW provisioned power budget.

Static power planning, the traditional approach, reserves capacity for peak demand, often leaving substantial headroom unused during typical operation. DSX MaxLPS, however, dynamically reallocates power based on workload needs, effectively reclaiming this previously underutilized capacity. “Every unused watt is capacity left on the table,” emphasises the core principle driving this approach.

Throughput per provisioned watt, a key metric for assessing efficiency, increased from 4.10 tokens/s/W with static provisioning to 6.12 tokens/s/W with DSX MaxLPS, mirroring the 49.2% gain in aggregate throughput. This indicates that the system doesn’t simply add capacity but does so with a significant improvement in performance per unit of power.

The team meticulously tracked not only aggregate throughput but also latency metrics, recognizing the importance of maintaining responsiveness alongside increased capacity. Median and P75 latency remained within 5% of the baseline, indicating minimal impact on typical response times. P99 time to first token, a measure of tail latency, representing the experience of the slowest 1% of requests, increased by 17% from 15.7 seconds.

This suggests that while the majority of requests experienced similar latency, the most demanding requests took longer to process. The workload mix influences the available headroom and the balance between capacity and performance.

NVIDIA’s commitment to sustainable AI infrastructure is highlighted by the choice of Nscale’s Iceland data center, powered entirely by renewable energy. This aligns with a broader industry trend toward environmentally responsible computing, as well as the company’s own investments in technologies like the Blackwell architecture and its focus on performance-per-watt improvements. Beyond the GB200 NVL72 systems used in this evaluation, NVIDIA plans to integrate DSX MaxLPS with future Vera Rubin NVL72 AI factories, further enhancing their efficiency and scalability.

The company’s broader portfolio includes software and hardware for hybrid quantum-classical computing, such as the CUDA-Q platform and NVQLink interconnect, demonstrating a long-term vision for integrating diverse computing paradigms, NVIDIA reports. The findings from Nscale demonstrate that DSX MaxLPS is a theoretical optimization and a practical solution for maximizing the use of AI infrastructure. By dynamically allocating power and reclaiming previously unused capacity, it allows operators to deploy more GPUs within existing power constraints, driving significant gains in throughput without increasing energy consumption.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Rusty Flint

Rusty Flint

Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things Quantum Physics especially. Rusty likes the more esoteric side of Quantum Computing and the Quantum world. Everything from Quantum Entanglement to Quantum Physics. Rusty thinks that we are in the 1950s quantum equivalent of the classical computing world. While other quantum journalists focus on IBM's latest chip or which startup just raised $50 million, Rusty's over here writing 3,000-word deep dives on whether quantum entanglement might explain why you sometimes think about someone right before they text you. (Spoiler: it doesn't, but the exploration is fascinating)

Latest Posts by Rusty Flint: