On a sweltering August evening in Silicon Valley, Silicon Valley Power successfully signaled an AI factory to adjust its power consumption, a first-of-its-kind interaction with a large-scale AI workload. Silicon Valley Power has since sent more than 200 demand signals, including the initial test signal, each successfully adjusting power without interrupting critical operations.
Emerald AI’s Conductor platform, an NVIDIA partner’s grid-orchestration technology, facilitated these adjustments across thousands of NVIDIA GPUs without manual intervention. “We were watching with bated breath,” said Varun Sivaram of Emerald AI, “It was our first time deploying across thousands of NVIDIA GPUs.”
Emerald AI & Silicon Valley Power Demonstrate Automated Demand Response
NVIDIA’s projections indicate DSX MaxLPS can enable up to 40% more GPU capacity for Vera Rubin NVL72 AI factories within the same megawatt power budget, contingent on suitable deployment environments. This increase stems from a demonstrated ability to dynamically adjust power consumption in large-scale AI workloads. The utility successfully signaled an AI factory to reduce its electricity demand from four to three megawatts, a feat accomplished automatically without interrupting critical AI operations.
Mansi Shah, head of product at Emerald AI, described the event as “This feels kind of like a SpaceX rocket launch,” highlighting the significance of the successful demonstration. Silicon Valley Power has since sent more than 200 demand signals to that AI factory, each one successfully modulating power usage, a reliability that moves beyond a single successful test.
The Emerald AI Conductor platform, integrating into Silicon Valley Power’s Flexible Load Interconnect Program, responded to each signal in under a minute, showcasing rapid grid flexibility in a production setting. This program marks the first commercial grid utility initiative designed to treat AI factories as dispatchable resources, allowing them to contribute to grid stability. NVIDIA DSX MaxLPS enables reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint, according to the company.
The Santa Clara demonstration validates the core concept at commercial scale. With this proof of concept, NVIDIA believes it has moved beyond the limitation of fixed power budgets, enabling AI factories to adapt to grid demands, the company says.
“A one-gigawatt factory will never become a two-gigawatt factory,” a spokesperson explained, emphasizing the importance of optimizing existing infrastructure rather than simply building more. The success in Santa Clara is now informing the first dedicated DSX Flex commercial deployment planned for a 96-megawatt Vera Rubin AI factory in Manassas, Virginia, building on five prior demonstrations across two continents.
With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets.
Dave Ward, president of cloud services at Lambda
Lambda Validates 24% Token Throughput Gain with DSX MaxLPS
Lambda’s initial validation of NVIDIA’s DSX MaxLPS in a deployment environment, released at the AI Infra Summit, demonstrates a 24% increase in token throughput within a fixed power budget, a result revealed at the AI Infra Summit. This performance gain, achieved on NVIDIA HGX B20 0 GPU Servers, signifies a substantial step toward optimizing resource utilization in AI factories without requiring immediate infrastructure expansion, according to the company.
The cloud provider, serving over 10,000 customers, ran the software on a 19-node cluster, effectively increasing computational density within the same physical footprint. DSX MaxLPS operates by dynamically monitoring GPU and rack-level power consumption, then reallocates available headroom based on workload priorities. This intelligent allocation addresses the differing power demands of training and inference tasks, recovering capacity often left unused by static provisioning methods.
Lambda’s tests showed a jump from approximately 4 million tokens per second to 5 million, while also improving performance per watt by 23%. These collaborations, alongside recent patent filings in the field, demonstrate a commitment to pushing the boundaries of AI and high-performance computing.
NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.
NVIDIA DSX Optimizes Power Allocation for Increased GPU Capacity
Training and inference tasks exhibit differing power profiles, and MaxLPS optimizes allocation within AI factories running both simultaneously, maximizing resource utilization. Beyond software optimization, NVIDIA is also addressing power delivery with an 800V DC power architecture, designed to reduce conversion complexity and improve efficiency for denser accelerated computing racks, the firm reports.
This architecture is being incorporated into NVIDIA DSX reference designs, signaling a broader commitment to optimizing the entire power infrastructure supporting AI workloads. NVIDIA was founded in 1993 and currently listed on Nasdaq as NVDA, has also been actively filing patents, with 10 families and 4 publications in the last twelve months, and forging partnerships with companies like Quantinuum and IQM to further advance the field, including recent work integrating quantum processors with GPU supercomputers through NVQLink and CUDA-Q.
We were watching with bated breath.
Varun Sivaram, Emerald AI
800V DC Architecture Enables Denser Accelerated Computing Racks
This architecture moves beyond traditional lower-voltage systems which introduce limitations as power demands escalate, enabling support for denser configurations and minimizing energy loss during conversion processes. The shift to 800V DC is not an isolated improvement, but part of a comprehensive approach to AI factory optimization. Traditional data center power delivery relies on multiple voltage conversions, each introducing inefficiencies and potential bottlenecks.
By increasing the voltage level, NVIDIA aims to reduce the number of conversions required, streamlining the power path and minimizing wasted energy. This is particularly important as racks become increasingly populated with high-powered GPUs like the HGX B20 0 GPU Servers, demanding substantial and reliable power delivery.
A faster GPU is ineffective if constrained by network bandwidth or insufficient cooling, and power can be wasted through inefficient provisioning. The DSX suite, including DSX Sim, DSX OS, DSX Exchange, and DSX Reference Designs, is intended to address these interconnected challenges, providing tools for simulation, operation, and validated architectural blueprints.
“No single component can optimize an AI factory on its own,” the company asserts, highlighting the need for a coordinated strategy. This integrated approach, coupled with the 800V DC architecture, aims to establish a new standard for AI factory efficiency, measured by useful work produced per megawatt consumed.
A one-gigawatt factory will never become a two-gigawatt factory.
Source: https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
