NVIDIA GPUs help AI factories get the most tokens per megawatt

With each AI factory now costing approximately $60 million to build, operators are intensely focused on return on investment. NVIDIA is engineering AI factories to maximize earnings, longevity and versatility, recognizing that strength in one area cannot compensate for weakness in another. The company achieves this through a platform designed to deliver the highest throughput per megawatt, extend hardware’s useful life and run any accelerated workload; NVIDIA CUDA-X libraries let a factory run any accelerated workload, the company says. This approach maximizes AI factory returns.

NVIDIA Vera Rubin NVL72 Boosts Tokens Per Megawatt

This performance leap directly impacts the earning capacity of AI factories, as tokens per second per megawatt govern revenue generation. Lowering the cost per token further increases profit margins, making the Vera Rubin NVL72 a compelling option for operators seeking to maximize returns on their substantial investments. Gains of this magnitude stem from an extreme codesign approach, optimizing models, workloads, software, compute, networking, and memory in unison.

Efficiency extends beyond raw speed; NVIDIA is also focused on extending the useful life of its hardware. Continuous software optimization ensures that installed NVIDIA systems remain productive for years after their initial deployment, a strategy that addresses the longevity component of AI factory ROI, according to the company. The NVIDIA A100 GPU shipped in 2020 and is still in commercial service six years later, demonstrating its continued economic value; CoreWeave recently extended bookings for units first introduced in 2020 through 2029.

Over the years, every major operator has extended the depreciation schedule on its servers, which is an estimate of when hardware stops earning, and one that keeps moving out. A September 2026 Sprout analysis, “ The Productive Life of a Data Center GPU,” tracks how that schedule has shifted across every major operator. Beyond performance and durability, NVIDIA is engineering AI factories to broaden the scope of potential demand.

NVIDIA CUDA-X libraries enable these facilities to run any accelerated workload, not solely AI-related tasks, the firm reports. This versatility deepens and broadens the demand they can serve, moving beyond a reliance on a single application. SemiAnalysis AgentX analysis indicates that even as each generation of hardware becomes more efficient, demand for compute power does not diminish. Cheaper tokens unlock new use cases that consume even more resources.

This suggests a virtuous cycle where increased efficiency fuels further innovation and demand, rather than leading to contraction. The company’s strategy centers on achieving a balance between productivity, durability and fungibility, all aimed at maximizing return on investment, NVIDIA reports.

“NVIDIA AI factories are engineered to be productive, durable and fungible: more profitable tokens, longer useful life and deeper demand,” the company states, highlighting the interconnectedness of these three factors. This approach positions NVIDIA as a key player in the rapidly evolving landscape of AI infrastructure, where capital expenditure is significant and long-term viability is paramount.

NVIDIA A100 GPU Demonstrates Extended Useful Life

Recent market data indicates that a six-year-old NVIDIA A100 GPU retains approximately 25% of its original value, a figure significantly exceeding typical five-year depreciation schedules according to Silicon Data. This sustained value challenges conventional accounting practices and demonstrates an extended operational lifespan for NVIDIA hardware within AI factories, by the company’s account. The findings suggest a shift in how operators assess the longevity and earning potential of their GPU investments, moving beyond simple book life calculations.

Microsoft’s NVIDIA V100 fleet, for example, operated for 8.4 years, surpassing its six-year accounting life and demonstrating a clear divergence between physical lifespan and financial depreciation. This extended use is not isolated; Barkr reports five-to-six-year lifespans for eight-GPU H100 systems, with newer GB300 NVL72 systems projected to operate for nine to ten years. The ability to extract value from hardware for longer periods directly impacts the overall return on investment for AI factory operators.

The demand for sustained GPU performance extends beyond simple longevity, with Ornn Data finding that the rental price for an A100 GPU on a five-year contract is 80% of the price for a one-month rental. This pricing structure indicates a robust secondary market and a willingness among customers to commit to longer-term contracts, reflecting confidence in the continued utility of the hardware.

This versatility is not limited to AI; the systems also handle non-AI workloads, further broadening their potential applications and extending their revenue-generating capacity, NVIDIA claims. “A factory that runs everything stays useful and revenue-generating even when the work changes,” the company states, emphasizing the importance of fungibility in maximizing long-term returns. NVIDIA founder and CEO Jensen Huang will further discuss these concepts during his GTC Berlin keynote.

NVIDIA Codesign Maximizes AI Factory Earning Capacity

This performance leap stems from an engineering approach focused on maximizing tokens per second per megawatt, directly influencing earning capacity by increasing revenue and lowering costs per token. Beyond raw performance, NVIDIA is engineering AI factories for operational longevity and broad applicability, recognizing that a single-purpose facility risks obsolescence, the company says. A factory capable of handling any accelerated workload, not solely AI tasks, expands potential demand and extends its revenue-generating lifespan.

NVIDIA CUDA-X libraries are central to this strategy, enabling the systems to process diverse tasks beyond traditional AI applications, thereby broadening their utility and attracting a wider range of clients. This approach positions NVIDIA as a provider of adaptable infrastructure, rather than a vendor of specialized hardware. Continuous software optimization further extends the productive life of installed hardware, ensuring sustained performance years after initial deployment.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of The Neuron

The Neuron

With a keen intuition for emerging technologies, The Neuron brings over 5 years of deep expertise to the AI conversation. Coming from roots in software engineering, they've witnessed firsthand the transformation from traditional computing paradigms to today's ML-powered landscape. Their hands-on experience implementing neural networks and deep learning systems for Fortune 500 companies has provided unique insights that few tech writers possess. From developing recommendation engines that drive billions in revenue to optimizing computer vision systems for manufacturing giants, The Neuron doesn't just write about machine learning—they've shaped its real-world applications across industries. Having built real systems that are used across the globe by millions of users, that deep technological bases helps me write about the technologies of the future and current. Whether that is AI or Quantum Computing.

Latest Posts by The Neuron: