Digital Event Horizon
NVIDIA AI Factories: Maximizing Return on Investment through Productivity, Durability, and Fungibility
NVIDIA AI Factories provide a scalable and efficient platform for generating strong returns on investment.AI factories focus on building a factory that can produce more, last longer, and serve more types of work.The three key factors that shape AI factory returns are earning capacity, useful life, and demand.NVIDIA AI factories are engineered to maximize earning capacity, durability, and fungibility.The platform is productive, durable, and fungible, which maximizes AI factory returns.Power is the binding constraint on an AI factory, governing earning capacity.The installed base keeps earning, and older generations remain valuable due to the right fit depending on workload complexity and shape.The NVIDIA platform is fungible and runs every type of AI workload, in every phase and place.The platform's versatility shows up in production across customers, with diverse use cases and applications.
NVIDIA AI Factories are revolutionizing the way artificial intelligence (AI) is developed and deployed, providing a scalable and efficient platform for generating strong returns on investment. The concept of AI factories, which was first introduced by NVIDIA, focuses on building a factory that can produce more, last longer, and serve more types of work. In this article, we will delve into the key aspects of AI factories and explore how NVIDIA's platform maximizes return on investment through productivity, durability, and fungibility.
According to NVIDIA, AI factories are built by the megawatt, even by the gigawatt, and each megawatt factory costs roughly $60 million. AI factory operators will only commit capital on that scale with a clear view of the return on investment. The three key factors that shape AI factory returns are earning capacity, useful life, and demand.
Earning capacity refers to the amount of money the factory can earn in a year if it sells every token it can produce. Useful life refers to how long the AI hardware keeps earning, while demand refers to how much demand there is for those tokens. It is crucial to note that these three factors are interdependent and cannot be fully offset by one another. A factory that can run more kinds of workloads finds more demand, keeping it earning year after year.
NVIDIA AI factories are engineered to maximize all three of these factors. They are productive, delivering the highest throughput per megawatt and the lowest cost per token, which maximizes their earning capacity. They are durable, with NVIDIA GPUs and systems keeping earning years after they ship, extending useful life. They are fungible, running every type of AI — in every phase and every place — as well as many workloads that don’t involve AI at all, which deepens and broadens the demand they can serve.
NVIDIA platform is productive, durable and fungible, which maximizes AI factory returns. Engineering codesign across the full stack maximizes AI factory throughput, and continuous software optimization keeps installed hardware productive years after it ships. NVIDIA CUDA-X libraries let a factory run any accelerated workload. A standardized architecture then puts all of it within reach of any operator, deployable from a validated reference design.
Power is the binding constraint on an AI factory. This makes tokens per second per megawatt the number that governs earning capacity. More tokens inside a fixed power envelope means more revenue. Lower cost per token means more margin on it. SemiAnalysis AgentX data shows NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than NVIDIA GB300 NVL72, and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model. Gains of that size come from extreme codesign across the full stack.
Two questions follow. If every generation makes tokens dramatically cheaper, does demand for compute shrink? No, it expands. Cheaper tokens make more use cases economical, and those use cases consume more tokens than the efficiency saved. The second question is about durability. If each generation is so much better than the last, what happens to the older generations?
The installed base keeps earning, not every workload needs the newest system. The right fit depends on a workload’s complexity and shape, which is why the last generation keeps earning after the next one arrives. The NVIDIA A100 GPU shipped in 2020 and is still in commercial service six years later, demonstrating its continued economic value. CoreWeave recently extended bookings for units first introduced in 2020 through 2029. Over the years, every major operator has extended the depreciation schedule on its servers, which is a guess about when hardware stops earning, and one that keeps moving out. A September 2026 Sprout analysis, “The Productive Life of a Data Center GPU,” tracks how that schedule has shifted across every major operator.
Every major operator has extended server life. Source: Sprout, “The Productive Life of a Data Center GPU,” September 2026. Data based on company disclosures and press reporting compiled by Sprout. Accounting life is a conservative proxy for physical life — Microsoft’s NVIDIA V100 fleet ran 8.4 years against a six-year book life. Barkr puts useful life at five to six years for an eight-GPU H100 system and nine to 10 years for GB300 NVL72, based on what those systems resell for. Silicon Data shows a six-year-old A100 GPU is still worth a quarter of what it cost, where a five-year depreciation schedule had it at zero more than a year ago.
CUDA, the software platform with which NVIDIA GPUs are programmed, runs across generations, so nothing an operator already owns is stranded when a new architecture arrives. Continuous software and kernel optimization keeps improving what existing hardware can do. The same platform runs machine learning, deep learning, generative AI, reasoning, agentic AI, and physical AI. Each new kind of work arrived on hardware that was already installed. That’s fungibility. The more kinds of work a system can take, the longer it keeps finding work.
AI factories built for one kind of work are a bet that the work stays. A factory that runs everything stays useful and revenue-generating even when the work changes. NVIDIA AI factories run every type of AI model — open and proprietary — across language, vision, biology, physics, and robotics. They run every phase, from data processing through pretraining, post-training, and inference. And they run in every place, from hyperscale and AI clouds to sovereign programs, enterprise data centers, and the edge.
The NVIDIA platform is fungible and runs every type of AI workload, in every phase and place. Not all of it is building and running AI models. The same infrastructure runs data processing, scientific computing, simulation, graphics, and more. All of those workloads, AI and non-AI alike, reduce to the same parallel math, and NVIDIA GPUs are built to run exactly that across thousands of cores at once. CUDA is why one chip can simulate light, fold a protein, and predict the next token. More than 1,000 ready-made CUDA-X libraries sit on top, covering everything from deep neural networks, computational lithography, and quantum circuit simulation to vector search and climate modeling, with more than 10 million developers building on them.
That’s what makes NVIDIA GPUs general-purpose accelerated computing rather than a custom ASIC built for one workload. Being general purpose does not mean being generic: Tensor Cores and the Transformer Engine put AI-optimized hardware inside a programmable architecture, delivering specialization and flexibility in one chip. One architecture running all of it is what keeps utilization high, and the return with it.
This versatility shows up in production across customers. Lilly builds and runs protein, small-molecule, and genomics models on a 1,016-GPU, on-premises cluster, plus chatbots and agentic workflows for its own teams. Pinterest post-training and deploying a vision language model on a hyperscale cloud, across 14,000 GPUs spanning NVIDIA Blackwell, Hopper, and earlier architectures. Revolut processes data for billions of transaction records with NVIDIA cuDF, then training and deploying a foundation model on an AI cloud. Runway trains a world model on NVIDIA Hopper and serves on the NVIDIA Blackwell platform using cloud infrastructure.
Texas A&M University runs molecular simulation and AI drug discovery on its supercomputer, at 95-98% utilization across 26 projects and seven institutions. The same is true beyond AI. Dassault Systèmes powers the virtual twin simulation behind aircraft certification at Wichita State and vehicle design at Lucid Motors. And Unilever builds product imagery from digital twins rather than photo shoots, cutting production costs in half.
No list of examples here would be complete — that’s the point. NVIDIA AI factories are engineered to be productive, durable, and fungible: more profitable tokens, longer useful life, and deeper demand. That’s what maximizes their return.
Related Information:
https://www.digitaleventhorizon.com/articles/NVIDIA-AI-Factories-Maximizing-Return-on-Investment-through-Productivity-Durability-and-Fungibility-deh.shtml
https://blogs.nvidia.com/blog/productive-durable-fungible-ai-factories/
Published: Thu Oct 1 10:37:40 2026 by llama3.2 3B Q4_K_M