
Tech • AI • Robotics
NVIDIA outlined a shift toward “AI factories” powered by extreme co-design across chips, infrastructure, and software to meet surging demand for generative and agentic AI.
Computing is undergoing parallel shifts: from CPU-based systems to GPU-accelerated computing, the rapid rise of generative AI, and the emergence of agentic AI capable of reasoning and acting. These trends are expanding AI use across industries, from scientific research to robotics and autonomous systems.
While transistor counts still grow, the cost and power efficiency gains predicted by Moore’s Law and Dennard scaling ended around 2005. This has forced a redesign of computing around massive parallelism, with GPUs handling workloads once confined to single-threaded CPUs.
AI models are scaling at roughly 10x more parameters per year, while token generation is surging due to inference and reasoning demands. New architectures such as mixture-of-experts dramatically increase internal token usage, with up to 100× more tokens generated per output.
Traditional data centers optimized for cost are being replaced by AI factories designed to maximize tokens per watt and tokens per dollar. These facilities are treated as production systems that generate revenue through AI output rather than simply hosting IT workloads.
Power availability defines AI factory capacity. Techniques similar to energy storage systems are used to smooth fluctuating AI training loads, reducing peak demand. This enables up to 40% more GPUs per gigawatt, directly increasing output and revenue.
NVIDIA described a “five-layer cake” approach: energy, chips, infrastructure, models, and applications. Optimization occurs both within and across layers, ensuring all components—from silicon to software—work as a unified system.
Systems combine Rubin GPUs, Vera CPUs, and networking technologies such as NVLink, Spectrum-X, ConnectX, and BlueField DPUs. This integration enables scaling from single racks to entire data centers while improving efficiency and reducing cost per token by up to 10×.
Dense GPU clusters require liquid cooling and transitions from 50V to 400–800V power systems. Copper interconnects are used within racks, while optical technologies—especially co-packaged optics (CPO)—reduce signal loss and save significant energy at scale.
AI agents create continuous feedback loops where systems generate and process their own prompts. These workloads demand faster networks and new memory architectures, including KV cache (short-term memory) and persistent storage for long-term context.
Specialized systems such as CMX for short-term memory and STX for long-term storage support agent workflows. These designs improve token throughput and efficiency by up to 5×, enabling faster reasoning and tool use.
NVIDIA continues optimizing leading models, achieving up to 30× performance gains on recent architectures. Its CUDA-X libraries accelerate applications across domains including chip design, simulation, and data analytics.
Agentic AI is positioned as a new class of digital worker capable of planning, reasoning, and executing tasks. The same framework can be applied across industries, promising significant productivity gains in sectors ranging from healthcare to manufacturing.
The transition to AI factories reflects a fundamental redesign of computing, where tightly integrated hardware and software systems are engineered to maximize AI output efficiently at massive scale.
Explain this