Full article — scored 9/10
NVIDIA Vera Rubin NVL72 Turns AI-Agent Efficiency Into the New Data-Center Battleground
NVIDIA says its Vera Rubin NVL72 rack-scale platform can deliver up to 30 times higher throughput per megawatt than GB300 NVL72 on agentic AI workloads, a claim that reframes the AI infrastructure race around power, latency, and cost per token rather than raw accelerator counts alone.
The 30x Claim: A Power Story, Not Just a Speed Story
NVIDIA’s latest Vera Rubin NVL72 update is not simply a faster-chip announcement. The company says newly measured data shows Vera Rubin NVL72 systems delivering up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic AI workloads, using SemiAnalysis AgentX workload traces made from real-world agentic coding sessions . In the same announcement, NVIDIA says that gain translates into up to 30x more agentic work for the same energy footprint in power-constrained AI factories .
That distinction matters. AI infrastructure buyers are no longer optimizing only for benchmark tokens per second on a single prompt. They are trying to run agents that keep state, call tools, spawn subagents, reuse context, and keep users waiting as little as possible. NVIDIA’s technical blog frames the same result as up to 30x higher AI-factory throughput per megawatt than GB300 NVL72 on AgentX workloads, while noting that the Vera Rubin results were measured by NVIDIA and are still pending SemiAnalysis review .
The caveat is important. The headline is not a universal “30x faster than everything” claim. It is a power-efficiency and serving-throughput result for a specific class of long-context, agentic inference. In NVIDIA’s own description, the performance curve depends on interactivity targets, model behavior, cache reuse, and system-level software optimizations . The strongest reading is that Vera Rubin NVL72 is being positioned as an AI-factory platform for agents, not merely as another GPU generation.
Why AI Agents Change the Infrastructure Math
The reason NVIDIA is emphasizing throughput per megawatt is that agentic workloads behave differently from chatbot workloads. NVIDIA says agentic AI workloads can consume 15 times more tokens than a simple chat request because agents repeatedly search, reason, invoke tools, delegate work to subagents, and fold prior context into later steps . Its technical blog describes modern agents as multi-step workflows that invoke tools, coordinate subagents, and carry growing context from one turn to the next .
This is the central economic problem. A customer-support chatbot may answer a short question with a bounded prompt and response. A coding agent might inspect a repository, generate code, run tests, read error messages, ask a subagent to evaluate a dependency, revise the patch, and explain the result. Each turn can add context. Each saved key-value cache can reduce recomputation. Each tool-call gap can leave expensive accelerators underused unless the serving system knows how to schedule around it.
That is why NVIDIA is using AgentX as the reference workload. The NVIDIA technical blog says AgentX evaluates agentic-coding inference by replaying production-style sessions and capturing long-context prefill, KV-cache reuse, tool-call gaps, dynamic concurrency, and distributed mixture-of-experts execution . NVIDIA’s blog adds that its measured inference-throughput data used recorded real-world agentic coding sessions with actual context growth, tool calls, and subagent spawning preserved .
For operators, this pushes the conversation away from peak FLOPS and toward useful work per watt. If a rack consumes a fixed power budget, the winning system is the one that keeps agents interactive while completing more sessions per megawatt. NVIDIA says Vera Rubin NVL72 can also lower cost per million tokens by up to 35x compared with GB300 NVL72 in the tested agentic context . That is the business case behind the engineering claim: more tasks completed inside the same electrical envelope.
Vera Rubin NVL72 Is a Rack-Scale System
NVIDIA is presenting Vera Rubin NVL72 as a full platform rather than a single processor. The current Vera Rubin NVL72 platform includes the Vera CPU, Rubin GPUs, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC as part of a seven-chip architecture for AI factories deploying agents at scale . That matters because agentic inference bottlenecks move across compute, memory, networking, scheduling, and storage.
The software stack is part of the pitch. NVIDIA says the efficiency gains come from system-level optimizations including MoE serving runtimes such as SGLang, TensorRT-LLM, and vLLM; DeepGEMM-based kernels; mixed-precision formats such as MXFP4 and MXFP8; NVIDIA Dynamo session-aware serving; and the high-bandwidth NVLink scale-up fabric connecting 72 GPUs for coordinated rack-scale inference . NVIDIA’s blog also points to disaggregated serving, rate matching, distributed KV-caching, KV-cache offloading, KV-aware routing, and fused CUDA kernels as key techniques for agentic scale .
The unifying theme is that agents create uneven work. Prefill may be heavy when a long context arrives. Decode may become the user-visible bottleneck when the system must emit tokens quickly. Tool execution may temporarily shift pressure onto CPUs. NVIDIA’s architecture tries to separate these phases, route repeated context intelligently, and keep the GPUs fed rather than waiting on memory, networking, or orchestration.
NVLink is central to that story. NVIDIA says the NVL72 scale-up domain enables high-bandwidth, low-latency inter-GPU communication for expert parallelism and distributed KV-caching, and that sixth-generation NVLink Switches deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives . In practical terms, NVIDIA wants customers to think of the rack as one coordinated inference machine, not 72 accelerators loosely tied together.
LPX Targets the Decode Bottleneck
NVIDIA also used the same news cycle to push Groq 3 LPX, an interactive AI inference accelerator designed to extend Vera Rubin inference for agents. NVIDIA says Groq 3 LPX is now in full production and extends the inference performance of Vera Rubin NVL72 systems by dramatically increasing token-generation rates . The company reports that Groq 3 LPX delivered 3,400 output tokens per second in Artificial Analysis benchmarking on Gemma 4 31B with a 100,000-token context, which it describes as the fastest recorded performance for that model .
The LPX announcement clarifies where NVIDIA sees the next bottleneck. Rubin GPUs are positioned for large-scale context processing, while LPX is positioned for latency-sensitive decode workloads where the agent must generate tokens quickly enough to remain interactive . NVIDIA says LPX can provide 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform .
Nebius is named as the first AI cloud to adopt NVIDIA Groq 3 LPX through its Nebius Token Factory inference platform . Groq also announced that it will be among the first adopters of NVIDIA Groq 3 LPX and Vera Rubin NVL72, working with Dell Technologies to deploy the platform for its AI inference cloud . Independent coverage of the launch also framed Groq 3 LPX as a purpose-built extension to Vera Rubin aimed at supercharging agentic AI inference .
This is strategically important because decode latency compounds across agents. A single slow response may be tolerable in chat. In an agent loop, every delay affects the next step: inspect, generate, test, verify, revise. If LPX can accelerate token generation while Rubin handles context, NVIDIA can sell the combination as a token factory rather than a conventional GPU rack.
Customer Signals and Ecosystem Momentum
NVIDIA is also connecting Vera Rubin to customer and partner adoption. In its broader Vera Rubin inference update, the company says Nebius is first to adopt Groq 3 LPX, CoreWeave has deployed Spectrum-X Multiplane in production to connect Vera Rubin racks, and SpaceXAI has announced that NVIDIA Vera CPUs will power its next generation of agentic AI .
The SpaceXAI announcement is particularly designed to show that Vera Rubin is more than a GPU platform. NVIDIA says SpaceXAI will deploy Vera CPUs for CPU-intensive orchestration, tool execution, code execution, data processing, and simulation tasks behind agentic AI . NVIDIA also says SpaceXAI plans to expand its AI infrastructure behind Grok on the Vera Rubin platform and base a planned first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack-scale system .
NVIDIA says the Vera CPU has 88 NVIDIA-designed Olympus cores, Spatial Multithreading technology, and LPDDR5X memory bandwidth of up to 1.2 TB/s, with up to 1.8x faster task completion compared with x86 CPUs across workloads including agentic AI, reinforcement learning, and data processing . The point is not just CPU performance in isolation. It is the idea that an agent platform needs fast CPUs to coordinate the non-GPU work around inference.
What Still Needs Proof
The Vera Rubin NVL72 announcement is significant, but it should be read carefully. First, the 30x figure is tied to AgentX and the DeepSeek V4 Pro workload at specific interactivity targets, not to every model or every deployment . Second, NVIDIA says the Vera Rubin measurements are pending SemiAnalysis review, which means buyers will want to see third-party-validated curves and production telemetry . Third, NVIDIA notes that the early Vera Rubin results do not yet reflect Vera CPU performance for tool calling, so the final platform story may change as more components are measured together .
There is also the usual difference between a benchmark result and a business result. Cost per token depends on hardware pricing, utilization, power contracts, cooling, software maturity, model choice, and customer traffic shape. NVIDIA can credibly argue that throughput per megawatt is the right metric for AI factories, but cloud providers still have to convert that efficiency into reliable capacity and prices that customers will pay.
Still, the direction is clear. NVIDIA is trying to make power efficiency, rack-level co-design, and token economics the central measures of AI infrastructure. If the 30x agentic-efficiency claim holds up under review and at customer scale, Vera Rubin NVL72 could mark a shift from the era of bigger model training clusters to the era of production agent factories. The race will not be only about who owns the fastest chip. It will be about who can turn a fixed megawatt into the most useful, responsive, and profitable intelligence.
Sources from the last 72 hours
- [1]Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI AgentsAug 24, 2026, 12:00 AM UTC
- [2]NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per WattAug 24, 2026, 12:00 AM UTC
- [3]With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for AgentsAug 24, 2026, 12:00 AM UTC
- [4]Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to MarketAug 24, 2026, 3:26 PM UTC
- [5]Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agentsAug 24, 2026, 3:00 PM UTC
- [6]NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AIAug 24, 2026, 3:00 PM UTC
- [7]SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive ScaleAug 24, 2026, 3:00 PM UTC
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.
