AI Business

Nvidia Reframes AI Infrastructure Around Token Economics and Power Efficiency

As AI systems grow more complex, Nvidia argues that data center economics must shift from measuring individual chip performance to optimizing entire systems for tokens per watt.

·3 min read
Nvidia ties AI factory economics to tokens and power efficiency
Nvidia ties AI factory economics to tokens and power efficiency

The economics of AI infrastructure are moving beyond simple access to powerful graphics processing units. When agentic systems rely on multiple models, databases and tools working in concert, the entire data center must function as a unified computing platform.

This shift is redirecting focus away from individual processors toward the broader infrastructure ecosystem that converts raw computing power into actionable intelligence. Networking, storage, processors and software must coordinate seamlessly at scale while maximizing the intelligence generated from each unit of electrical power consumed, according to Ian Buck, vice president and general manager of hyperscale and HPC at Nvidia Corp.

Buck articulated this new framework during remarks at the Fully Connected event, where he spoke with theCUBE Research's Dave Vellante and John Furrier. "Instead of cars or devices or PCs, it's tokens," he said. "These assets are not IT; they're not cost. They're actually appreciating, revenue-generating, fungible, durable, productive parts of an economy."

Inference reshapes AI factory economics

The commercial value of an AI factory flows from inference, where deployed models respond to user requests and generate tokens. However, inference does not eliminate the need for training, since organizations must continuously refresh their deployed models as new data arrives and market conditions evolve, Buck explained.

"It's not just fire and forget on all these services," he said. "As companies are using these models, they're refining them, they're aligning them, they're adding more data to them. Having them up to date and aware — that actually is a little bit of training. We're seeing the work in reinforcement learning and online alignment."

Latency considerations introduce an additional economic dimension for workloads where faster reasoning commands higher value. Nvidia's Groq 3 LPX inference accelerator pairs with its Vera Rubin platform to boost per-user token throughput for time-critical applications.

"If there's value in those tokens to have the fastest possible thinking, LPX can be boosted on top of Vera Rubin to make that possible," Buck said. "We're seeing a lot of interest in areas like fintech and other areas where things are happening in real time."

Power makes efficiency a system-level priority

Electrical power capacity represents the fundamental constraint on how much computing infrastructure a data center can accommodate. This bottleneck elevates tokens per watt to a critical metric for AI factory economics and compels hardware manufacturers to deliver efficiency gains with successive product generations.

"Data centers have a natural cap, and that cap is actually their power," Buck said. "With every generation of GPU, we make sure that our tokens per watt is upwards of 10 times more efficient. In fact, we saw that with Blackwell — we got, in the end, a 30x improvement in tokens per watt."

This transition also requires rethinking the scope at which systems are architected and managed. CoreWeave enables customers to select specific configurations or leverage higher-level inference services that fine-tune the tradeoff between throughput and token speed, according to Buck.

https://www.youtube.com/embed/64vTNUuQaeQ?feature=oembed

"CoreWeave can do that for customers," he said. "They don't have to feel overwhelmed by all the choices. That's where our partner ecosystem is so important."