Nvidia Bridges Data Centers With Giga-Scale Networking, Accelerates AI Model Serving
Nvidia unveiled Spectrum-XGS networking technology and new inference serving methods designed to connect multiple data centers and speed up AI model deployment across distributed infrastructure.

Nvidia Corp. revealed a suite of enhancements to its artificial intelligence software stack and networking infrastructure, targeting faster deployment and operation of AI systems at scale. The chipmaker, whose graphics processors underpin the modern AI industry, introduced Spectrum-XGS—marketed as "giga-scale"—an extension of its Spectrum-X Ethernet switching platform built specifically for AI computing tasks. While Spectrum-X enables data to flow across clusters housed in a single data center to feed AI models, Spectrum-XGS adds the ability to orchestrate and link multiple data centers together.
"So, you've heard us use terms like scale up and scale out. Now we're introducing this new term, 'scale across,'" said Dave Salvator, director of accelerated computing products at Nvidia. "These switches are basically purpose built to enable multi-site scale with different data centers able to communicate with each other and essentially act as one gigantic GPU."
The distinction matters operationally. Expanding "up" involves deploying larger machines, while expanding "out" means adding more machines within the same facility. Yet most data centers face hard constraints: they can only draw so much electrical power or dissipate so much heat before performance degrades. These physical limits restrict how many machines or how much processing capacity can fit into any one location.
Salvator emphasized that the system reduces jitter and latency—the unpredictability in when packets arrive and the time lag between sending and receiving data. These metrics matter greatly for AI networks because they directly affect the bandwidth achievable when GPUs operate across geographically separated sites.
The new offering complements NVLink Fusion, a network fabric technology Nvidia released in May that lets cloud operators scale up their data centers to manage millions of GPUs simultaneously. Together, the two technologies address scaling at different levels: NVLink Fusion handles growth within a single data center, while Spectrum-XGS manages growth across multiple data centers.
Researching better methods to serve AI models
Dynamo represents Nvidia's inference serving framework—the software layer responsible for deploying models and executing inference tasks. The company has been investigating a disaggregated serving approach through Dynamo that separates "prefill," the phase where context is built, from "decode," the phase where tokens are generated, distributing these workloads across different GPUs or machines.
This research direction addresses a critical shift in AI workloads. Inference, once treated as a secondary concern compared to model training, has become a major bottleneck in the era of agentic AI, where reasoning models produce far more tokens than their predecessors. Dynamo tackles this challenge by providing a faster, more efficient, and more economical approach to inference handling.
"If you look at both interactivity on a model like GPT OSS, OpenAI's most recent community model they just released, we're able to achieve, about a 4X increase in tokens per second," said Salvator. "You look at DeepSeek, we're also able to achieve really significant bumps there in terms of a 2.5X increase."
Nvidia is also exploring "speculative decoding," a technique that deploys a second, smaller model to predict the outputs of the primary model for a given input, aiming to accelerate overall performance. "The way that this works is you have what's called a draft model, which is a smaller model which attempts to sort of essentially generate potential next tokens," said Salvator.
Since the smaller model trades accuracy for speed, it can produce multiple candidate tokens for the main model to validate.
"The ability here is that the more that that draft model can speculatively correctly guess what those next tokens need to be, the more performance you can pick up," explained Salvator. "And we've already seen about a 35% performance gain using these techniques."
According to Salvator, the primary AI model performs verification in parallel against its learned probability distribution. Only tokens that pass verification are retained; rejected tokens are thrown away. This architecture maintains latency beneath 200 milliseconds, a threshold Salvator characterized as "snappy and interactive."


