Skip to main content
Back to Blog
AI/MLNetworkingCloud Computing
7 October 202611 min readUpdated 8 October 2026

Growing Pains: How Distributed AI Training Is Changing Inter-Datacenter Networks

Large scale AI training is increasingly moving beyond individual datacenters. Google has said that Gemini was trained synchronously across clusters in multiple locations. Micros...

By Hardware Team

Large-scale AI training is increasingly moving beyond individual datacenters. Google has said that Gemini was trained synchronously across clusters in multiple locations. Microsoft has connected AI datacenters in Wisconsin and Georgia as part of what it describes as one distributed AI supercomputer. AWS has connected AI compute clusters across wide areas to support Anthropic’s development of Claude models. Meta has built high-capacity datacenter interconnects for model training, while CoreWeave and Google Cloud have announced cross-cloud training through a private interconnect. Azure may follow later in the year.

This geographic expansion is partly a response to the limits of individual datacenters and campus clusters. New AI models can require substantial amounts of compute and power. Cisco estimates that current training workloads can use clusters containing tens of thousands of GPUs. Researchers at Epoch AI estimate that the largest individual frontier training runs could require 4-16GW of power by 2030. Distributing compute allows hyperscalers and datacenter operators to use sites with more available power or space and fewer planning restrictions.

How to Train an LLM

At the center of a large language model is a neural network containing billions of numerical parameters, or weights. During training, a batch of data passes through the model, which produces a prediction. The system measures the prediction error and calculates how the weights should change. It then updates the weights and repeats the process.

Modern models distribute this work across thousands of GPUs and other accelerators, including AWS Trainium and Google TPUs. Different accelerators can process separate data batches or different sections of a model, but they cannot operate entirely independently. At several points, they must exchange results and synchronize before training can continue.

The network generally carries large arrays of numerical data, such as gradients and intermediate results, rather than the original text or images used for training. Thousands of accelerators may need to exchange and combine this information at nearly the same time.

“If you look at any training job, it is basically a repetition of compute and synchronization,” explains Ramesh Sivakolundu of Cisco’s Silicon One architecture team. “You compute the weights, communicate those weights to synchronize every GPU in the cluster, and then restart the computation. That means bursty traffic happens periodically throughout the training process.”

When thousands of accelerators complete a computation phase and begin communicating simultaneously, the network can experience an “incast” event. Many senders converge on the same destination, causing traffic to arrive faster than the receiving link can process it.

The underlying challenge is synchronization. In a synchronous training job, one group of GPUs cannot simply ignore a slower group and continue independently. The network-wide exchange must finish before participating GPUs can begin the next stage. As a result, network delay becomes compute delay.

“You don’t want the network to be the bottleneck in the training process,” says Sivakolundu. “The idea is to keep the GPUs fully occupied and not have the network introduce additional latency into the overall training program.”

Congestion or packet loss does not necessarily stop an entire training job. However, a delayed or dropped flow may require retransmission, extend a collective communication operation, and leave other accelerators waiting. More serious failures can force training to return to a saved checkpoint.

A Different Kind of Interconnect

“Inside a single datacenter, you can assume you have built a full network connecting the GPU racks with the bandwidth you planned for,” says Itamar Gold, director of product management at Cisco.

With scale-across deployments, coordinated AI workloads may have to cross a more constrained inter-site network than the network inside either facility. Traffic heading to another datacenter may be funneled through what is effectively a narrower pipe, allowing a synchronized burst to become a temporary choke point.

Traditional datacenter interconnects connect otherwise independent facilities. They are typically designed for redundancy, reach, and workload distribution. Their traffic often consists of replication data or application data, and flows are usually asynchronous, meaning that one flow does not need to finish in lockstep with others.

In scale-across AI training, the inter-site network becomes part of a coordinated computation rather than simply transferring data between independent systems. Sufficient bandwidth and predictable delivery are therefore essential.

From Rack to Region

Within a rack, bandwidth requirements can reach 800 Gb/sec and 1.6 Tb/sec, although the distances are short. High-speed serial lanes combine to make a GPU cluster operate as a single compute unit.

When an AI fabric extends across a facility, scale-out racks can still require connections of up to 1.6 Tb/sec. At these data rates, copper is generally limited to short distances. Connections across rack rows and datacenter floors instead use pluggable optics, with scale-out links extending from hundreds of meters to approximately 2km.

Scale-across networking extends a coordinated AI fabric across multiple facilities, making inter-site network behavior part of workload performance. Depending on the architecture, links may span metropolitan or regional distances, and potentially farther. Cisco modeling of large distributed AI deployments indicates that aggregate bandwidth requirements can reach approximately 14 times a conventional DCI baseline.

Short-range datacenter optics are not designed to carry 400-800 Gbps signals across hundreds of kilometers. For those distances, operators need coherent optics. These systems use sophisticated modulation and high-performance digital signal processing to keep high-rate optical signals usable across metro and regional fiber spans, compensating for physical impairments that become more pronounced with distance.

Port counts also become significant. Cisco estimates that connecting two 100-megawatt AI sites could require 12,000 to 32,000 coherent optical ports at the scale-across layer, compared with roughly 1,000 to 2,000 ports for conventional DCI between comparable facilities. At 400 Gb/sec or 800 Gb/sec per port, the resulting capacity reaches multiple petabits per second.

Moving this volume of data also creates power and space challenges. Coherent pluggables place the coherent transponder function directly in router or switch ports, avoiding a separate bank of standalone DWDM transponders. Cisco says this can reduce the additional rack space, power, and cooling required by the inter-site optical layer, which is important when the AI facilities themselves are power-constrained.

The Latency Problem

Even with sufficient bandwidth and suitable optics, distance adds latency. When congestion develops, the network needs time to signal the sender to slow down. In a local AI fabric, feedback can arrive quickly. If a connection spans 100 kilometers of fiber, however, considerably more data can be transmitted before the sender learns that a problem exists.

Cisco calculates that an 800 Gb/sec connection spanning 100 kilometers can have approximately 100 MB of data in transit during this feedback interval. Cisco argues that longer feedback loops increase the value of networking silicon with substantially deeper buffers. Shallow-buffer switching architectures designed for very low latency inside a datacenter have less capacity to absorb traffic while congestion feedback travels across a long-distance link.

“The reason you need deeper buffers as the distances get longer is that you don’t get feedback about the transfer of packets for a longer period of time,” says Sivakolundu. “If you don’t buffer deeply enough, you can end up dropping and retransmitting packets, increasing latency.”

Deep buffering can absorb a temporary burst or disruption while traffic management and congestion signaling address the underlying problem. Sustained congestion still requires additional capacity or a change in traffic distribution.

Buffers can also help with oversubscription. If the combined capacity feeding an inter-site connection exceeds the capacity of the link itself, a synchronized burst can arrive faster than the link can drain it. Buffering gives excess traffic somewhere to wait without being dropped while congestion controls respond.

Rather than assigning fixed buffer amounts to individual ports, a shared buffer provides a common pool of packet memory that can be allocated where congestion occurs. This allows a busy link to absorb a larger burst without dropping packets.

“The longer you’re reaching out, the more you may need to buffer,” says Gold. “You also need the flexibility to put as much buffer as possible where it is needed, perhaps on one port or a few ports where you’ve identified a problematic flow. Our approach is a single, fully shared buffer.”

AI training offers one useful characteristic: much of its collective communication is repetitive and therefore more predictable than conventional DCI traffic. Cisco’s approach combines proactive congestion management, which can steer or schedule traffic before a predictable burst creates a problem, with deep buffering as protection against transient congestion and unpredictable events such as link failures. The two methods serve different purposes: one attempts to avoid queues, while the other provides space for traffic when queues form.

Cisco’s Co-Design Approach

Cisco’s scale-across architecture uses the Silicon One P200, a 51.2 Tb/sec programmable, deep-buffer routing processor included in the Cisco 8223 and Cisco N9000 platforms. These systems can be paired with Cisco’s 400G and 800G coherent pluggable optics for long-distance links. Open line systems such as Cisco Open Transport 3000 and Cisco NCS 1014 can handle the optical transport layer where required.

Coherent pluggables generate DWDM wavelengths directly from the routing system, while the line system amplifies and carries those wavelengths across the fiber plant. For links requiring multiple parallel fiber pairs, Cisco’s Open Transport 3000 uses a multi-rail design that combines optical components for multiple rails on one line card. Cisco claims that this reduces power per rail by 75 percent and rack space by 80 percent. Where a separate transport system is required, the NCS 1014 can provide 12.8 Tb/sec of capacity from a 1RU line card.

The P200 includes a programmable run-to-completion network processor and P4 tooling. Cisco says this allows protocol support, telemetry, and other packet-processing features to be developed in software rather than requiring every change to wait for a new silicon generation.

For scale-across deployments, where operators are still testing different topologies and traffic-management methods, Sivakolundu says this programmable flexibility can allow network functions to evolve without every new requirement forcing a chip replacement.

Cisco also includes hardware-based protection for inter-site traffic. The company’s position is that encryption should not create another processing bottleneck for the training fabric.

A central part of Cisco’s approach is co-design. A networking ASIC is defined years before a finished router reaches a datacenter. Cisco argues that having its silicon, system, and optics teams within one company allows requirements for port density, buffering, power, telemetry, and optical interfaces to influence chip design. This differs from building a system around a commercial ASIC and determining afterward what the system can support.

Scale-across deployments have not settled on one standard architecture. Cisco divides the problem into campus, metro, and regional deployments, while operators are testing different combinations of networking and training techniques.

Implementations vary by operator, including distance, topology, and workload design. The common challenge is maintaining coordinated AI performance across sites. The technology remains at an early stage.

“We are seeing early deployments now, but it is a process and I think it will take a few years,” says Gold.

Network design cannot eliminate the latency caused by distance, but it can prevent the interconnect from becoming the primary bottleneck. That allows hyperscalers and datacenter operators to add compute where power and space are available while continuing to treat multiple sites as part of one training infrastructure.