When One Data Center Is No Longer Enough
Training the largest AI models is no longer only a matter of adding GPUs. As clusters expand, access to power is becoming a physical constraint. Infrastructure teams therefore n...
By Hardware Team
Training the largest AI models is no longer only a matter of adding GPUs. As clusters expand, access to power is becoming a physical constraint. Infrastructure teams therefore need to consider how a single training workload can run across multiple data centers.
This creates a different networking challenge. Traditional data center interconnects were designed to move traffic between sites. AI training requires huge, synchronized flows, minimal packet loss, and tightly coordinated communication between GPUs. When one part of the cluster stalls, the delay can affect the entire job.
Networking as Part of the Compute System
In an interview, The Register’s Tim Phillips speaks with Rakesh Chopra, SVP, Silicon and Systems Architecture, Cisco Fellow, about what Cisco calls the “Scale-Across” imperative and its view that distributed AI infrastructure requires a new approach to routing.
Chopra describes how AI workloads are changing the role of the network. Instead of serving only as a transport layer, the network becomes an integral part of the compute system. A key challenge is making geographically separated facilities operate more like a single, deterministic machine. That requires careful attention to buffering, high-speed coherent optics, and tightly integrated silicon when workloads travel across long-distance fiber links.
Managing Synchronized GPU Traffic
The discussion also examines communication within the cluster. Chopra explains how Cisco’s Silicon One architecture and Intelligent Collective Networking are designed to manage synchronized bursts of GPU traffic, reduce bottlenecks, and improve job completion time.
Power efficiency is another central issue. Network infrastructure consumes energy that must be balanced against the power available for GPUs, making efficiency an important consideration as distributed AI environments grow.
Security and Long-Term Adaptability
The interview also covers hardware-accelerated MACsec and IPsec. It examines how programmability can help infrastructure adapt as AI workloads continue to change.
For infrastructure and data center leaders preparing for larger, more distributed AI environments, the discussion presents the networking challenges that arise when a single site is no longer sufficient.