Upscale AI Partners With Nvidia While Competing With It
Why GPU Interconnects May Need an Open Standard CPU manufacturers have traditionally used proprietary methods to connect multiple processors into shared memory clusters. That re...
By Hardware Team
Why GPU Interconnects May Need an Open Standard
CPU manufacturers have traditionally used proprietary methods to connect multiple processors into shared-memory clusters. That remains largely true despite efforts by the CCIX and CXL consortia to create universal interconnects that allow multiple CPUs to operate like one larger processor.
Attempts to establish broadly adopted CPU SMP and NUMA standards largely failed. Customers wanted a common way to connect CPUs, but workloads combining processors from different architectures and vendors appeared relatively uncommon. Each manufacturer could also advance its SMP and NUMA chipsets independently, eventually integrating those functions into the CPUs themselves. This allowed vendors to coordinate interconnect improvements with changes to their cores, memory subsystems, and I/O without waiting for a standards organization to approve new features.
A similar standard for GPUs and XPUs may nevertheless emerge in what is increasingly called scale-up networking. Competition can reduce costs, and the market may benefit from alternatives to proprietary interconnect systems.
Nvidia currently charges a significant premium for NVSwitch and NVLink technology. These products provide high bandwidth and relatively low latency for connecting as many as 72 GPUs into a shared-memory system, which Nvidia calls the NVL72 rack-scale machine.
Through NVLink Fusion, Nvidia allows custom CPU designers to add NVLink ports to their processors so they can share memory with Nvidia GPUs. It also allows hyperscalers, cloud builders, and AI model developers to add NVLink ports to their XPU accelerators and place them within Nvidia rack-scale systems. Nvidia has not, so far, allowed customers to put NVLink ports on both their own CPUs and XPUs while purchasing NVSwitches and an NVLink protocol license to build independent shared-memory systems.
That could change in the future. Nvidia might eventually adopt an industry-standard scale-up network with proprietary protocol extensions, particularly for AI inference. The company has less incentive to give up control of AI training interconnects, where it holds most of the market.
Upscale AI's SkyHammer and SkyFabriX
Upscale AI is one company pursuing an alternative. The company previously raised $200 million in Series A funding and has raised $500 million in total. Since the January funding announcement, its valuation has doubled to $2 billion, while its employee count has more than doubled to over 300 people.
The company has provided limited information about the performance of its SkyHammer scale-up switch ASIC, which has been under development for several years. Upscale AI co-founder and executive chairman Rajiv Khemani and chief executive officer Barun Kar have extensive experience in high-performance systems and networking.
A presentation by Arvind Srikumar, Upscale AI's senior vice president of product and marketing, indicates that SkyHammer will provide aggregate bandwidth of 115.2 Tb/sec. That places it alongside scale-out network ASICs from Cisco Systems, Broadcom, and Nvidia. The SkyFabriX architecture built around SkyHammer is expected to support switch equipment with aggregate bandwidth measured in multiple petabits per second.
This will likely require multiple SkyHammer ASICs in a single switch, a design approach used by Broadcom and Cisco in fixed-port and modular switches. A two-tier design using six independent ASICs can double the bandwidth available from a single switch ASIC architecture. Four ASICs can provide downlinks, while two additional ASICs cross-connect the system. This arrangement increases latency for some paths, although not all of them. Specialized fabric connectors can also be used to create larger modular switches.
Upscale AI is expected to develop devices with higher aggregate bandwidth over time because the economics can support that approach. A switch with twice the native bandwidth per port can be built from six ASICs, but the cost per port is approximately 1.5 times that of a single device offering twice the bandwidth. This estimate reflects street pricing rather than ASIC manufacturing costs. The company has not provided a roadmap for future SkyHammer ASICs.
Scale-Up Capacity and Latency
SkyHammer's architecture can support up to 576 accelerators in one networking tier. Multiple tiers can extend the system to thousands of accelerators. The exact upper limit has not been disclosed, although 1,024 and 1,156 are potential targets.
The UALink consortium is targeting 1,024 XPUs in a single tier. Nvidia plans to expand NVSwitch beyond its current 576-GPU limit to 1,152 GPUs in the Rubin Ultra generation, using a multi-tier rather than single-tier network. Nvidia currently supports 72 GPUs in one domain, while 576 GPUs are possible across multiple tiers but are not intended for production workloads.
At 576 devices, SkyHammer would provide approximately 200 Gb/sec per device. That is substantially below the 1.8 TB/sec offered by the NVSwitch 5 and NVLink 5 generation in Grace-Blackwell NVL72 systems. Vera-Rubin NVL72 systems are expected to use NVSwitch 6 and NVLink 6, providing 3.6 TB/sec.
With 72 XPUs, SkyHammer provides 1.6 Tb/sec per port, equivalent to 200 GB/sec of memory bandwidth across the scale-up network. That is one-sixth of the bandwidth provided by the NVSwitch 5 stack.
Bandwidth may not be the only important measure, especially for AI inference. Predictable latency can matter more than minimum latency. Upscale AI says SkyFabriX is designed to provide predictable latency and a lossless fabric.
According to Arvind Srikumar, latency is a critical metric for token generation. The SkyHammer ASIC includes mechanisms intended to provide predictable jitter, while standards-based protocols provide link-level reliability. Within the fabric, additional mechanisms are intended to make transmission fully lossless.
Srikumar also said that SkyFabriX uses substantial parallelism rather than relying on the fixed pipeline arrangements common in many Ethernet fabrics. Upscale AI designed the system specifically for scale-up networking instead of adapting an existing Ethernet fabric after the fact.
UALoE, ESUN, and Ethernet Switching
The practical differences among these approaches will become clearer as deployments begin using UALink over Ethernet, commonly called UALoE, and ESUN memory protocols on high-speed Ethernet switches adapted for scale-up networking.
SkyHammer uses UALoE rather than native UALink. Its support for ESUN also indicates that SkyHammer is an Ethernet switch ASIC. The design could instead have been a higher-speed, higher-radix ASIC capable of running native UALink and emulating ESUN when necessary, but Upscale AI chose the Ethernet-based approach.
Nvidia Partnership for Scale-Out Networking
Upscale AI has partnered with Nvidia as its preferred supplier for scale-out fabrics, specifically selecting Nvidia's Spectrum-X Ethernet switches. The partnership does not use Nvidia's Quantum-X InfiniBand switches.
Upscale AI will acquire Nvidia Spectrum-X ASICs and build its own scale-out switches around them. It will also adapt SkyOS, the network operating system used on SkyFabriX scale-up switches, to run on Spectrum-X ASICs.
SkyOS is an AI-optimized version of SONiC and its associated Switch Abstraction Interface, or SAI. Microsoft created SONiC in 2016, and the software has since become a de facto standard in many hyperscaler and cloud-builder data centers. Upscale AI's Spectrum-X-based systems will support 400 Gb/sec, 800 Gb/sec, and 1.6 Tb/sec ports.
The company says SkyOS will cover both scale-up and scale-out networks, while its SkyCMD telemetry and orchestration software will manage them as a unified environment. Customers will be able to combine CPUs, GPUs, and XPUs connected through DPUs or NICs when building AI clusters.
There is currently an important limitation: Nvidia does not support SkyHammer or SkyOS as an alternative to NVSwitch for scale-up networking. As a result, SkyHammer cannot currently provide scale-up connectivity for Nvidia GPUs.
AMD is relying on companies such as Upscale AI to offer scale-up networking based on UALink, UALoE, or ESUN. The availability of these alternatives could influence how accelerator vendors approach proprietary and open interconnects.
Nvidia could also reduce the bandwidth per port in a future NVSwitch ASIC while increasing its radix, allowing a flatter NVSwitch network. That would require sacrificing bandwidth. Another possibility would be to port NVLink to SkyHammer, potentially making Nvidia a major supplier of UALoE or ESUN switching technology.