Skip to main content
Back to Blog
AI/MLEnterpriseInnovation
21 August 20265 min readUpdated 26 August 2026

Cerebras Presents Rack-Scale WSE Systems at Hot Chips 2026

Cerebras presented its rack scale hardware at Hot Chips 2026 following the announcement of its WSE 3 Turbo accelerators and CS 4 racks. The CS 4 is Cerebras's first dedicated ra...

By Hardware Team

Cerebras presented its rack-scale hardware at Hot Chips 2026 following the announcement of its WSE-3 Turbo accelerators and CS-4 racks. The CS-4 is Cerebras's first dedicated rack-scale system designed to operate within a single scale-up domain. At its center are three WSE-3T wafer-scale engines, whose clock speeds and performance are approximately twice those of their predecessors.

CS-4 and the Nexus Platform

The CS-4 rack is significant not only because it combines three WSEs in one system, but also because it introduces the Nexus platform. Nexus is intended to serve as the hardware foundation for Cerebras systems over the next several years. It also represents a step toward combining WSEs with GPU racks, pairing the different accelerator types according to their respective strengths. Cerebras is working toward systems that pair WSEs with AMD's forthcoming Instinct MI455X accelerators.

Cerebras describes CS-4 as an improvement in both throughput and energy efficiency. Compared with CS-3, the company states that CS-4 provides twice as many tokens and up to 10 times more tokens per watt. Cerebras also presented comparisons with GPUs, including a claim that CS-4 can be up to 30 times faster than a GPU for the workloads discussed.

WSE-3T Memory and Rack Architecture

Cerebras identifies memory bandwidth as a key factor in inference performance. The WSE-3T relies entirely on on-chip SRAM and provides 43,000 TB per second of memory bandwidth.

The Nexus rack places power infrastructure at the front and compute modules at the rear. The compute hardware is installed in removable backpacks, which provide several functions:

  • I/O connectivity for the rack
  • Power delivery
  • Cooling connections
  • Leak detection
  • Energy monitoring
  • Valved quick-disconnect connections

The CS-4 backpack supports twice the power and cooling capacity of the CS-3 design while using 50% fewer components.

Power Delivery and Cooling

CS-4 receives power from the front of the system. Each power supply unit is protected by a 30A circuit breaker and accepts inputs of up to 277V AC. Each backpack can contain up to 30 PSU modules. Power is delivered from the top of the rack, and the AC power system is fully phase-balanced.

The rack's vertical design places DC-DC converters close to the busbar. The WSEs are positioned close to these components as well, reducing resistive losses in the power path.

The backpacks also connect the WSEs to the rack's liquid-cooling network. Each unit includes leak detection and an energy meter, along with valves that support quick disconnection. A next-generation wafer I/O interface uses a new module structure.

Wafer Fabric and Interconnect

Cerebras emphasized that the system does not require a large number of external cables. Processing elements within each wafer communicate across an internal fabric providing 53 PB per second of bandwidth. The backpacks also provide direct links between the wafers in a CS-4 system.

These wafer-to-wafer connections can provide latency as low as 2 microseconds, with 2.4 Tb per second of aggregate bandwidth available from each wafer.

CS-4 Specifications and Model Scaling

The CS-4 rack uses three WSE-3T engines. On paper, the combined system provides a sixfold improvement across many throughput and bandwidth metrics compared with the preceding configuration.

Cerebras also showed how a model such as GPT-5.6 SOL can flow through the system. The company states that CS-3 can already run the largest frontier models, while CS-4 is intended to support even larger models in the future.

In its summary of CS-4, Cerebras described the system as delivering twice the token performance and 10 times the performance on a tokens-per-second-per-watt basis compared with CS-3.

CS-5 and CS-6 Roadmap

The Nexus platform is also planned as the foundation for the CS-5 and CS-6 systems, which will use newer WSE designs.

Cerebras expects CS-5 to arrive in 2027. The company has stated that the system could deliver up to 3 million tokens per megawatt, or up to 10,000 tokens per second per user.

CS-6 is planned to introduce DRAM for the first time in a Cerebras system. The design will stack DRAM above the wafer-scale engines. Based on the information presented, the approach may allow Cerebras to use smaller processors with memory capacities similar to those of the largest SRAM-equipped WSEs.

Cerebras provided limited detail about CS-6, but the general concept involves reducing the amount of SRAM on each WSE to make room for more compute hardware, while stacked DRAM supplies additional memory capacity.

Supporting wafer-scale processing requires changes across the system, including 3D packaging, power delivery, cooling, and interconnect design. The CS-4 is the first rack-scale implementation of this approach, while CS-5 and CS-6 extend the Nexus platform with newer WSE generations.