Cerebras Overclocks WSE-3 for the Nexus CS-4 Inference System
Cerebras Systems has introduced the CS 4, a new system and rack design built around an overclocked version of its existing WSE 3 waferscale engine. The systems are codenamed “Ne...
By AI Engineering Team
Cerebras Systems has introduced the CS-4, a new system and rack design built around an overclocked version of its existing WSE-3 waferscale engine. The systems are codenamed “Nexus” and are intended to provide the foundation for several future Cerebras generations.
The CS-4 does not introduce a new WSE-4 compute engine. Instead, its compute wafer retains the WSE-3’s 900,000 cores and 44 GB of on-wafer SRAM. It is also manufactured using Taiwan Semiconductor Manufacturing Co. 5-nanometer processes. The manufacturing process is more mature than when Cerebras first delivered the WSE-3 in March 2024.
The updated compute engine is called the WSE-3 Turbo. Cerebras says it can perform twice as much work as its predecessor. The clock speed appears to have increased from 1.4 GHz on the standard WSE-3 to approximately 2.8 GHz. Achieving this increase requires roughly twice the power and more than twice the cooling capacity through the wafer packaging.
Cerebras System Generations
The CS-4 continues Cerebras’ focus on waferscale computing, while changing the way its components are packaged and connected. Earlier systems used a 16U chassis containing one WSE and its host CPU. The Nexus design separates the compute wafer and host CPU from the power supplies and network interfaces, allowing these sections to be upgraded independently.
Cerebras co-founder and chief executive officer Andrew Feldman and co-founder and chief technology officer Sean Lie presented roadmaps showing at least three generations based on the Nexus rack design. The company expects system throughput to increase by 2X each year through 2029. That target could refer to peak performance or effective performance.
Future improvements may come from increasing SRAM capacity relative to compute. Waferscale engines, like GPUs, can be limited by available memory capacity. Adding more SRAM could allow the cores to perform more useful work. A future design may eventually require 3D SRAM stacking to better align memory capacity with compute resources.
The Nexus Rackscale System
The Nexus rack is the principal architectural change in the CS-4. Each rack contains three power shelves at the front and three vertically oriented CS-4 compute “backpacks” at the rear. The backpacks plug into the power shelves and contain the compute, cooling, and networking components.
Cerebras says the Nexus design uses 60 percent more manufacturing automation than previous systems. It provides three times the compute per rack compared with the CS-1 through CS-3 systems, which used one 16U chassis per rack. The new rack is also designed for large clusters, with 50 percent fewer components and deployment times up to three times faster than earlier Cerebras systems, according to Lie.
The modular structure allows Cerebras to ship the Nexus rack and power shelves before the compute backpacks are ready for installation. This arrangement provides flexibility in coordinating component supply and customer deliveries.
Each CS-4 backpack includes a power module, a WSE-3 Turbo module, and a cold plate secured by an array of screws. The WSE-3 Turbo sits above the power module, while the cooling assembly helps dissipate the additional heat generated by the higher clock speed.
Networking and Cluster Expansion
The backpack includes network interface cards at both the top and bottom. The CS-4 specifications refer to “new higher speed wafer links.” Based on those specifications, the wafer I/O module may use six Ethernet ports operating at 200 Gb/sec. That would provide twice the bandwidth per port of the 100 Gb/sec I/O modules used in the CS-1 through CS-3 systems.
The configuration of the earlier systems was not fully clear, including whether they used one Ethernet I/O fabric interface or two. Cerebras had not confirmed all of these networking details at the time of the system announcement.
The updated networking hardware supports a programmable, low-latency packet pipeline and direct wafer-to-wafer links. The number of ports available for wafer interconnection remains unclear. A 2X increase in network radix could allow CS-4 systems to scale beyond the previous limit of 2,048 nodes in a single data-parallel cluster, although Cerebras may continue to cap supported clusters at 2,048 nodes.
Arista Networks is providing the Ethernet switching technology for both scale-out networking between CS-4 systems and front-end connections to users and storage.
Cerebras has offered select customers early access to CS-4 systems. General availability is expected later in the third quarter of 2026.