Cerebras Launches Faster WSE-3 Turbo Processor and Rack-Scale CS-4 System
In 2019, Cerebras Systems introduced the first version of its unusual AI accelerator, the Wafer Scale Engine (WSE). Rather than dividing a 300mm wafer into separate chips, the w...
By Hardware Team
In 2019, Cerebras Systems introduced the first version of its unusual AI accelerator, the Wafer Scale Engine (WSE). Rather than dividing a 300mm wafer into separate chips, the wafer-scale design uses almost the entire wafer as one accelerator, creating the largest possible processor from a single wafer.
Cerebras has since produced three WSE generations, with each increasing performance and supporting broader developer and commercial adoption. The latest previous version, the WSE-3, launched in 2024 and provided 125 PFLOPS of sparse FP16 performance in one processor.
In 2026, Cerebras is introducing the WSE-3 Turbo, a faster version designed to double the original WSE-3's performance. The company is also launching the CS-4, its first true rack-scale system, along with a new networking architecture that allows multiple WSE processors to operate together in one rack.
WSE-3 Turbo: Higher Performance Through Higher Clock Speeds
The WSE-3 Turbo is an updated version of the WSE-3 rather than a completely new wafer-scale design. Cerebras specifies twice the performance of the original processor, with most of the increase coming from higher clock speeds across the processor's major subsystems. Unlike earlier WSE generations, which generally followed TSMC process generations, the WSE-3 Turbo remains on the same 5nm process as the WSE-3.
Cerebras Wafer Scale Engine Generations
| Specification | WSE-3 Turbo | WSE-3 | WSE-2 |
|---|---|---|---|
| AI cores | 900,000 | 900,000 | 850,000 |
| Sparse FP16 performance | 250 PFLOPS | 125 PFLOPS | 75 PFLOPS |
| SRAM | 44GB | 44GB | 40GB |
| Memory bandwidth | 43.2PB/sec | 21PB/sec | 20PB/sec |
| Fabric bandwidth | 53.5PB/sec | 26.8PB/sec | 27.5PB/sec |
| Network bandwidth | 300GB/sec | 150GB/sec | ? |
| Transistor count | 4 trillion | 4 trillion | 2.6 trillion |
| Process node | TSMC 5nm | TSMC 5nm | TSMC 7nm |
| Power consumption | ? | ~27kW | ~23kW |
The WSE-3 Turbo contains 900,000 AI cores and 44GB of on-processor SRAM. Like the WSE-3, it contains 4 trillion transistors and is manufactured using TSMC's 5nm process.
The processor's sparse FP16 throughput rises to 250 PFLOPS, while SRAM bandwidth increases to 43.2PB/sec. Fabric bandwidth reaches 53.5PB/sec, and external network bandwidth rises from 150GB/sec to 300GB/sec. With these elements operating at approximately twice the previous speed, Cerebras has a direct route to doubling WSE-3 performance.
Cerebras has not disclosed the WSE-3 Turbo's power consumption. The company says the CS-4 system enables twice as much power to be delivered to the processor. If the original WSE-3 consumed approximately 27kW, that statement suggests a Turbo power level of about 54kW, although the figure has not been confirmed. Doubling clock speeds can substantially increase power use, so the final consumption will be an important characteristic of the new processor.
CS-4: Three WSE-3 Turbo Processors in One Rack
The second announcement is the CS-4 rack-scale compute system, which succeeds the CS-3 systems that housed WSE-3 processors. The CS-4 redesigns the server, rack, cooling, power, and networking hardware around the wafer-scale processors, creating a system intended to scale beyond a single WSE.
The CS-3 was a single-WSE system and served as the smallest hardware unit in the WSE ecosystem. Each CS-3 occupied 16U and used liquid cooling. Cerebras also offered a CS-3 Rack configuration, but it essentially placed two independent CS-3 systems in one rack. Although CS-3 racks could be deployed in larger clusters, the systems were not connected as a conventional scale-up system.
Cerebras System Generations
| Specification | CS-4 | CS-3 |
|---|---|---|
| Wafers | 3x WSE-3 Turbo | 1x WSE-3 |
| Sparse FP16 performance | 750 PFLOPS | 125 PFLOPS |
| SRAM | 132GB | 44GB |
| Memory bandwidth | 129.6PB/sec | 21.6PB/sec |
| Fabric bandwidth | 160.5PB/sec | 26.7PB/sec |
| I/O bandwidth | 900GB/sec | 150GB/sec |
| I/O latency | 2µs | 5µs |
A complete CS-4 rack contains three WSE-3 Turbo processors. Compared with a single WSE-3, the rack has three times as many wafer-scale processors. Compared with the two-system CS-3 Rack configuration, it represents a 50% increase in WSE processors per rack. Combined with the WSE-3 Turbo's doubled performance, Cerebras specifies six times the performance of a CS-3 system, or approximately three times the performance of a CS-3 Rack.
The Nexus Rack Architecture
The hardware surrounding the WSE processors is a major part of the CS-4 design. Cerebras created a new rack-scale architecture called Nexus to accommodate the higher power and cooling requirements of the WSE-3 Turbo while supporting future processors and networking technologies.
Nexus separates the rack's supporting equipment from the compute modules. Power supplies, fans, and other infrastructure are positioned at the front, while the WSE systems are installed at the rear. The design is intended to span multiple generations, allowing WSE processors and other components to be replaced without replacing the entire rack.
Each WSE is installed in a vertically oriented, self-contained module called a backpack. The backpack connects the processor to the rack's power and liquid-cooling loops, as well as to networking and other I/O systems.
Modular Networking and Connectivity
The CS-4 also increases networking capacity. The WSE-3 Turbo provides 300GB/sec of external fabric and networking bandwidth, twice the 150GB/sec available on the WSE-3. This supports higher bandwidth and lower latency between wafers.
Each backpack is paired with a wafer I/O module containing the networking hardware. Separating the I/O modules from the WSE processors allows Cerebras to update networking components independently of the processors. The modular arrangement is intended to support newer networking technologies in later systems.
The initial CS-4 rack uses RDMA over Converged Ethernet v2 (RoCEv2) for scale-up and scale-out networking. Cerebras has also designed an optional topology that connects backpacks in a chain rather than using Ethernet switches in a hub-and-spoke arrangement.
According to Cerebras, the chain topology reduces total inter-wafer latency to 2 microseconds. It also simplifies the rack by removing the need for major networking components beyond those contained in the backpacks. Direct backpack-to-backpack connections replace switches and the associated cabling needed to connect each backpack to them.
Nexus Roadmap
Cerebras plans to use the Nexus platform as the foundation for several generations of rack-scale systems. In addition to the CS-4, the company has committed to using the architecture for the CS-5 and CS-6 systems planned for later in the decade.