
Key Points
- 01Cerebras (CBRS) unveils the CS-4 rack-scale AI inference system built from three WSE-3T processors
- 02Company specs cite 750 PFLOPS of AI compute with very high memory and fabric bandwidth
- 03Claims include more than 4,400 tokens per second per user on GPT-OSS-120B
- 04Cerebras (CBRS) targets higher efficiency than its CS-3 system, with shipments set for this quarter
Cerebras debuts CS-4 rack-scale AI system
Cerebras Systems introduced the CS-4 rack-scale AI accelerator on August 18, 2026, presenting it as its latest system for large-scale inference workloads. The CS-4 is built as a rack solution composed of three Wafer Scale Engine 3 Turbo (WSE-3T) processors, integrating them into a single appliance aimed at high-throughput deployment.
For this three-wafer configuration, Cerebras states that the CS-4 delivers 750 PFLOPS of AI compute. The company positions this as its highest-performing rack system to date, designed to support demanding generative AI and large language model applications that require both speed and scale.
Key performance and bandwidth specifications
Cerebras publishes detailed performance specifications for the CS-4. The system is specified to provide 129.6 petabytes per second of memory bandwidth and 160.5 petabytes per second of on-chip fabric bandwidth, highlighting the data-movement capacity within and across the wafers.
The CS-4 is also described as offering 7.2 terabits per second of system I/O bandwidth, intended to connect the appliance to external networks and storage. Wafer-to-wafer latency is stated to be as low as two microseconds, which is designed to support scaling to larger clusters and the deployment of very large models.
WSE-3T processor architecture
At the heart of the CS-4 is the WSE-3T processor, which Cerebras characterizes as its latest wafer-scale engine. Each WSE-3T is said to contain four trillion transistors and approximately 900,000 AI-optimized cores, along with 44 GB of on-wafer SRAM.
Cerebras cites the WSE-3T as delivering 250 PFLOPS per wafer, forming the basis for the aggregate 750 PFLOPS figure in the three-wafer CS-4 rack. The dense on-wafer memory and large core count are presented as supporting high parallelism and fast access for AI inference workloads.
Inference performance claims on GPT-OSS-120B
In system-level comparisons, Cerebras highlights CS-4 performance on the GPT-OSS-120B model. The company claims the system can deliver more than 4,400 tokens per second per user, a figure aimed at demonstrating throughput for large language model inference.
Cerebras also asserts that, in its head-to-head testing, the CS-4 offers up to 30 times faster inference than GPU-based solutions. These comparisons are used to position the rack system as a high-speed alternative for enterprises and organizations running intensive AI inference workloads.
Efficiency gains and rollout timeline
Relative to its prior-generation CS-3, Cerebras states that the CS-4 delivers up to 10 times more throughput per watt, emphasizing improved energy efficiency. The company also indicates that a single CS-4 can achieve up to twice the speed of a single CS-3 system, underscoring generational performance gains.
Cerebras describes the CS-4 as integrating into a modular Nexus rack architecture, intended to simplify deployment and future upgrades. The company plans to begin first shipments of the CS-4 this quarter, signaling the transition of the platform from announcement to commercial availability.
Key Takeaways
- 01The CS-4 consolidates three WSE-3T wafers into a rack-scale system, targeting large, latency-sensitive AI inference workloads.
- 02Cerebras emphasizes memory, fabric, and I/O bandwidth, along with very low wafer-to-wafer latency, as central to CS-4 scalability.
- 03Claimed gains in tokens-per-second and throughput-per-watt over both GPUs and the CS-3 frame the CS-4 as a generational performance and efficiency step.
References
- https://www.manilatimes.net/2026/08/19/tmt-newswire/globenewswire/cerebras-unveils-cs-4-up-to-30-times-faster-than-gpu-based-solutions/2408047
- https://www.theregister.com/systems/2026/08/19/cerebras-cs-4-rack-systems-juice-chips-for-every-last-drop-of-ai-performance/5289286
- https://www.cerebras.ai/cs4
- https://finance.yahoo.com/technology/article/cerebras-says-latest-offering-is-fastest-ai-accelerator-in-the-industry-as-it-takes-aim-at-nvidia-000000090.html