Cerebras's Next Generation CS-4: Fast Just Got Faster

Myron Xie · SemiAnalysis · August 19, 2026 at 01:32 · ⏱ 12 min read  | Read on Substack ↗
Summary
Cerebras's new CS-4 system doubles performance and memory bandwidth over the CS-3 by dramatically increasing power consumption and clock speeds on the same 5nm silicon. While this architecture provides a massive interactivity advantage over GPUs for decode-heavy workloads, its inherent on-wafer memory capacity constraints require expensive scale-out or complex disaggregated partnerships with HBM-based systems to run frontier models with large context windows.
  • The CS-4 uses the same 5nm WSE-3 silicon and 44GB of SRAM capacity as the CS-3, but extracts double the performance by doubling clock speeds and power delivery.
  • A single CS-4 rack holds three wafer-scale engines and draws 125-135kW TDP, nearly double the 23kW power draw of a single CS-3 wafer.
  • Cerebras claims 43 PB/s of total on-chip memory bandwidth, which it markets as roughly 2,000x more than Nvidia's upcoming Rubin architecture.
  • The CS-4 is expected to hit near 4,000 tokens/sec/user on frontier models, compared to an estimated realistic 100 tokens/sec/user for Nvidia Blackwell GPUs.
  • Running a 1.6T parameter model like DeepSeek V4 Pro at a 1M context window requires around 20 to 40 Cerebras WSE systems, costing over $20M in CAPEX and 1MW of power.
  • Cerebras is partnering with AMD and AWS Trainium to build heterogeneous, disaggregated inference setups that pair the CS-4 with HBM-based XPUs to overcome memory limits.
Read time 12 min
Length 12,812 chars
Category finance
Ideas
Myron Xie Substack author, SemiAnalysis
Cerebras is directly attacking Nvidia's inference dominance, claiming 2,000x more memory bandwidth than Rubin and delivering 20-40x more interactivity than Blackwell GPUs for decode workloads.
Cerebras is directly attacking Nvidia's inference dominance, claiming 2,000x more memory bandwidth than Rubin and delivering 20-40x more interactivity than Blackwell GPUs for decode workloads. Risk: Nvidia GPUs maintain a significant advantage in flexibility, allowing dynamic reallocation between prefill and decode workloads, whereas Cerebras disaggregated setups lock in fixed hardware ratios.
Myron Xie Substack author, SemiAnalysis
Cerebras is explicitly partnering with AMD for open, heterogeneous, disaggregated inference setups, pairing the CS-4's decode capabilities with AMD's HBM-based XPUs.
Cerebras is explicitly partnering with AMD for open, heterogeneous, disaggregated inference setups, pairing the CS-4's decode capabilities with AMD's HBM-based XPUs. Risk: Disaggregated setups introduce network latency bottlenecks and require complex software orchestration to function efficiently.
Myron Xie Substack author, SemiAnalysis
The article notes the new CS-4 I/O module seems designed specifically for AWS, allowing them to use EFA NICs to interface CS-4 with AWS Trainium servers for disaggregated inference.
The article notes the new CS-4 I/O module seems designed specifically for AWS, allowing them to use EFA NICs to interface CS-4 with AWS Trainium servers for disaggregated inference. Risk: AWS may prefer to keep customers entirely within its proprietary Trainium/Inferentia ecosystem rather than heavily subsidizing third-party Cerebras hardware.
Myron Xie Substack author, SemiAnalysis
The author explicitly notes that latency through the 2-layer fat-tree network using Arista ethernet switches is reduced to 3 microseconds for the CS-4 system.
The author explicitly notes that latency through the 2-layer fat-tree network using Arista ethernet switches is reduced to 3 microseconds for the CS-4 system. Risk: Competitors are quoting all-in switch latencies in nanoseconds, making a 3-microsecond latency a potential bottleneck for pipeline parallelism.
More from SemiAnalysis

This newsletter, published August 19, 2026, features Myron Xie discussing NVDA, AMD, AMZN, ANET. 4 trade ideas extracted by AI with direction and confidence scoring.

Speakers: Myron Xie  · Tickers: NVDA, AMD, AMZN, ANET