DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - Huawei, GB300 NVL72, MI355X, B200

Bryan Shan · SemiAnalysis · June 09, 2026 at 12:15 · ⏱ 39 min read  | Read on Substack ↗
Summary
DeepSeekV4 inference performance improved dramatically across hardware platforms in the first 43 days, with NVIDIA's CUDA stack delivering reliable Day 0 support while AMD's ROCm suffered early setbacks before a 100x recovery. The article argues that software ecosystem maturity, not raw hardware specs, determines deployable inference performance, reinforcing NVIDIA's competitive moat and highlighting Huawei's growing Day 0 readiness as a challenger.
  • DeepSeekV4 1.6T model achieved 100x throughput improvement on AMD MI355X within 26 days, from unusable Day 0 to exceeding H200 at low interactivity.
  • NVIDIA's CUDA stack (vLLM/SGLang) supported DeepSeekV4 out of the box on Day 0, while TensorRT-LLM had a hidden bug (hardcoded hidden size 4096 vs 7168) that corrupted inference for over a week.
  • Huawei Ascend 950DT was the only non-NVIDIA stack with Day 0 inference support, using CANN and a dedicated AI CPU for metadata planning and MC² fused compute-communication operators.
  • GB300 NVL72 rack-scale inference achieved best-in-class cost per million output tokens at $0.156 (50 tok/s/user) using Wide EP and MTP speculative decoding.
  • B200 tokens per all-in provisioned utility megawatt improved from ~300k to ~500k tok/s/MW (1.7x) between Day 0 and June 5th, driven entirely by software optimizations.
  • AMD's ATOM inference engine had zero production customers and initially ran with batch size 1; its FP4 MoE support required fallback kernels on Day 0.
  • DeepSeekV4 architecture replaces MLA with Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), achieving 50x KV cache reduction at 1M context length.
  • MegaMoE kernel from DeepSeek achieves 1.92x speedup over naive MoE by wave-scheduling experts to overlap dispatch/combine communication.
Read time 39 min
Length 39,249 chars
Category finance
Ideas
Bryan Shan Substack author, SemiAnalysis
NVIDIA's CUDA ecosystem delivered Day 0 support for DeepSeekV4 across B200, B300, and GB300, with vLLM/SGLang working out of the box. Despite TensorRT-LLM bugs, the overall software maturity and GB300
NVIDIA's CUDA ecosystem delivered Day 0 support for DeepSeekV4 across B200, B300, and GB300, with vLLM/SGLang working out of the box. Despite TensorRT-LLM bugs, the overall software maturity and GB300's rack-scale performance advantage ($0.156 per million output tokens) reinforce NVIDIA's inference leadership. Risk: Huawei's Day 0 readiness and AMD's rapid 100x improvement suggest competitive pressure is rising; Nvidia's moat is not absolute.
Bryan Shan Substack author, SemiAnalysis
AMD's MI355X and ROCm stack had disastrous Day 0 performance—only FP8 mode worked, distributed inference did not function, and ATOM engine ran at batch size 1 with no production customers. Although a
AMD's MI355X and ROCm stack had disastrous Day 0 performance—only FP8 mode worked, distributed inference did not function, and ATOM engine ran at batch size 1 with no production customers. Although a 100x improvement was achieved by Day 26, the article highlights that AMD SGLang and vLLM progress lags far behind CUDA, and the company is focusing on ATOM instead of the widely-used vLLM. Risk: If AMD continues to prioritize proprietary engines over open-source ecosystem, it may cede inference market share to NVIDIA and Huawei.
Bryan Shan Substack author, SemiAnalysis
CoreWeave is explicitly thanked for contributing compute to the open-source community, scrambling to provide two spare GB300 NVL72 racks for benchmarking. This validates CoreWeave's role as a key enab
CoreWeave is explicitly thanked for contributing compute to the open-source community, scrambling to provide two spare GB300 NVL72 racks for benchmarking. This validates CoreWeave's role as a key enabler of cutting-edge AI infrastructure and positions it as a go-to cloud provider for frontier inference workloads. Risk: Heavy dependency on NVIDIA hardware supply and potential competition from other GPU cloud providers.
More from SemiAnalysis

This newsletter, published June 09, 2026, features Bryan Shan discussing NVDA, AMD, CRWV. 3 trade ideas extracted by AI with direction and confidence scoring.

Speakers: Bryan Shan  · Tickers: NVDA, AMD, CRWV