Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis

Alec Ibarra · SemiAnalysis · July 23, 2026 at 00:47 · ⏱ 24 min read  | Read on Substack ↗
Summary
Nvidia's Vera Rubin NVL72 delivers substantial inference performance gains over Blackwell-based GB200/GB300, achieving up to 5.4x more output tokens per megawatt and 5x lower cost per token at high interactivity levels, driven by architectural improvements like HBM4, LUT-based weight compression, and kernel reuse. Early results from CoreWeave show strong gains, but the article cautions that the benchmark uses an older model and a 2025 baseline, and that Rubin's true advantage over modern competitors like AMD MI455X and Google TPUv7 remains unverified until third-party benchmarks land in late 2026.
  • Vera Rubin NVL72 delivers 5.4x performance per MW and 5x performance per dollar over GB200 NVL72 on DeepSeek R1, according to CoreWeave benchmarks.
  • Rubin can reuse Blackwell SM100 kernels, easing software bringup and time-to-market vs the Hopper-to-Blackwell transition.
  • Key microarchitecture upgrades: 328 KiB SMEM, 256 KiB TMEM, inline TMA descriptor updates for MoE, doubled BF16/FP16 exponential throughput, 2x FP8/FP4 Tensor Core throughput, and 2.8x higher memory bandwidth via HBM4.
  • New LUT-based weight compression achieves ~3.125 bits per weight using a 3-bit index and 8-entry codebook, reducing HBM footprint and bandwidth demand.
  • Rubin adds 2:4 activation sparsity support, but no accuracy data or benchmark results using it have been published yet.
  • Compared to July 2026 GB300 NVL72, Rubin's performance lead ranges from 2x at low interactivity (up to 100 tok/s/user) to 4x at 200 tok/s/user, ballooning to 5.4x at 300 tok/s/user where GB300 can barely serve.
  • Rubin's TCO per GPU-hour is $3.57 vs $1.84 (GB200) and $2.36 (GB300), but its higher throughput makes it 1.5-5x cheaper per million output tokens depending on interactivity.
  • Nvidia has committed to submitting verifiable Rubin benchmarks to InferenceX by Q3 CY2026; Google and AMD plan to submit TPUv7 and MI455X results, setting up an objective comparison.
Read time 24 min
Length 24,413 chars
Category finance
Ideas
Alec Ibarra Substack author, SemiAnalysis
The article quantifies strong inference performance gains for Nvidia's Vera Rubin architecture over its own prior generations (GB200/GB300), validating continued competitive advantage in AI inference
The article quantifies strong inference performance gains for Nvidia's Vera Rubin architecture over its own prior generations (GB200/GB300), validating continued competitive advantage in AI inference hardware. The detailed architectural improvements (HBM4, LUT decompression, kernel reuse) suggest Nvidia maintains leadership. Risk: Early benchmarks are on an older model (DeepSeek R1) and engineering samples; real-world performance with modern models and mature software may differ. Competitor submissions (AMD, Google) could narrow the gap.
Alec Ibarra Substack author, SemiAnalysis
The article explicitly states that CoreWeave's Vera Rubin NVL72 benchmarks were conducted on a Dell Engineering Sample rack, indicating Dell is providing early hardware for Nvidia's next-generation ra
The article explicitly states that CoreWeave's Vera Rubin NVL72 benchmarks were conducted on a Dell Engineering Sample rack, indicating Dell is providing early hardware for Nvidia's next-generation rack-scale system, positioning it as a key partner in deploying Rubin-based solutions. Risk: Dell's reliance on Nvidia's product cycles means any delays or issues in Rubin's production ramp could impact Dell's server revenue. The article notes Rubin's simpler cableless design may improve ramp, but risks remain.
More from SemiAnalysis

This newsletter, published July 23, 2026, features Alec Ibarra discussing NVDA, DELL. 2 trade ideas extracted by AI with direction and confidence scoring.

Speakers: Alec Ibarra  · Tickers: NVDA, DELL