Reiner Pope gives a blackboard lecture on how frontier LLMs are trained and served, using roofline math to derive latency/cost trade-offs from batch size, context length, sparsity, and memory bandwidth. He explains mixture-of-experts parallelism, rack-scale NVLink/scale-up vs scale-out networking, and why memory bandwidth is the binding constraint on long-context AI. He then uses public API prices to infer memory bottlenecks and storage tiers, and estimates models are heavily over-trained relative to Chinchilla. The discussion points to advantages for Nvidia rack-scale systems, Google's earlier scale-up lead, and persistent memory/HBM tightness.
This Dwarkesh Patel video, published April 29, 2026, features Reiner Pope discussing NVDA, GOOG, HBM. 3 trade ideas extracted by AI with direction and confidence scoring.
Speakers: Reiner Pope · Tickers: NVDA, GOOG, HBM