Google deployed very large scale-up domains for a long time while others were limited to smaller domains. Reiner says this hardware/scale-up lead explains why Gemini seemed to have successful pre-training for longer than some other labs, giving Google's AI infrastructure an early advantage for very large or sparse models.
Reiner reinforces the memory wall thesis: hyperscaler capex on memory is enormous, memory is a huge constraint for AI buildouts, and HBM bandwidth is the critical bottleneck for frontier inference, long context, and latency. He adds HBM is not getting hugely better, implying the bottleneck and pricing power persist.
Nvidia's rack-scale NVLink/NVL72 architecture gives every GPU all-to-all connectivity within one rack, matching Mixture-of-Experts expert parallelism. Reiner argues one rack bounds the size of an expert layer, so larger scale-up domains are a huge unlock for bigger sparse models, lower weight-loading latency, and longer context; he credits Nvidia with a genuine ~4x scale-up increase via difficult rack design.