Summary
Steve Hou, head of research at Silicon Data, discusses the evolving AI compute market, focusing on token efficiency, model routing, GPU rental pricing, and memory bottlenecks. He explains how falling token costs and rising efficiency could shift value from frontier AI labs to the compute layer, and presents data showing firming GPU rental rates and tight supply, supporting a bullish outlook for semiconductors and AI infrastructure.
- Silicon Data builds indices to bring futures and hedging to the physical AI compute market.
- The token expenditure index measures usage-weighted AI model pricing, showing substitution from expensive frontier models to cheaper open-weight models.
- GPU rental indices for H100, A100, and B200 reveal strong inference demand, with even older A100 rates holding firm.
- Forward curves for GPU rentals are steepening, indicating cloud providers are confident enough to avoid long-term discounts, signaling underlying tightness.
- Memory (DRAM) demand continues to grow due to longer context lengths and multimodal data, with efficiency improvements unlikely to derail secular demand growth.
- The value chain may shift from frontier model labs toward the compute layer or end-user enterprises as models become more substitutable.
- Enterprise AI adoption is expected to accelerate as cheaper models enable true token maxing and experimentation, eventually delivering measurable ROI.
- US-China decoupling and both nations doubling down on AI capex could further boost demand for compute and hardware.