Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

Watch on YouTube ↗  |  June 30, 2026 at 12:00  |  1:10:15  |  Sequoia Capital
Speakers
Dylan Patel — Founder, CEO, and Chief Analyst at SemiAnalysis

Summary

Dylan Patel of SemiAnalysis explains why AI progress increasingly comes from hardware-software co-design across silicon, systems, and models, not just faster chips. He describes InferenceX, a living benchmark showing roughly 60x annual cost declines, and argues inference will become a huge market. He discusses Nvidia vs. TPUs, custom ASICs, neoclouds, memory/optics bottlenecks, and the persistent compute crunch. He also outlines where SpaceX, space data centers, and analog compute fit on a longer horizon.

  • Dylan Patel argues co-design across silicon, kernels, and models can turn multiple 2x gains into 100x.
  • InferenceX tracks inference performance across hardware/models, showing ~60x annual cost decline and ~40x intelligence-per-watt improvement.
  • Nvidia and Google TPUs each have advantages tied to model architectures; custom ASICs and general-purpose GPUs will coexist.
  • The compute crunch persists because model TAM and useful AI tasks grow faster than compute capacity.
  • Neoclouds exist because AI GPU rental economics differ from traditional cloud; CoreWeave and Crusoe are highlighted.
  • Memory bandwidth and co-packaged optics are key bottlenecks/technology transitions.
  • Cerebras faces large-model/long-context risk despite fast-inference demand.
  • SpaceX and space data centers are seen as long-term opportunities, not near-term trades.
Ideas
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 14:11
Inference will dwarf oil.
Use of tokens and inference will become one of the largest markets in the world, much bigger than oil and many percentage points of GDP, because AI adoption and token value creation are expanding rapidly.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 27:42
Google TPU vertical integration scales.
Google's TPU program is a massive vertically integrated custom AI compute effort: Google will make 10+ million TPUs and $100B+ per year, with TPUs optimized for Gemini and Anthropic's denser models, energy-efficient networking, and multiple design programs, though TPUs struggle with models optimized for Nvidia/DeepSeek.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 28:05
Co-design drives 100x semiconductor gains.
The biggest AI gains come from software-hardware co-design across silicon, systems/kernels, and model layers; separate 2x improvements can compound into 100x when optimized together, favoring companies that co-optimize across the abstraction stack.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 28:38
Nvidia wins general-purpose AI compute.
Nvidia is advantaged as the general-purpose AI compute platform because leading models are co-designed around GPU shapes, expert dimensions, and interconnects, labs do not know their next-year architecture and need a general-purpose bucket, its moat is the downstream ecosystem more than CUDA itself, and Jensen is backing neoclouds/labs to preserve a multipolar world.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 28:46
TSMC co-optimizes across supply chain.
TSMC is a co-design winner because it optimizes not just fabrication but also components, consumables, tools, and customer chip designs, extending co-optimization across many layers of the abstraction stack.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 29:46
HBM bottleneck drives memory innovation.
Memory capacity and bandwidth are a major bottleneck: NAND/DRAM cell breakthroughs have been slow and HBM progress has been mostly more stacks and speed, but upcoming direct-on-chip memory stacking and other innovations could explode bandwidth and create opportunities for advanced memory suppliers.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 38:47
Cerebras faces large-model SRAM risk.
Cerebras is innovative and fast inference is a big market, but its SRAM-based architecture faces a major risk: the best models are the ones users want fast mode on, and if models scale to trillions of parameters with million-token context, they may not fit or be cost-justified on Cerebras.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 44:56
Co-packaged optics coming by 2030.
Co-packaged optics is a secular adoption theme: everyone knows it will happen by the end of the decade, with debate only over whether timing is 2027, 2028, 2029, or 2030, and timing shifts can create large market movements.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 49:51
Custom AI ASIC proliferation continues.
Every hyperscaler and large lab will have its own ASIC program and deploy billions to hundreds of billions on custom AI silicon because they want to co-optimize models and hardware and avoid dependence, even though they will keep general-purpose compute buckets because model architectures keep changing.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 51:10
Compute crunch persists on model progress.
The compute crunch persists because demand for useful AI tasks and long agents is expanding faster than compute capacity; 20 GW is coming online this year and more than 30 GW next year, so compute prices and suppliers should stay supported as long as model progress continues, though delays and capital constraints are risks.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 59:08
Amazon Trainium improves on Anthropic use.
Trainium is really good hardware and getting better, Anthropic is using it heavily, and although it currently rents at sub-$10B per GW versus $12-13B for GPUs, growing adoption could lift its value and support Amazon's custom AI silicon strategy.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 63:38
CoreWeave outperforms hyperscaler AI compute.
CoreWeave's GPU compute is objectively better than Amazon/Google/Microsoft in performance and reliability, and its team is phenomenal; the neocloud opportunity exists because AI GPU rental economics differ from traditional cloud, with customers renting whole racks long-term and Nvidia supporting a multipolar ecosystem.
Dylan Patel Founder, CEO, and Chief Analyst at SemiAnalysis 64:29
Neoclouds capture AI compute demand.
The neocloud opportunity exists because AI GPU rental demand is for whole racks under long-term contracts, making hyperscaler multi-tenant security and CPU-era advantages less relevant, while nimble neoclouds with aligned incentives can deliver faster and better-performing compute; Nvidia supports them to create a multipolar world, though many neoclouds will fail.
Up Next

This Sequoia Capital video, published June 30, 2026, features Dylan Patel discussing AI-SECTOR, GOOG, SMH, NVDA, TSM, HBM, CBRS, CPO, Custom AI ASICs, AI compute, AMZN, CoreWeave, Neoclouds. 13 trade ideas extracted by AI with direction and confidence scoring.

Speakers: Dylan Patel  · Tickers: AI-SECTOR, GOOG, SMH, NVDA, TSM, HBM, CBRS, CPO, Custom AI ASICs, AI compute, AMZN, CoreWeave, Neoclouds