OpenAI Jalapeño: Better Than Nvidia Blackwell

Bryan Shan · SemiAnalysis · 25 августа 2026, 14:00 · ⏱ 31 мин чтения  | Читать в Substack ↗
Резюме
OpenAI has successfully developed a generalized AI inference ASIC named 'Jalapeño' in partnership with Broadcom, achieving industry-leading performance-per-watt that beats Nvidia's Blackwell and rivals the upcoming Vera Rubin. The rapid 16-month development cycle and AI-assisted software bring-up demonstrate that frontier AI labs can successfully co-design hardware to bypass Nvidia's CUDA moat and drastically reduce their reliance on merchant silicon.
  • OpenAI developed the Jalapeño ASIC in ~16 months in partnership with Broadcom, utilizing TSMC's N3P node and CoWoS packaging.
  • Jalapeño beats Nvidia's Blackwell on performance-per-watt and competes head-to-head with Vera Rubin on performance-per-TCO, even without using speculative decoding.
  • The chip features HBM4 memory, likely from Samsung, delivering 15.4TB/s of memory bandwidth per package to outpace current HBM3E accelerators.
  • At concurrency 1, Jalapeño achieves over 700 tokens/sec/user on DeepSeek R1 and ~1,400 tok/sec/user on Kimi-K2.5 and GPT-OSS using single-token prediction.
  • OpenAI's AI-assisted coding tool, Codex, rapidly wrote functional kernels for the chip, challenging the perceived dominance of Nvidia's CUDA software moat.
  • The rack architecture, designed with Celestica, houses 128 ASICs drawing 130kW, with a scale-up network supporting up to 2,048 XPUs via Broadcom Tomahawk 6 switches.
  • OpenAI deliberately avoided prefill-decode disaggregation to maintain high global hardware utilization and eliminate KV cache movement overhead across the network.
Время чтения 31 мин
Объём 31,375 симв.
Категория finance
Идеи
Bryan Shan Автор Substack, SemiAnalysis
The article argues that OpenAI's custom chip beats Blackwell on perf/W and that 'The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon.'
The article argues that OpenAI's custom chip beats Blackwell on perf/W and that 'The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon.' Risk: Nvidia's software ecosystem remains deeply entrenched outside of frontier labs, and their merchant silicon is still the default for the broader market.
Bryan Shan Автор Substack, SemiAnalysis
Broadcom is explicitly named as OpenAI's ASIC design partner for Jalapeño and the provider of the 102.4T Tomahawk 6 switches used in the Chana switch trays for the scale-up network.
Broadcom is explicitly named as OpenAI's ASIC design partner for Jalapeño and the provider of the 102.4T Tomahawk 6 switches used in the Chana switch trays for the scale-up network. Risk: Custom ASIC margins are typically lower than merchant silicon margins, and volume depends entirely on OpenAI's deployment success.
Bryan Shan Автор Substack, SemiAnalysis
TSMC is manufacturing the Jalapeño compute die on its N3P node, the I/O chiplet on N3E, and handling the CoWoS packaging for the high-volume ramp.
TSMC is manufacturing the Jalapeño compute die on its N3P node, the I/O chiplet on N3E, and handling the CoWoS packaging for the high-volume ramp. Risk: Capacity constraints on advanced nodes and CoWoS packaging could limit the volume of OpenAI's rollout.
Bryan Shan Автор Substack, SemiAnalysis
The article explicitly notes that 'The system level design is done in partnership with Celestica,' giving them direct exposure to OpenAI's rack-scale hardware rollout.
The article explicitly notes that 'The system level design is done in partnership with Celestica,' giving them direct exposure to OpenAI's rack-scale hardware rollout. Risk: Hardware manufacturing and system integration is a low-margin business subject to intense competition.
Bryan Shan Автор Substack, SemiAnalysis
AMD is supplying the host CPUs for the system, with the article noting each host rack houses '16 host CPU trays... each houses two Turin-class AMD EPYC CPUs.'
AMD is supplying the host CPUs for the system, with the article noting each host rack houses '16 host CPU trays... each houses two Turin-class AMD EPYC CPUs.' Risk: Host CPU revenue is a relatively small portion of total AI cluster spend compared to the accelerators and networking.
Ещё от SemiAnalysis

This newsletter, published August 25, 2026, features Bryan Shan discussing NVDA, AVGO, TSM, CLS, AMD. 5 trade ideas extracted by AI with direction and confidence scoring.

Speakers: Bryan Shan  · Tickers: NVDA, AVGO, TSM, CLS, AMD