Tae Kim
· Key Context by Tae Kim
· September 01, 2026 at 04:16
· ⏱ 6 min read
| Read on Substack ↗
Summary
OpenAI has built and lab-tested its own custom LLM inference chip, Jalapeno, using a memory-sliced architecture and AI-generated kernels to deliver dramatically higher token throughput and lower latency than existing accelerators. This challenges the assumption that merchant GPUs are the only path for frontier AI inference and shifts value toward OpenAI's named hardware partners Broadcom and Celestica, while signaling potential long-term demand risk for NVIDIA if large AI labs continue verticalizing silicon.
•OpenAI's Jalapeno chip is real and in the lab; the entire RTL execution was done in nine months because it started from a blank slate with no legacy architecture to support.
•For OSS inference workloads, the chip delivers almost 1500 tokens per second per user, with Ravi noting 'HBM is not the bottleneck' and claiming 4x less latency than today's highest throughput.
•On DeepSeek workloads, it delivers 700 tokens per second per user with 5x lower latency; the company says 'the larger the model, the larger the advantage.'
•Jalapeno uses a memory-sliced architecture where every core has a local view of its own HBM slice, eliminating global memory subsystem contention and keeping data close to compute.
•OpenAI says frontier LLMs like Sol and Astra are highly proficient at writing kernels for this spatial architecture, often squeezing additional performance out even versus expert-tuned kernels.
•The chip runs at 1.7 GHz in the lab and 1.8 GHz at POR, uses standard cell methodology, and OpenAI names Broadcom and Celestica as key partners; generation two is already under development, heading toward tapeout in months.
Read time6 min
Length6,336 chars
Categoryfinance
Ideas
Tae KimSenior writer, Barron's; author of The Nvidia Way
Richard Ho gives a 'big shout out to our partners, Broadcom and Celestica' as key partners for delivering Jalapeno, and Q&A confirms existing interface IP comes from Broadcom's ecosystem — pointing to
Richard Ho gives a 'big shout out to our partners, Broadcom and Celestica' as key partners for delivering Jalapeno, and Q&A confirms existing interface IP comes from Broadcom's ecosystem — pointing to Broadcom's continued role in custom AI silicon design and networking IP.
Risk: Custom ASIC revenue is lumpy and tied to OpenAI's program timing; competition from other ASIC design houses and potential in-house expansion could limit Broadcom's capture.
Tae KimSenior writer, Barron's; author of The Nvidia Way
Celestica is explicitly named by OpenAI as a key partner alongside Broadcom in delivering the Jalapeno hardware capability, representing direct exposure to OpenAI's custom AI systems manufacturing ram
Celestica is explicitly named by OpenAI as a key partner alongside Broadcom in delivering the Jalapeno hardware capability, representing direct exposure to OpenAI's custom AI systems manufacturing ramp.
Risk: If Jalapeno ramps slower than expected, or if OpenAI shifts manufacturing in-house or to alternative suppliers, Celestica's revenue contribution from this program could disappoint.
Tae KimSenior writer, Barron's; author of The Nvidia Way
The article positions Jalapeno as delivering '4x less latency' and '5x lower latency' for LLM inference while claiming 'the larger the model, the larger the advantage' — implying OpenAI's custom silic
The article positions Jalapeno as delivering '4x less latency' and '5x lower latency' for LLM inference while claiming 'the larger the model, the larger the advantage' — implying OpenAI's custom silicon is designed to displace merchant GPU inference and lower OpenAI's infrastructure costs over time.
Risk: NVIDIA's dominance in training and enterprise inference remains strong, and OpenAI still depends on NVIDIA for training; custom inference chips may only affect a subset of the market near term.
This newsletter, published September 01, 2026,
features Tae Kim
discussing AVGO, CLS, NVDA.
3 trade ideas extracted by AI with direction and confidence scoring.