AI Czar David Sacks Explains the DeepSeek Freak Out

Watch on YouTube ↗  |  February 02, 2025 at 19:30  |  12:28  |  All-In Podcast
Speakers
David Sacks — General Partner, Craft Ventures
Chamath Palihapitiya — CEO, Social Capital
David Friedberg — CEO, The Production Board

Summary

David Sacks gives his synthesis of the DeepSeek reaction, explaining why an open source reasoning model from a Chinese lab became a global story tied to a one day trillion dollar market cap decline. He argues the six million dollar training figure should be debunked because it covers only the final training run and ignores cumulative R&D and a compute cluster of roughly 50,000 Nvidia Hopper GPUs that cost well over a billion dollars. Chamath Palihapitiya focuses on the engineering, noting DeepSeek replaced PPO with the lighter GRPO algorithm and compiled around CUDA using PTX, which he treats as evidence that Nvidia software lock-in can be bypassed under constraint. David Friedberg closes by arguing that cheaper, faster, commoditizing models push value away from the model layer toward applications, users and the broader economy.

  • Sacks says the story landed hard because it combined US versus China competition with the open source versus closed source debate.
  • R1 compressed industry estimates of China AI lag from six to twelve months down to three to six months.
  • The cited six million dollar figure covers only the final training run, not fully loaded R&D or hardware.
  • American labs final training runs were already in the tens of millions of dollars nine to ten months earlier.
  • Dylan Patel estimates roughly 50,000 Nvidia Hoppers, 10,000 H100s, 10,000 H800s and 30,000 H20s, implying over a billion dollars of compute.
  • Those chips were accumulated before and around export controls through the founder affiliated hedge fund.
  • Chamath points to GRPO replacing PPO and to PTX bypassing CUDA as constraint driven innovations that also expose Nvidia lock-in risk.
  • Friedberg expects commoditizing models to shift value downstream to applications, users and the wider economy.
Ideas
David Sacks General Partner, Craft Ventures 4:32
Frontier training still needs billion-dollar GPU clusters
Sacks argues the widely repeated $6 million figure for DeepSeek R1 should be debunked, and that the conclusion people drew from it about AI compute spending is wrong. The number cannot be validated empirically, and even taken at face value it covers only the final training run, which is not an apples-to-apples comparison with the fully loaded, soup-to-nuts numbers quoted for American labs. On a like-for-like basis the Anthropic founder and Brad Gerstner put OpenAI and Anthropic final training runs in the tens of millions of dollars nine or ten months ago, and a fully loaded DeepSeek number would have to include all prior R&D, experiments and training runs plus the compute cluster. Dylan Patel estimates DeepSeek plus the founder affiliated hedge fund holds roughly 50,000 Nvidia Hoppers, about 10,000 H100s, 10,000 H800s and 30,000 H20s, accumulated ahead of and around export controls, which adds up to well over a billion dollars, and there may be further chips they would not admit to. So the scrappy company trained a frontier model for six million dollars story that drove a one day trillion dollar market cap decline is overhyped, and frontier training still depends on very large Nvidia GPU clusters.
Chamath Palihapitiya CEO, Social Capital 10:18
DeepSeek bypassed CUDA, threatening Nvidia lock-in moat
Chamath reads DeepSeek as a case of necessity being the mother of invention, and the part that matters for Nvidia is that the team compiled entirely around CUDA, writing PTX straight to the bare metal, which is effectively assembly. He says he has argued several times that CUDA is Nvidia biggest moat but also its biggest threat factor precisely because it is a lock-in mechanism. A compute constrained Chinese team demonstrating that the CUDA layer can be bypassed, alongside replacing the orthodox PPO algorithm with GRPO to cut compute memory use, shows the moat is circumventable once developers are forced to economize, so the durability of Nvidia software lock-in is now a live question worth monitoring.
Up Next

This All-In Podcast video, published February 02, 2025, features David Sacks, Chamath Palihapitiya discussing NVDA. 2 trade ideas extracted by AI with direction and confidence scoring.

Speakers: David Sacks, Chamath Palihapitiya  · Tickers: NVDA