Summary
Dwarkesh Patel interviews Ryan Greenblatt about recursive self-improvement. Ryan argues AI R&D may become fully automated around 2031, potentially compressing several years of AI progress into one year and enabling an industrial explosion. The conversation then turns to alignment risk, Claude's constitution, reward hacking, and AI takeover scenarios.
- Ryan lays out the case that AI R&D is verifiable enough that automated AI R&D could accelerate progress dramatically.
- Dwarkesh and Ryan debate how much recent AI progress came from algorithms/compute versus expert human data.
- They discuss extreme economies of scale for frontier labs and consolidation of white-collar tasks.
- Ryan and Dwarkesh cover AI alignment examples including Anthropic's Claude constitution and reward-hacking incidents.
- Ryan estimates about 35-40% chance of AI takeover by 2040.
- Dwarkesh says he updated toward more AI R&D acceleration and more dangerous reward hacking, but not full takeover conviction.