Summary
Dwarkesh Patel narrates his essay on the next AI training paradigm, arguing that RLVR scaling may build competent agents but that real-world continual learning will require new techniques such as on-policy self-distillation and 'dreaming' to update model weights from deployment experience.
- Labs are betting that training on millions of verifiable RL tasks can produce general problem-solving agents.
- Verifiability alone is insufficient; domains also need grindable, replayable simulators.
- Computer use has lagged because websites and apps are not easily cloned into parallel training environments.
- RLVR generalization to long-horizon, real-world skills may be limited.
- Deployment inference compute is not currently used to improve model weights.
- On-policy self-distillation could compress session learning back into weights.
- Dreaming or test-time training could become a fourth compute-scaling axis.
- By 2027-2028, AIs may improve mainly through online learning from broad deployment.