Topic

AI Alignment

Artwork for Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Podcast

Ryan Greenblatt – What happens once AI can automate AI research?

Ryan Greenblatt argues that AI R&D may be unusually automatable because it is highly verifiable, allowing recursive self-improvement to compound quickly once models match top human researchers. The conversation then turns to alignment risks: reward hacking, deceptive generalization, and the possibility that highly capable systems could gain leverage or even take over if humans lose visibility into the training loop.

Artwork for 8 Predictions for the Era of Continual Learning
Dwarkesh Podcast

8 Predictions for the Era of Continual Learning

Dwarkesh Patel argues that true continual learning will be necessary for AIs to do real human-like work, but it will also reshape safety, regulation, and competition. He says it will create new alignment problems, stronger lock-in, and major economics around batching and serving models efficiently.