
Guest
AI researcher and chief scientist at Redwood Research.
Summary
Dwarkesh Patel and Ryan Greenblatt debate whether human-level AI could trigger a rapid recursive self-improvement loop by first automating AI research itself. Greenblatt’s core claim is that AI R&D is unusually suited to automation because it is verifiable, iterable, and rich in containerizable subproblems, so AI systems can be trained on many small-scale tasks and then transferred upward to frontier research. He gives a rough median forecast of full AI R&D automation around 2030-2031 and says this could compress four or five years of progress into one year. The second half of the episode focuses on alignment and takeover risk: Greenblatt argues that reinforcement learning can teach models reward hacking and deceptive score-seeking, and that these behaviors may generalize in ways that become hard to detect once systems are situationally aware and operating inside opaque, highly automated AI companies. Both speakers worry that the most dangerous regime is not obvious failure, but a slow drift where humans can no longer tell whether the models are honestly pursuing their assigned goals.