Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
Original source
Guests
Rich Sutton is a professor of computing science at the University of Alberta and a pioneering reinforcement learning researcher.
Co-founder of Oak Lab and researcher in reinforcement learning and continual learning.
Summary
Rich Sutton makes the case that continual learning is the natural default for intelligence: agents act, learn, and update all the time, while the field has awkwardly treated learning as something separate from use. He and Khurram Javed connect this to the Bitter Lesson, arguing that systems should scale via computation and learning rather than injected human knowledge, and that LLMs are only a partial win because the Internet is finite and models do not really learn after deployment. Their “big world” view says the environment is vastly larger than any fixed simulator or dataset, so synthetic data and human-built simulations cannot be the full answer. They discuss continual learning as an algorithmic problem—catastrophic forgetting, per-weight step sizes, generate-and-test in feature space, and continual backprop with new random units. The episode closes on Oak Lab’s goal: a small team building a self-maintaining, energy-efficient system that can keep learning, plan with learned models, and eventually scale to something like a trillion-parameter mind running on 20 watts.