OpenAI's Dan Roberts: Why AI Can Now Make Discoveries

Original source
Artwork for OpenAI's Dan Roberts: Why AI Can Now Make Discoveries

Guest

Dan RobertsOpenAI RL team lead

OpenAI researcher leading the Foundations of Reinforcement Learning team.

Summary

Dan Roberts, who leads OpenAI’s Foundations of Reinforcement Learning work, makes the case that RL is now a core mechanism for turning compute into intelligence, especially once models are strong enough to use test-time compute to reason rather than merely autocomplete. He uses OpenAI’s recent mathematics result as an example of explorer-style search: the model assumed a conjecture was false, pursued a long contrarian chain of reasoning, and found a disproof with an unexpected cross-field connection. Roberts contrasts OpenAI’s informal-language approach with DeepMind’s Lean/formal-proof route, and explains RLHF as preference comparisons used to train a proxy reward model because humans cannot provide real-time training feedback. A recurring theme is that scientific discovery requires more than scale: research taste, problem selection, and the ability to sustain long-horizon reasoning are still missing pieces. He argues the trajectory is smooth rather than discontinuous, but expects more math, science, and engineering breakthroughs soon, with AI increasingly participating in the research loop itself.

Notes

Guests

Dan Roberts

Hosts

Matt Turck

Mentioned

Noam BrownYann LeCunRich Sutton