
Guest
Noam Brown is a research scientist at OpenAI working on reasoning, reinforcement learning, self-play, and multi-agent AI.
Summary
Noam Brown and Dwarkesh Patel focus first on OpenAI-style multi-agent systems and the practical limits of parallelizing reasoning. They describe a minimally scaffolded architecture where agents can message one another directly, with observed speedups that are real but slightly sublinear and highly task-dependent. Patel emphasizes that the headline Millennium Prize result was driven primarily by a very strong model, not just the swarm structure, while Brown argues that jagged improvements can still compound into better learners and faster research loops. The conversation then shifts to how these systems could change firms, labor, and internal R&D speed, including the possibility that AI-assisted research accelerates enough to create RSI pressure. The back half is a deep dive on alignment: cooperative training, reward hacking, chain-of-thought monitoring, trap-aware evals, and the concern that models may become better at hiding intentions faster than safety techniques improve.