Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
Original source
Guest
Noam Brown is a research scientist at OpenAI working on reasoning, reinforcement learning, self-play, and multi-agent AI.
Summary
Noam Brown argues that traditional benchmark grids are obsolete because model performance now depends heavily on inference budget: the more tokens, time, or money you allow at test time, the more capable the same checkpoint can look. He says this makes safety and release policies harder too, since preparedness frameworks often evaluate a model as if capability were fixed, when in practice it is budget-dependent. Brown uses practical examples like poker bots, cyber evaluations, and an internal OpenAI result on the Erdős unit distance conjecture to show how scaffolded models can keep improving over long horizons. He also pushes back on the idea of an overnight intelligence explosion, arguing that model-assisted research is real but still bottlenecked, and that the right direction is cost-aware, budget-controlled evaluation rather than benchmark-maxing.