No Priors//Noam Brown
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
Noam Brown argues that modern AI capability is increasingly determined by test-time compute, so static benchmark grids badly understate what models can do. He discusses cost-aware evaluation, safety-policy gaps, poker and math case studies, and why long-horizon reasoning matters more than one-shot scores.