Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

OpenAI Research Scientist Noam Brown explains that traditional AI benchmarks fail to capture the dynamic performance improvements of modern models like GPT-5.5, which can leverage increased test-time compute to enhance capabilities over extended periods, necessitating new evaluation frameworks that account for compute budgets. He also highlights the implications for safety and responsible deployment, advocating for compute-aware assessments to better understand risks, while noting that AI progress remains incremental and benefits from coordinated, nuanced research approaches.

In the discussion with OpenAI Research Scientist Noam Brown, the central theme revolves around the inadequacy of traditional AI benchmarks in evaluating modern models, especially given the rise of large-scale test-time compute. Brown explains that unlike earlier models like GPT-3, which couldn’t effectively leverage increased compute during inference, current models such as GPT-5.5 can significantly improve their performance if allowed more “thinking time” or computational budget at test time. This shift challenges existing evaluation frameworks and responsible scaling policies, which typically assess model capability as a fixed attribute without accounting for how performance scales with inference compute or time.

Brown highlights the practical difficulties this creates, noting that modern models can continue to improve on certain tasks for weeks or even months before their performance plateaus, making it infeasible to run exhaustive evaluations within typical model release cycles. He advocates for benchmarking approaches that incorporate an explicit x-axis representing test-time compute, tokens, or cost, enabling clearer comparisons between models. This approach would also help address issues like benchmark maxing, where models artificially inflate scores by repeated attempts or ensemble methods without considering the total compute budget used.

The conversation also delves into the implications for safety evaluations and responsible AI deployment. Current preparedness frameworks do not adequately consider how a model’s capabilities can scale with inference budget, raising concerns about underestimating risks associated with powerful models when given sufficient compute resources. Brown stresses the importance of explicitly defining the compute budget for evaluations to better understand potential capabilities and risks, especially as models become capable of running complex, long-horizon tasks that could have both beneficial and harmful applications.

Brown shares insights from his own work using AI models to develop poker solvers, illustrating how model reasoning and optimization abilities have improved across versions. While earlier models struggled and often produced misleading outputs, newer iterations can perform complex tasks zero-shot and optimize code significantly faster. However, he notes that models still lack “research taste” and cannot yet autonomously generate novel algorithms or fully replace human researchers, though this capability is gradually improving with each model release.

Finally, Brown reflects on the broader AI research landscape, emphasizing that progress is currently incremental rather than explosive. He argues against the notion of an imminent overnight intelligence explosion, citing the bottleneck imposed by the need for extensive test-time compute. He also discusses the potential of multi-agent systems and coordinated AI efforts to emulate human civilization’s cumulative knowledge-building. Despite intense competition among AI labs, Brown expresses optimism about shared awareness of risks and the collective goal of steering AI development toward positive outcomes, while encouraging the community to adopt more nuanced and compute-aware evaluation practices.