Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

Nicholas Kang and Michael Aaron from Google DeepMind discuss the challenges of fragmented and outdated AI evaluations and present solutions like hackathon platforms, standardized agent exams, Game Arena, and collaborative Benchmarks to create open, scalable, and transparent AI assessment tools. Their goal is to democratize AI evaluation, foster community involvement, and enable continuous, meaningful benchmarking that keeps pace with rapid AI advancements.

In their talk, Nicholas Kang and Michael Aaron from Google DeepMind discuss the challenges and solutions related to AI agent evaluations at scale. They begin by highlighting the current problems with AI evaluations: they are scattered, decentralized, and quickly become outdated. Many benchmarks emerge daily, but tracking and understanding them is difficult, and leaderboards often become irrelevant as authors move on. Additionally, evaluations lack transparency and accessibility, with configurations and testing conditions often unclear, leading to inconsistent results. They emphasize the disparity between the small number of AI researchers creating benchmarks and the vast global population expected to benefit from AI, which results in uneven AI capabilities across different domains.

To address these issues, the team is developing several solutions. One is a hackathon platform that enables anyone to host and participate in hackathons focused on building AI benchmarks. This approach channels diverse expertise and creativity, encouraging open-source contributions to tackle important but underrepresented problems, such as a wastewater treatment benchmark created by an industry expert. Another solution is standardized agent exams, designed to democratize AI evaluation by allowing users to easily test their AI agents and compare results on leaderboards. This tool aims to bridge the gap between sophisticated research labs and consumer-level AI developers, promoting safer and more reliable AI deployment.

Michael Aaron then delves into two additional projects: Game Arena and Benchmarks. Game Arena is a platform where AI models compete in PvP games like poker, chess, and Werewolf to provide continuous, unsaturated evaluation through Elo-style scoring. This method prevents benchmark saturation by ensuring there is always a winner and a loser, allowing ongoing performance measurement. The platform is open source and integrates with Kaggle, but faces challenges such as high computational costs, maintaining community engagement, and ensuring fair and meaningful comparisons over time as models evolve.

The Benchmarks platform allows the community to build, run, and share AI evaluations collaboratively. It supports tasks with assertions and LLM-based judging, aggregating results into benchmarks that can be openly accessed and analyzed. While this fosters community involvement and transparency, challenges remain in motivating users to create high-quality evaluations, handling the complexity of testing AI agents versus models, and managing the rapid release and deprecation cycles of AI models, which complicate longitudinal comparisons.

Overall, Kang and Aaron emphasize the importance of open, accessible, and scalable evaluation tools to democratize AI progress and ensure equitable benefits across diverse fields. They invite the community to engage with their platforms on Kaggle and contribute to advancing AI evaluation methodologies. Despite the technical and logistical hurdles, their work aims to create a more transparent, collaborative, and dynamic ecosystem for AI benchmarking that can keep pace with the fast-evolving AI landscape.