12 Days of OpenAI: Day 12

The final day of OpenAI’s 12-day event unveiled two new reasoning models, O3 and O3 Mini, showcasing significant advancements in AI capabilities, particularly in coding and mathematics, with O3 achieving a state-of-the-art score on the Arc AGI benchmark. The event emphasized safety and encouraged researchers to apply for testing access, with plans for a public launch in early 2024.

The final day of OpenAI’s 12-day event introduced two new models, O3 and O3 Mini, marking a significant advancement in AI reasoning capabilities. The event began with a recap of the launch of O1, the first reasoning model, which has received positive feedback for its ability to handle complex tasks. The new models, while not publicly launched yet, are available for public safety testing, allowing researchers to apply for access and contribute to the testing process. This initiative reflects OpenAI’s commitment to safety as their models become increasingly capable.

Mark, the head of research at OpenAI, highlighted the impressive performance of O3 on various technical benchmarks, particularly in coding and mathematics. O3 achieved a remarkable 71.7% accuracy on software style benchmarks, significantly outperforming O1. In competitive programming and mathematics, O3 also demonstrated superior capabilities, achieving scores that suggest it is nearing the performance level of expert PhD candidates. The discussion emphasized the need for more challenging benchmarks to accurately assess the capabilities of frontier models like O3.

The event also featured Greg from the Arc Foundation, who discussed the Arc AGI benchmark, which has remained unbeaten for five years. O3 achieved a state-of-the-art score of 75.7% on this benchmark, showcasing its advanced reasoning skills. When tested under high compute conditions, O3 scored even higher, surpassing human performance levels. This achievement represents a significant milestone in the pursuit of artificial general intelligence (AGI) and highlights the importance of rigorous benchmarks in evaluating AI progress.

Huran, an OpenAI researcher, introduced O3 Mini, an efficient reasoning model designed for cost-effective performance. O3 Mini supports adjustable reasoning efforts, allowing users to optimize for speed or accuracy based on their needs. Initial evaluations showed that O3 Mini outperformed O1 Mini in coding tasks while maintaining lower costs. The model also demonstrated strong performance in mathematical evaluations, indicating its versatility and effectiveness in various applications.

Finally, the event concluded with a discussion on safety advancements, including a new technique called deliberative alignment, which enhances the model’s ability to assess the safety of prompts. OpenAI encouraged safety researchers to apply for early access to O3 and O3 Mini for testing purposes, with plans for a public launch in early 2024. The team expressed excitement about the future of AI and the collaborative efforts to ensure the safe deployment of these advanced models, wishing everyone a Merry Christmas as they wrapped up the event.