OpenAI's GPT 6 Astra Is Near "AGI" Level?

OpenAI’s internal model Astra has made significant strides by solving 10 longstanding mathematical problems with formally verified proofs, marking a major advancement in AI-assisted research, though it still relies on human input and struggles with open-ended, creative problem formulation. While Astra demonstrates impressive problem-solving capabilities, it has not yet achieved artificial general intelligence (AGI), as it cannot independently identify or pursue novel research questions.

OpenAI recently announced that their internal model, Astra, solved 10 longstanding open problems across eight diverse fields of mathematics, including group theory, quantum complexity, and coding theory. These problems had remained unsolved for years or even decades, with one example being the construction of the first non-sofic group, a question open for 27 years. Unlike previous AI claims, OpenAI published the formal proofs in Lean 4, a formal proof verification system, allowing anyone to mechanically verify every logical step without gaps or skipped parts. This transparency marks a significant advancement in verifying AI-generated mathematical results.

However, while Lean verifies the logical correctness of the proofs, it cannot confirm whether the formalized problems perfectly match the original human mathematical questions. Translating complex conjectures into formal code is a manual and challenging process, leaving room for potential weakening or misinterpretation of the original problems. Independent mathematicians, including Fields medalists and experts who previously scrutinized OpenAI’s claims, have reviewed and supported Astra’s results, lending credibility to the announcement. This contrasts with earlier misrepresentations by OpenAI, making the current claims more trustworthy.

The computational cost reported by OpenAI for generating these proofs was approximately $2,000, but this figure only accounts for successful runs and excludes numerous failed attempts, leaving the true cost unknown. Additionally, human researchers played a role in polishing and preparing the final papers, raising questions about the extent of human versus AI contribution. Interestingly, Anthropic’s publicly available model, Fable, reproduced half of Astra’s proofs within a day without internet access, suggesting that multiple AI systems are approaching similar capabilities in solving well-defined mathematical problems.

Despite these breakthroughs, Astra and similar models still struggle with open-ended research problems that require formulating new questions, recognizing valuable directions, and navigating dead ends—tasks that human mathematicians excel at. Tests involving live, ongoing research problems and specially designed benchmarks show that current AI models cannot yet independently drive the research process or generate genuinely novel mathematical insights without human guidance. This distinction highlights the difference between AI’s ability to solve given problems and its capacity for autonomous, creative problem formulation.

In conclusion, while Astra’s achievements represent a major step forward in AI-assisted mathematical research, they do not yet constitute artificial general intelligence (AGI). The model excels at searching vast solution spaces for answers to well-defined problems but lacks the broader cognitive abilities to identify which problems matter or to innovate independently. The real milestone may come when AI can autonomously select, formulate, and pursue meaningful research questions, marking a shift from problem-solving to genuine scientific discovery. Until then, Astra’s success signals that AI is moving from merely solving problems to actively participating in research, a profound evolution in its capabilities.