OpenAI’s latest model, Astra, has reportedly solved ten longstanding and complex mathematical problems by producing fully verified proofs in Lean 4, enabling independent validation and marking a significant milestone in AI-assisted mathematics. While these achievements demonstrate impressive progress, questions remain about the authenticity of the problems solved, the broader capabilities of AI in general mathematical problem-solving, and the implications for authorship and credit in research involving AI.
Eighteen months ago, OpenAI’s model made significant strides in mathematics by solving problems previously deemed unsolvable, culminating in their latest model, Astra, which reportedly cracked ten longstanding math problems that had stumped researchers for decades. These problems were not typical exam questions but deep structural gaps in mathematical understanding, some unsolved since as far back as 1946. OpenAI published the proofs in Lean 4, a formal proof language that compiles code to verify every logical step, ensuring no gaps or hand-waving in the reasoning. This transparency allows anyone to independently verify the correctness of the proofs without relying solely on OpenAI’s claims.
The use of Lean 4 means the proofs are rigorously checked by a compiler that demands explicit logical steps, with no room for ambiguity or skipped reasoning. A zero count of “sorry” commands—placeholders that bypass proof checks—confirms the completeness of each proof. However, while the machine verifies logical consistency, it cannot assess whether the original problem statements accurately represent the historic open questions. This raises concerns about whether the AI solved the genuine problems or slightly altered versions that are easier to prove, a critical distinction in validating the breakthrough.
Independent verification has begun, with notable mathematicians like Tim Gowers reviewing earlier results from the same model family and endorsing their validity. Additionally, rival AI labs have replicated some of Astra’s proofs using their own models, suggesting that these breakthroughs might not be exclusive to OpenAI’s technology. Despite this, real-world blind tests on fresh, uncurated research problems show that AI models still struggle, indicating that while Astra excels on selected problems, general mathematical problem-solving remains a challenge.
The announcement has sparked debate within the academic community, not over the correctness of the proofs—which remain unchallenged—but over issues of credit and authorship when AI contributes significantly to research. This has led to discussions like the Leiden Declaration, addressing how to attribute work done by AI. Meanwhile, some skepticism persists among AI models and researchers about the broader implications of these results, with many viewing the claims as extraordinary and requiring thorough peer review.
In conclusion, Astra’s achievements represent the most verifiable and transparent mathematical milestone reached by AI to date, though it does not yet constitute artificial general intelligence (AGI). The proofs stand as a strong claim that can be independently tested and either confirmed or refuted by the mathematical community. The ongoing peer review process will determine the true impact of these results, marking a significant step forward in AI-assisted mathematics while highlighting the complexities of evaluating and integrating AI contributions into human knowledge.