Claude AI Failed 650 Times…Then Beat The Human Record

An unreleased version of the AI Claude, after around 650 failures and motivated by non-expert encouragement, surpassed the human record by improving a mathematical bound related to the Riemann hypothesis, marking a significant advancement in mathematics. The project showcased AI’s evolving ability to generate and critically assess novel insights, highlighted the importance of transparency and verification, and underscored the need for new tools like Weights & Biases’ Weave to support AI development.

An unreleased version of the AI Claude was tasked with tackling the Riemann hypothesis, a famously difficult and unsolved problem in mathematics concerning the distribution of prime numbers. While Claude did not solve the hypothesis itself, it managed to improve a related mathematical bound beyond the existing human record, marking a significant advancement in the field. This achievement impressed many mathematicians, who regarded it as a major leap forward, even though the video’s narrator, a student, refrained from making deep mathematical judgments.

Interestingly, the process behind Claude’s success was unconventional. The AI was prompted not by expert mathematicians but by a non-mathematician who mostly sent messages of encouragement such as “keep going” and “believe in yourself.” Claude failed around 650 times before finally succeeding, highlighting a unique dynamic where motivational support played a key role in the AI’s persistence and eventual breakthrough. This approach suggests a future where AI progress in complex fields might be driven as much by human encouragement as by technical expertise.

The technical paper detailing Claude’s findings is publicly available, though it is quite complex. To aid understanding, Claude was also asked to explain its own results, and a formalized, machine-verifiable version of the proof was produced. This transparency allows others to verify and build upon the work. The video’s narrator also reviewed extensive transcripts from the project, revealing that Claude had internet access but did not rely on it during the critical breakthrough. The AI initially explored many incorrect paths but learned from these mistakes, with the first key result emerging after about 37 minutes of silence.

Claude itself expressed skepticism about the quality of its result, describing it as “too strong to be new,” which reflects the AI’s pattern recognition rather than human-like surprise. This moment underscores how AI systems are evolving to not only generate novel insights but also to critically assess their own outputs. The narrator reflects on the broader implications, noting that we are entering a new era where AI systems, built through human ingenuity, are actively pushing the boundaries of human knowledge.

Finally, the video highlights the need for new tools to support the development and debugging of large language model (LLM) applications. It promotes Weights & Biases’ Weave, a toolkit designed to help developers trace data flow and evaluate progress in AI projects. This call to action emphasizes the growing ecosystem around AI research and development, encouraging viewers to explore these resources as the field rapidly advances.