OpenAI Claims AI Solved 700+ Math Problems in One Prompt

OpenAI's Surprise Math Drop
Imagine giving an AI a single math problem, walking away for three hours, and coming back to a brand-new mathematical proof. That is essentially what OpenAI claims to have achieved.
The artificial intelligence company recently published 722 math manuscripts on GitHub. According to OpenAI, almost all of these papers came from a single unreleased AI model responding to a single prompt. While that sounds like a massive technological leap, many top mathematicians are asking for hard proof before celebrating.
Breaking Down the Numbers
To understand what happened, it helps to look at how OpenAI gathered these results. The company did not just solve 722 standalone problems. Instead, they organized the manuscripts into 372 result families.
A single family might bundle a main theorem together with supporting arguments, alternative proofs, or logical consequences. Here is how the process worked:
- OpenAI fed roughly 4,000 open math problems into its internal AI model.
- The team selected and published only the outputs they judged significant.
- On average, each result required roughly three hours of ChatGPT Pro thinking compute.
This setup looks very different from OpenAI's earlier experiment on the Navier-Stokes problem. That project relied on 10,000 coordinating AI agents working together over 88 hours.
Why Mathematicians Want Receipts
Despite the headlines, many experts urge caution until independent researchers can test the AI themselves. Andrew Sutherland, a mathematician at MIT, told Scientific American that one-prompt claims remain unverified until the model is publicly released.
Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified. We should ask for receipts.
A major issue involves verification. Only 162 of the 722 manuscripts—about 22%—have been mechanically checked using Lean, a specialized software tool that verifies logical steps mechanically.
OpenAI itself admitted that unformalized papers could contain errors. Passing a Lean check confirms that the math logic follows the written code, but it does not prove that the statement matches the original problem or that the discovery is truly useful.
Alien Math and Locked Repositories
As researchers began examining the papers, some found the AI's logic puzzling. Researcher Dmitry Rybin looked at OpenAI's proof regarding the chromatic number of a plane and described it as unexpected alien math that introduced strange, unpredicted steps.
Others pointed out practical issues with OpenAI's GitHub repository. Mathematician Keith Adler attempted to formalize OpenAI's proof of Saxl's Conjecture in Lean 4, but noted that OpenAI had disabled repository issues and pull requests.
If you publish 722 manuscripts and ask for Lean formalisations, you need somewhere for people to send them.
Breakthroughs and the Human Element
The Institute for Advanced Study (IAS) in Princeton highlighted broader concerns about human understanding in modern science. In a statement, IAS noted that AI can now output complex math that humans cannot easily verify or explain.
However, Professor Abhishek Saha viewed the event more positively, calling it a very big day for mathematics. He noted that while most papers offer solid progress on existing problems, one paper addresses a major landmark: the Quasi-Riemann Hypothesis.
Transparency and Competitor Approaches
Transparency remains a central debate. IAS previously recommended disclosing full details, including:
- The exact AI model name
- The specific prompts used
- Summarized chain-of-thought reasoning
- Computation time and precise compute costs
OpenAI provided average compute numbers and 10 short reasoning summaries, but omitted the prompts. Meanwhile, mathematician Daniel Litt from the University of Toronto argued that sharing open math answers publicly benefits the community regardless of the prompts.
By contrast, rival company Anthropic recently took a different approach. Anthropic released 13 million lines of Lean code on GitHub to verify Andrew Wiles' iconic 1995 proof of Fermat's Last Theorem, focusing on verifying known human breakthroughs rather than generating unverified new claims.
OpenAI plans to add more Lean formalizations over time. For now, the math community continues to examine the 162 checked papers while waiting for full access to the underlying model.
Latest blog posts

XRP Ledger Upgrade Arrives as Price Awaits Next Move
The XRP Ledger launches its highly anticipated Batch amendment, while XRP trades around $1.39 following macroeconomic pressure from Federal Reserve updates.

XRP Ledger Payment Volume Drops by 400,000 Transactions
The XRP Ledger saw a sudden drop of 400,000 daily payments alongside broader crypto market volatility. Here is what technical levels reveal.

Hyperliquid Whale Reloads After $69M Ethereum Liquidation
A crypto whale lost $69.69 million in an Ethereum liquidation on Hyperliquid, then deposited $10 million in fresh margin 30 minutes later.