OpenAI Math Proofs: 372 AI Results Land on GitHub
OpenAI has released 372 new mathematical results produced by one of its internal frontier models. It did not send them to academic journals. Instead, it posted all of them in a public GitHub repository, along with revision logs and citations. The company says each result either solves an open problem or makes substantial progress toward solving one.
The collection covers a wide range. Some results improve major computer algorithms. Others relate to the Riemann hypothesis, one of the best-known unsolved problems in mathematics. According to OpenAI, the same model has already produced a solution to a Navier-Stokes problem, and that solution has been under formal review for weeks.
One prompt, one agent, three hours
The production method is what stands out most. OpenAI says nearly every result came from a single prompt given to a single AI agent, although some needed more than one attempt. On average, each result used about three hours of ChatGPT Pro Thinking compute.
The Navier-Stokes work was very different. It required a swarm of 10,000 agents and millions of dollars in compute. The new batch suggests the cost per result has fallen sharply, at least for the problems OpenAI chose to attempt.
OpenAI also published notes on its methodology:
- summaries of the model's reasoning process
- statistics on how many problems the model attempted
- estimates of compute costs
Lean as a review shortcut
Many of the proofs come with formalizations in Lean, a programming language designed for mathematical proofs that a machine can check. OpenAI says more formalizations are coming.
The reason is practical. Hundreds of AI-generated proofs could easily overwhelm the capacity of human mathematicians to review them by hand. Other fields have run into similar problems. Google, for example, paused an open source bug bounty after being flooded with AI-generated reports. Machine-checkable proofs are OpenAI's attempt to avoid that kind of bottleneck in mathematics.
Skipping the journals
Choosing GitHub over peer-reviewed journals sends a message. It implies that traditional academic publishing is too slow for this volume of potentially new results.
OpenAI consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (IAS), the research center in Princeton. Fields Medal winner Timothy Gowers is among its members. OpenAI only loosely followed the group's public recommendations. It did not publish any of its prompts, and it shared only average compute costs rather than figures for each problem.
The company also set a limit before the consultation began. The mathematicians could advise on how the results were communicated, but not on whether they were produced or how quickly.
OpenAI says it will fund workshops and conferences focused on understanding AI-produced results. It admits that its citations and presentation need work. It also says it is preparing a responsible release of the model to "directly empower scientists with state-of-the-art capabilities."
Correct is not the same as meaningful
Lean can confirm that a proof is logically correct. It cannot tell whether a result is original or whether it matters mathematically. OpenAI is betting that its results push the boundaries of human knowledge. It is not yet clear whether mathematicians agree, and early reactions range from excitement to frustration.
Some of the field's most prominent figures have already raised concerns. In an open letter titled "A Severe Misalignment of AI in Mathematics," 25 Fields Medal winners described a deep disconnect between the goals of the AI industry and the goals of mathematics. In their view, solving problems is only a tool and a proxy. The real aim is conceptual understanding and insight. They argued that mass-producing true statements could destroy fertile ground instead of creating new ideas, and that the damage would spread to other fields.
Gowers has warned that within one to two decades, the mathematical literature could grow enormously while no human community remains that truly understands it. Terence Tao, also a Fields Medal winner, has said that training young mathematicians should emphasize the human side of the discipline and strictly limit the use of AI tools, so that real learning and understanding survive.
Our Take
This release appears to be as much about process as about mathematics. A single agent running for three hours per problem changes the economics of producing proofs. If that cost holds, output will keep growing faster than human reviewers can keep up. Lean helps with one part of the problem, which is correctness. The harder questions are whether a result is relevant and whether it is original, and those still depend on people who are already stretched thin.
The ground rules also deserve attention. OpenAI asked leading mathematicians for advice but did not let them influence whether or how fast the work was produced. It also withheld its prompts and per-problem costs. That leaves the research community reacting to the release rather than shaping it. Similar tension is showing up elsewhere in academia, as our coverage of an MIT report on eroding faculty trust has shown.
There are several things to watch next:
- Whether independent mathematicians confirm that some of the 372 results are truly new and significant.
- How the Navier-Stokes review concludes.
- Whether OpenAI's promised model release gives researchers a tool they can steer themselves, as harnesses like BootLoops aim to do with other models.
If the release only adds to an unread pile of proofs, the Fields medalists' warning may turn out to be the more accurate prediction.
