OpenAI’s Math Flood Leaves Researchers Racing to Verify It
OpenAI’s Math Flood Leaves Researchers Racing to Verify It
The dispute had been building since September, when OpenAI said an internal model had resolved more than 100 longstanding problems and, after criticism over its earlier Navier-Stokes announcement, turned to the new Advisory Group on Mathematics and Artificial Intelligence. The group urged labs to publish promptly, disclose methods and avoid using mathematical results as marketing. “Refrain from treating the release of mathematical results as marketing vehicles,” it warned.
On Tuesday, OpenAI raised the stakes, publishing 722 manuscripts across 372 families of findings. The company said it had consulted the independent advisory group on how to release them, a process echoed in an OpenAI announcement shared by Lilian Weng on X. For advocates, the cache is evidence that AI can accelerate research rather than merely imitate it. Toronto mathematician Dan Litt called the new work “great for mathematics,” while warning that society must continue supporting human expertise if it wants to benefit from the results.
But the scale of the release became the point of conflict. The Association for Human Mathematics called it “not a demonstration of scholarship, but a demonstration of power,” arguing that hundreds of simultaneous claims can overwhelm peer review. Skeptics also question whether the findings are genuinely independent of researchers’ prior work and whether they can be trusted before formal verification.
Those concerns sharpened as errors surfaced: OpenAI withdrew at least three published solutions by Thursday. Its release offered reasoning summaries for only 10 manuscripts, while critics noted gaps between natural-language arguments and Lean formalizations in a Navier-Stokes-related proof. Harvard’s Melanie Wood captured the central objection: at release, “there is not human understanding of them, and now the work begins.”
The promise remains real, but so does the bottleneck. As Stephen Wolfram put it, AI can produce “a trillion theorems easily”; the harder question is which ones anyone will care about.
Write a comment