
722 papers from a closed OpenAI model on math. How many did a machine actually check?
Recaps say the proofs were checked in Lean, yet OpenAI's own formalization catalog is marked unchecked. We compared the October 6 release with what AGMAI asked.
In this article5
On October 6, 2026, OpenAI published 722 mathematical manuscripts, grouped into 372 result families, in the openai/math repository. They came from an unnamed internal model: it was given about 4,000 problems and spent on average three hours of ChatGPT Pro-level compute per result. The claims include a quasi-Riemann hypothesis (no zeros of the zeta function for Re s > 7/8), the bound ω ≤ 9/4 for matrix multiplication and Barnette's conjecture. By our count, 235 of the 372 families link to Lean, where a program checks the proof. But the formalization catalog itself is marked «Partial progress» and review: unchecked, and the README says the unformalized results could have issues. On September 29 the AGMAI advisory group of mathematicians asked for the model, prompts, reasoning, time and cost to be published for every result. OpenAI delivered part of that.
722 manuscripts, 372 result families, a single commit titled Initial commit. That is what OpenAI's math release looked like on October 6: several weeks of output from an internal model landed in a public repository in one day. On September 21 the company had been talking about «more than 100» solved open problems.
A month ago we looked at how a swarm of 10,000 OpenAI agents went after the Navier–Stokes problem, and the real story there turned out to be the machine checker: Lean. The question for the new release is the same: where was the checker, and where is it missing.
What OpenAI published on October 6
According to the repository README, almost every result went through one procedure: the model was given open problems, about 4,000 of them, and spent on average three hours of ChatGPT Pro «thinking» per result. The answers were then grouped into families and filtered by significance. The exceptions are named outright: the zero-free region for the zeta function and the Hodge conjecture for CM abelian varieties did not follow the general pipeline, and the text on Re(s) > 11/12 was edited by humans.
The names on the list are big ones. A quasi-Riemann hypothesis: the zeta function and all Dirichlet L-functions have no zeros for Re s > 7/8. A matrix multiplication exponent of ω ≤ 9/4 over the complex numbers. Barnette's conjecture on Hamiltonian cycles. Counterexamples to several of Kaplansky's conjectures. The Birch and Swinnerton-Dyer formula for a class of elliptic curves. OpenAI says it will release the model itself «responsibly», with no date.
One more detail shows up only in the folder names. Every manuscript carries a date, and 564 of the 722 are dated September 23–27, with 370 of them falling on just two days, September 23 and 24. Another 112 are marked October 5, the day before publication.
result families with a Lean link in CONTENTS.md. 137 have none, including the Birch and Swinnerton-Dyer formula
«Has Lean» does not mean «verified»
Recaps say that «many of the proofs are formalized». That is true, but the wording hides three caveats OpenAI itself left in the repository.
First: 235 families link to Lean, and 137 rest on text alone. The README says it plainly: «Some of the unformalized results could have issues», and promises quick fixes.
Second: formalization often covers the main theorem, not the whole paper. The description of the Lean part for the quasi-Riemann hypothesis says the 7/8 bound is proved formally, while the later appendices of the paper are not formalized.
Third: the lean/formalization.yaml catalog lists 162 papers with a formalized main result. The scope field says «Partial progress», and the review field says status: unchecked. Lean guarantees that the derivation is sound. Whether the formal statement is exactly the claim in the paper's title is something a human has to check, and by OpenAI's own marking that check has not been done.
What mathematicians asked for and what OpenAI did
On September 21 an advisory group, AGMAI, was formed at the Institute for Advanced Study in Princeton: nine mathematicians. It has no say over how fast OpenAI pursues mathematics, only advice on how to publish. On September 29 it issued recommendations based on more than 600 responses from colleagues, and opened with a sentence that is hard to read any other way: «we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models».
The release came a week later. We compared them point by point.
| AGMAI recommendation, September 29 | What the October 6 release has |
|---|---|
| Model name for every result | «Unreleased internal OpenAI model», no name |
| Prompts | Not in the repository |
| Reasoning summary for every result | For 10 of 372 families |
| Compute time and cost | An average: three hours of ChatGPT Pro per result, no dollar figure |
| A repository not controlled by an AI lab | OpenAI's GitHub; independent venues are «being explored» |
| Formalization wherever possible | 235 families with a Lean link, catalog marked unchecked |
The group's own response on October 7 was measured: the release is «the beginning, not the completion, of the process of human understanding». It reaffirmed its recommendations.
How to check an OpenAI proof yourself
The release's strong side is elsewhere. For the formalized results OpenAI included tasks for comparator, a Lean FRO tool that checks a proof against a statement written down in advance. Anyone can repeat the check. The README advises against building the whole library at once: it has 122,140 .lean files, and the build can hit the vm.max_map_count system limit, so results are checked one at a time.
lake updatelake exe cache getlake env comparator ComparatorChallenges/QuasiRiemannHypothesis.json
The next day an OpenAI forum member, davegoldblatt, did exactly that, with Claude Code running the check. The result: 2,924 modules built with no errors or warnings, and the proof was accepted both by comparator and by the independent nanoda kernel. His own caveat is more precise than any headline: «this checks the Lean proof, not the paper». One run, one machine, and what was verified is the formal text, not the paper.
What people who work with agents should take from this
Our opinion, and you can argue with it: the main result of the release is not the 722 manuscripts but how easy it has become to tell the verified from the claimed. Where there is a comparator task, checking the formal part comes down to one command, and an agent can run it. Where there is not, all that is left is the company's reputation and «could have issues». We will be wrong if independent mathematicians find serious errors in the formalized part over the coming months: then the line runs somewhere else.
For your own agents the takeaway is the same as with the Navier–Stokes swarm. Tests, types and the build play the role of your Lean, and a separate reviewer plays the role of comparator: it checks that what was asked for got done, not just that the code compiles. How to build that kind of check for an AI feature in a product is shown in our breakdown of an eval and hillclimb for the Claude API. Splitting the executor and the checker into separate subagents is what the Claude Code course teaches.
Figures as of October 7, 2026
The count of families with a Lean link is based on CONTENTS.md in the first version of the repository. OpenAI promises to add formalizations and keep a version history, so the verified share will grow. Check the current README and formalization.yaml.
Sources7expand
- OpenAI, «Sharing AI progress in mathematics», October 6, 2026 — https://openai.com/index/sharing-ai-progress-in-mathematics/
- OpenAI, openai/math repository: README, CONTENTS.md, lean/formalization.yaml, lean/docs/003.md, October 6, 2026 — https://github.com/openai/math
- AGMAI, «Responsible Release of AI-Generated Mathematics», September 29, 2026 — https://agmai.org/general-sep29/
- AGMAI, «On OpenAI's Release of Mathematical Results», October 7, 2026 — https://proofsandprompts.com/2026/10/07/on-openais-release-of-mathematical-results/
- davegoldblatt, re-check of family 003, OpenAI forum, October 7, 2026 — https://community.openai.com/t/first-look-at-mathematics-manuscripts-from-an-internal-frontier-model-at-openai/1403886
- TechCrunch, «OpenAI forms math advisory group as its AI resolves more than 100 open problems», September 21, 2026 — https://techcrunch.com/2026/09/21/openai-forms-math-advisory-group-as-its-ai-resolves-more-than-100-open-problems/
- Gizmodo, «OpenAI Dumps 377 New Math Results on GitHub», October 6, 2026 — https://gizmodo.com/openai-dumps-377-new-math-results-on-github-publishes-hand-wringing-blog-post-2000822613
Read next
10,000 agents did in 88 hours what nobody managed in 90 years. Why did their creators immediately ask for the brakes?September 14, 2026
Welcome to the AGI era? What OpenAI is really keeping quiet about GPT-6 AstraSeptember 7, 2026
98.9% on the exam, 90.5% on questions the model never saw. Which number should you trust for your AI feature?October 5, 2026
Comments