OpenAI Withdraws Three AI-Written Math Papers a Day After Release, and 300 of 719 Results Have Lean Proofs
A.I. / news
OpenAI Withdraws Three AI-Written Math Papers a Day After Release, and 300 of 719 Results Have Lean Proofs
A sign error in one manuscript took down two dependent papers, while OpenAI's own README warns that unformalized results could have issues.

OpenAI withdrew three AI-generated mathematics manuscripts on Wednesday, a day after publishing a catalogue of 719 results from an unreleased internal model. A sign error invalidated an argument they share.
The company recorded the change in the dated Oct. 7 entry of history.md in its openai/math repository. It said "a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers." The page names no people and does not say who found the error.
The three manuscripts and the 14 repairs
The error sits in the first paper below. The other two depend on its construction.
| Withdrawn manuscript | Role |
|---|---|
| Algebraicity of Weil classes on split abelian eightfolds | Source of the sign error |
| Algebraicity of Kuga–Satake Correspondences for K3 Surfaces | Dependent paper |
| The rational Hodge conjecture for products of K3 surfaces | Dependent paper |
OpenAI also revised 14 other manuscripts: four on Lipschitz heights and Ashkin–Teller currents, six on Kähler minimal model programs and abundance, two on taming and hypersymplectic deformation, and one each on incompressible box transport and on the Birch–Swinnerton-Dyer formula. A further 13 were updated only to cite the revised editions. The withdrawn papers carry notices that link to archived copies.
The history page does not say whether any of the three had Lean proofs.
How much of the catalogue is machine-checked
OpenAI's README says 300 of 719 top-line results, about 42%, are formalized in Lean. The same page warns that "some of the unformalized results could have issues."
- Formalized in Lean300 results
- Not formalized419 results
Source: OpenAI, openai/math history.md, entry dated 2026-10-07, accessed 2026-10-08
The 419 is the remainder of 719 after the 300, not a figure OpenAI printed. The Oct. 7 update added six formalizations and five other supporting additions.
What OpenAI says produced the results
The README says most results came from one unreleased internal model, given roughly 4,000 posed problems and about three hours of ChatGPT Pro thinking per result. Two exceptions are the zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for CM abelian varieties. The zeta writeup was edited by a human for readability.
Implicator.ai reported that an OpenAI spokesperson said nearly every result came from a single prompt to a single agent, and that many results are not yet understood by OpenAI's own mathematicians. Both claims are OpenAI's. With the model unreleased, nobody outside the company can test them.
The disclosure standard the mathematicians set
On Sept. 29 the Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, asked labs to disclose model names, prompts, summarized reasoning, and per-result time and compute cost. Implicator reported that OpenAI published no prompts and gave only an average compute figure. It counted ten abridged reasoning summaries; the README describes summaries for nine families.

Andrew Sutherland of MIT said claims of one-shotting problems with a single agent should be treated as "unverified" until the model is released and results are replicated, according to Implicator. Daniel Litt of the University of Toronto said "I see no reason why we should ask the company to keep them secret from us." Both were quoted by Implicator, which is the only source here for them.
This is not the first audit of OpenAI's mathematics. An arXiv paper by Mikołaj and Krzysztof Sienicki reviewed ten results OpenAI announced on Aug. 1 and found no confirmed substantive error remaining in a principal result, with one apparent polarity error withdrawn after a lost overbar was recovered. The arXiv page lists no affiliations for the authors.
OpenAI has not said how it will handle the 419 unformalized results. Its README says corrections will arrive as new versions, with earlier versions kept accessible. The lab's earlier math claim is covered in the Navier-Stokes data dispute, and rival labs' model pricing in the Claude Opus 5.5 report.
Sources
More in A.I.
- 01OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 02Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M EmbedderThe Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.
- 03Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on OneThe Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.
- 04Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One LineAlibaba's 27B dense model scores 61.7 on SWE-bench Pro by its own table. The card does not say what it was trained on or what hardware it needs.