OpenAI Posts Hundreds of AI-Written Maths Papers, Keeps the Prompts
A.I. / news
OpenAI Posts Hundreds of AI-Written Maths Papers, Keeps the Prompts
The repository holds 719 manuscripts by README count, a 42% Lean share by OpenAI's measure and 22% by Decrypt's, and no model name.

OpenAI published a public GitHub repository on October 6 holding AI-written mathematics manuscripts that it says came from an unreleased internal model, and withheld the model's name, the prompts and the per-problem compute. The repository's own README warns that "some of the unformalized results could have issues."
The repository is named math, sits under OpenAI's account and carries an Apache-2.0 licence, according to Let's Data Science. As of Saturday its README tallies 719 manuscripts in 372 families of related results. Decrypt and Let's Data Science both count 722. The two sources differ by three papers, and the unit is manuscripts, not solved problems.
What is in the repository
OpenAI posed roughly 4,000 problems and published the outputs it judged significant, Decrypt reported. The average result used about three hours of ChatGPT Pro thinking, per OpenAI.
An OpenAI spokesperson told Scientific American that nearly all results came from one prompt handed to one agent, though some took several attempts, as Let's Data Science and Decrypt relayed it. The company's September 8 claim on the Navier-Stokes problem reportedly used about 10,000 coordinating agents for 88 hours.
| Entry | Claim in the repository | Caveat reported |
|---|---|---|
| 003 | No Dirichlet L-function has a zero with real part above 7/8 | OpenAI calls it "resolving the quasi-Riemann hypothesis"; it is weaker than the Riemann Hypothesis |
| 102 | Proof of Khot's Unique Games Conjecture | Not peer reviewed |
| 109 | Integer multiplication in O(n(log n)^(1−κ)) time, κ = 2^(−182) | Speedup too small to matter on real hardware |
How much of it is checked
The README says the collection "includes results at different stages of verification" and that "not all have accompanying Lean formalizations". It states that "~42% top-line results" are formalized.
Decrypt counts differently and found 162 of the 722 papers, about 22%, with a Lean-formalized main result. The two figures use different denominators and OpenAI's page does not reconcile them.
A passing Lean check proves a proof follows from the statement as encoded. It does not show that the statement matches the original problem, which still takes a human reader.

What was withheld
The Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, was announced on September 21. On September 29 it recommended that labs disclose the model, prompts, chain-of-thought summaries, time taken and compute cost, Let's Data Science and Decrypt reported. Its nine members include Fields Medalists Timothy Gowers and Martin Hairer and the physicist Edward Witten.
OpenAI released average compute figures and abridged reasoning summaries for 10 results. It released no prompts and no model name, and said it is working to release the model responsibly.
MIT mathematician Andrew Sutherland said single-agent, one-shot claims stay unverified until the model is out. "We should ask for receipts," he said.
Reaction
Levent Alpöge, a researcher at Anthropic, called it on X "obviously the most significant moment in mathematical history" and later added caveats. OpenAI chief executive Sam Altman wrote that he was "looking up at the stars with extra awe tonight."
Dmitry Rybin questioned an OpenAI proof on the chromatic number of the plane, Decrypt reported, finding a key step unexpected. Keith Adler said the repository has Issues turned off. The GitHub page shows no Issues tab, and its README says only collaborators can open pull requests.
Anthropic posted a Lean-checked proof of Fermat's Last Theorem in September, about 13 million lines, per Decrypt, but it formalized Andrew Wiles's 1995 result rather than claiming a new one. Anthropic has its own AI-accuracy story on this site, and OpenAI's handling of research information is the subject of another.
The README says the team is "exploring community-hosted repositories" for the material. OpenAI has not given a date for releasing the model.
Sources
More in A.I.
- 01Nvidia Is Reportedly Weighing a Takeover of Reflection AI, Whose 501B Beam Model Has No Public Weights YetThe Financial Times reported early-stage talks on October 10. Reflection's own benchmark table has Beam behind Kimi K3 and Qwen 3.8 Max on every coding test where all three report.
- 02Qwen3.8-27B Is Apache 2.0. The 2.4T Max Weights Have No Published Licence TexteWeek reports a $50 million revenue trigger on the Max licence; RuntimeWire calls the terms unresolved, and the Hugging Face card gives only a name.
- 03Cloudflare Open-Sources Clef, a 27B Drop-In for TypeSafe's JevThe Apache 2.0 decision models cut median latency from 524 ms to 39 ms on Cloudflare's own tests, and lose to Jev by 30 points on GPQA Diamond.
- 04Microsoft-Decision-1 Is a Post-Trained Qwen3.5-9B at $0.042 per Million Tokens, With No Licence GivenMicrosoft's launch post claims a win across 36 benchmarks and 35 times the speed of GPT-6 Sol, and every figure in it is the company's own.