Nine Mathematicians Tell AI Labs How to Release Machine-Made Proofs After OpenAI Asked for Advice
A.I. / news
Nine Mathematicians Tell AI Labs How to Release Machine-Made Proofs After OpenAI Asked for Advice
The Advisory Group on Mathematics and Artificial Intelligence published its recommendations on Sept. 29, after more than 600 replies from the field.

The Advisory Group on Mathematics and Artificial Intelligence, a nine-member panel of mathematicians, published a statement on Sept. 29 telling AI labs how to release mathematical results their systems produce. The document, titled Responsible Release of AI-Generated Mathematics, draws on more than 600 replies from the mathematical community.
The group came together after OpenAI approached some of its members about an external advisory board, Terence Tao wrote on his blog on Sept. 21. He described the immediate task as advising OpenAI on releasing "a large number of significant results in mathematics that they report have been produced by their internal model."
Who is on the panel
The group's site lists nine members: François Charles of ENS-PSL, Camillo De Lellis of the Institute for Advanced Study and GSSI, Timothy Gowers of the Collège de France and Cambridge, Martin Hairer of EPFL and Imperial College London, Nikhil Srivastava of Berkeley and the Simons Institute, Ulrike Tillmann of Oxford, Ravi Vakil of Stanford, Edward Witten of the Institute for Advanced Study, and Melanie Matchett Wood of Harvard.
The site says members do not accept payment for the work and that the group operates independently of any AI company.

The three principles
The statement opens with three principles. Labs should release significant results responsibly and promptly. A lab that releases large amounts of mathematics nobody yet understands must take responsibility for human understanding following, including funding. That understanding must stay "organic and community led."
The statement warns that proprietary models create "a two-tier system where labs outrun the rest of the field, effectively alienating the mathematical community."
What labs are asked to do
Results that the lab's own staff understand should follow ordinary academic norms: preprints, peer review, conference talks. Results nobody understands get a two-step process.
| Step | What the lab does |
|---|---|
| I. Initial release | Improve exposition with LLMs and cite related literature; deposit in a repository the lab does not control; disclose model name, prompts, chain of thought, computation time and cost; formalise proofs where possible; report failed attempts |
| II. Supporting understanding | Fund conferences, workshops or expository writing; support postdocs and students in long-term work groups |
The statement also asks labs to "refrain from treating the release of mathematical results as marketing vehicles."
Reaction on Hacker News
Reaction on Hacker News split. One commenter, kingstnap, called most of the document uncontroversial and said the contested part is a request that long-standing problems not be used as benchmarks for proprietary models. Another, simianwords, argued OpenAI should be free to publish whatever it likes and that mathematicians can ignore it.
A third commenter, Animats, summarised the statement as asking AI companies to pay human mathematicians to understand AI-generated results, which they called an unusual ask.
What is not known
The statement does not name OpenAI, and the Tao post is the only fetched source linking the group to a specific lab. Neither document lists the results OpenAI intends to release or says when. The group has not said whether OpenAI accepted the recommendations.
The statement is a recommendation with no enforcement. Any lab can ignore it, and the group's site invites further community input at agmai.org/input.
OpenAI's feed for Sept. 29 and 30 lists GPT-6.1 Sol and dots, neither of them a mathematics release.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.