Math Advisory Group Asks OpenAI to Formalize Proofs and Publish Prompts, With No Way to Enforce It
A.I. / news
Math Advisory Group Asks OpenAI to Formalize Proofs and Publish Prompts, With No Way to Enforce It
Nine mathematicians published release norms on September 29 for AI-generated proofs no human yet understands. OpenAI has already said the group will not advise on its pace.

The Advisory Group on Mathematics and Artificial Intelligence published recommendations on September 29 for how AI labs should release proofs that no human yet understands. For those results, it asks the lab to disclose the model name, the prompts used, a summarized chain of thought, the time taken and the estimated cost of computation.
The nine-member group is hosted at the Institute for Advanced Study in Princeton. Its recommendations page says the text follows feedback from the mathematical community, and refers to over 600 replies. The page does not name individual signatories.

Group member Timothy Gowers of the Collège de France and Cambridge, at the GOSIM conference in Paris on May 5.
What the September 29 document asks for
The document splits AI-produced mathematics in two. Category 2.A covers papers a human mathematician fully understands. For those, the page says the mathematician "should post a preprint, submit a paper for peer review at a journal, and give talks", which is the ordinary route.
Category 2.B covers results with no human understanding yet, and it is where the new requirements sit. A first step covers the release itself. A second step covers what the lab owes the community afterwards.
| Step | What the lab is asked to do | Attached to |
|---|---|---|
| Repository | Deposit results in scholarly repositories "not controlled by any AI lab" | Release |
| Disclosure | Publish model name, prompts, summarized chain of thought, time and estimated compute cost | Release |
| Formalization | Formalize the proof, meeting community standards | Release |
| Understanding | Provide support, including funding, for human understanding | After release |
The page also warns of a two-tier system in which "labs outrun the rest of the field, effectively alienating the mathematical community from its own discipline", and asks for broad access to the models involved.
Why OpenAI is the immediate audience
The group exists because of OpenAI. In its September 21 announcement, the group said OpenAI approached certain members about an external advisory board, after which they set up an independent organization and invited others. The announcement says the group is independent of any AI company and that members are not paid.
The named task is advising OpenAI "on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model." TechCrunch reported that OpenAI puts that number at more than 100 open problems, on top of its claimed solution of the Navier-Stokes Millennium Prize problem announced on September 8.
The members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood.
What the group cannot do
The recommendations page contains no enforcement mechanism and no consequence if a lab ignores it. OpenAI has also limited the group's remit. TechCrunch quoted OpenAI as saying "the group will not be responsible for advising us on how to pace our internal progress on mathematics." OpenAI's own post returned an error when this run tried to fetch it, so that wording rests on TechCrunch's report.
The disclosure item is the one with teeth, because the material it asks for is already in OpenAI's hands. The dispute over the Navier-Stokes claim turned on exactly that kind of record: our report on the data dispute covers the mathematician who asked whether his unpublished Codex sessions fed the training run, and neither side has published the logs.
The group is also not a uniform bloc. TechCrunch reported that De Lellis was the only initial member to sign the open letter from Fields Medalists objecting to the pace of AI proof-solving. We covered the Fields Medalists' warning separately.
The group has not said when it expects OpenAI to release the first batch of results, or whether OpenAI has agreed to the September 29 norms. The next thing to watch is whether the first 2.B release arrives with a formalized proof and a prompt log attached.
Sources
More in A.I.
- 01Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 02OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 03Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use ItGoogle priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.
- 04Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.