Nine Mathematicians Advising OpenAI Ask AI Labs to Stop Testing Hard Problems on Models Nobody Else Can Use
A.I. / news
Nine Mathematicians Advising OpenAI Ask AI Labs to Stop Testing Hard Problems on Models Nobody Else Can Use
The Advisory Group on Mathematics and AI published its rules on Sept. 29: release fast, fund human understanding, disclose prompts and costs.

The Advisory Group on Mathematics and Artificial Intelligence, a nine-member panel hosted by the Institute for Advanced Study, published recommendations on Sept. 29 for how AI labs should release mathematical results. The group opens by saying it does not endorse what some labs are doing now. "We do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models," the statement says, referring to "frontier AI labs" testing problems on models "that remain inaccessible to the broader scientific community."
The group says it was informed by more than 600 survey responses. Its website says it formed after OpenAI approached some of its members about an external advisory board, and that members "do not accept payment for this work."

Who is on the panel
The members listed on agmai.org are François Charles of ENS-PSL, Camillo De Lellis of the Institute for Advanced Study and GSSI, Timothy Gowers of the Collège de France and Cambridge, Martin Hairer of EPFL and Imperial College London, Nikhil Srivastava of Berkeley and the Simons Institute, Ulrike Tillmann of Oxford, Ravi Vakil of Stanford, Edward Witten of the Institute for Advanced Study, and Melanie Matchett Wood of Harvard.
ENGtechnica attributes to Hairer the statement that the group "receives no funding or technical support from OpenAI" and operates independently.
What OpenAI has claimed
The survey, the statement says, asked about "a specific situation in which OpenAI announced the existence of many results without giving details." ENGtechnica quotes OpenAI as saying that "an unreleased model has resolved more than 100 longstanding mathematical problems."
The group's current job is advising OpenAI on coordinating the release of "a large number of significant results in mathematics" from that internal model, according to agmai.org.
The rules, in the group's own order
The statement rests on three principles. Labs should release significant results promptly. Labs that release large volumes of output without matching human understanding must fund efforts to build it. That work must stay "organic and community led" and not be directed by the labs.
| Area | What the statement asks |
|---|---|
| Understood results | Follow academic norms: preprints, peer review, talks |
| Literature | Scour it for ideas related to the proofs |
| Deposit | Independent repositories with persistent identifiers |
| Disclosure | Model name, prompts, reasoning chains, compute time, costs |
| Verification | Formalise proofs where feasible |
| Failures | Report failed attempts on comparable problems |
| Funding | Conferences, workshops, postdocs, expository writing |
The statement also warns against "a two-tier system" in which research depends on proprietary models, and asks for broad access to publicly available models.
Why this follows the September dispute
OpenAI's earlier math claim drew a public fight. On Sept. 8, a New York University mathematician accused the company of unfair conduct over a Navier-Stokes result, as covered in our report on OpenAI's Navier-Stokes data dispute. ENGtechnica reports that researchers worry about "a sudden flood of claims that would demand effort to verify and understand," and that past OpenAI manuscripts were criticised as "poorly written, thin on relevant scholarship, or unclear about earlier contributions."
The same access gap runs through Google's Gemini 4 Argon release, which only a narrow set of users can reach.
What is not known
The statement has no enforcement mechanism, and no lab has said it will adopt the disclosure list. OpenAI has not said, in the material fetched for this report, when its 100-plus results will be released.
Sources
More in A.I.
- 01GPT-Synopsys: OpenAI Gets Paid Only When the Chips It Helps Design Beat the Customer's BaselineThe Sept. 30 deal pairs an OpenAI model with Synopsys' design software, with no price, no release date and no named customer.
- 02Ataraxos Beats Stratego's Top Player 15-1-4 After Training on 16 H100s for a WeekA Nature paper from MIT, Carnegie Mellon, NYU and Stanford puts the compute bill under $8,000, against an estimated $3 million to $4.5 million for DeepMind's DeepNash.
- 03Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 04OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.