Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
Back to Originals
Research · Mathematics

OpenAI Posted 722 AI-Written Math Papers to Its Own GitHub. Its Advisers Had Asked for a Repository No Lab Controls.

Adrian Vale··7 min read

On Tuesday, October 6, OpenAI pushed a single commit to a new public repository, openai/math. It holds 722 mathematics manuscripts, grouped into 372 result families, all produced by what the README calls "an unreleased internal OpenAI model." The commit is timestamped 14:58 Pacific, which is 21:58 UTC.

The headline claims in OpenAI's own catalog are enormous. A proof of Khot's Unique Games Conjecture. A zero-free half-plane at Re(s) > 7/8 for every Dirichlet L-function, which the catalog calls the quasi-Riemann hypothesis. An isomorphism between the free group factors L(F2) and L(F3). The rational Hodge conjecture for every complex CM abelian variety. Integer multiplication below n log n, with a savings exponent of 2-182.

We are not going to grade the proofs. Nobody can yet, and the advisory group closest to the release says that assessment belongs to the mathematical community. What we can grade is the release itself, because the yardstick was published a week earlier by the group OpenAI says it consulted.

The Yardstick

The Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, was announced on Monday, September 21, in a guest post on Terence Tao's blog. It has nine members, including Timothy Gowers, Martin Hairer and Edward Witten, is hosted at the Institute for Advanced Study in Princeton, and says its members "do not accept payment for this work." Its own site says the group formed after OpenAI approached some members about an external advisory board, and they chose to make it independent instead.

On Tuesday, September 29, AGMAI published guidelines for the responsible release of AI-generated mathematics, informed by more than 600 survey responses. It opened with a line aimed squarely at labs like OpenAI: "we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." Then, acknowledging that labs are doing it anyway, it listed what a responsible release should contain. Those items are concrete enough to check line by line against what landed on GitHub.

AGMAI asked forWhat openai/math containsOur grade
Name of the model"An unreleased internal OpenAI model"Missed
Prompts usedNone we could find in the repositoryMissed
Summarized chain of thoughtAbridged reasoning summaries for 10 of 372 familiesPartial
Time takenAn average: three hours of ChatGPT Pro thinking compute per resultPartial
Estimated cost of computationNo dollar figure in the READMEMissed
Formalization, a formalization.yaml, Comparator challenge filesformalization.yaml lists 162 papers with a formalized main result; 405 Comparator challenge configsMet, for about 22 percent of papers
How many problems were tried and failed, and how they were chosenAbout 4,000 problems posed; kept results meeting "an appropriate level of significance"Partial
Deposit in a repository no AI lab controlsOpenAI's own GitHub organization; "exploring community-hosted repositories"Missed, for now
Recorded revisions and citable versionsPromises preserved release history; BibTeX per manuscriptPartial

Our count: one clear pass, four partials, four misses. The grades are ours, not AGMAI's. AGMAI says it is up to the mathematical community to judge how well its recommendations were followed.

The Denominator Finally Showed Up

The best line in the README is a number we said was missing in August. When OpenAI published its ten Astra proofs on August 1, it told us what the winning runs cost and nothing about how many problems it tried. This time it says the model "was posed approximately 4,000 problems."

That is real progress, but it is not a success rate. A family can bundle a principal result with companions, consequences and alternative proofs, and OpenAI does not map families back to the 4,000 prompts. So 372 divided by 4,000, about 9 percent, tells you only the order of magnitude. AGMAI asked for something sharper: how many problems "of comparable difficulty" the model tried and failed. That number is still missing.

The compute disclosure has the same shape. "On average, each result used three hours of ChatGPT Pro thinking compute," the README says. That is a unit of product, not a unit of cost, and it comes with exceptions. The README says work on a zero-free region for the Riemann zeta function and the Hodge conjecture proof for CM abelian varieties did not follow the fixed procedure, and that the Re(s) > 11/12 writeup was human edited for readability.

The folder names tell their own story about pace. By our count, 564 of the 722 manuscript folders carry dates from September 23 to September 27. Another 112 are dated October 5, the day before release.

The Lean Files Are the Strongest Part

The formalization work is where OpenAI went furthest. The repository ships a Lean library, a formalization.yaml in the community schema, and Comparator challenge configurations, including ones named for the quasi-Riemann hypothesis and the Unique Games theorem. Anyone with the toolchain can run them.

Read the metadata closely, though. The yaml lists the formalization method as "agent," the review status as "unchecked," and the scope as "Partial progress." The README itself warns: "Some of the unformalized results could have issues." With 162 of 722 papers carrying a formalized main result, roughly 78 percent of the papers have no formalized main result in the catalog. And a Lean certificate proves the formal statement, which is only as good as the match between that statement and the conjecture humans care about.

A Community Repository Opened the Same Day

The miss that stings most is the venue. AGMAI asked that results be deposited in repositories "not controlled by any AI lab," with persistent identifiers and recorded modifications. On the same Tuesday, Ben Antieau announced Hexagon, a repository run by a new nonprofit foundation and built for the wave of LLM-assisted results, in a guest post on Tao's blog. Its board includes AGMAI member François Charles and Northwestern's Bryna Kra. Hexagon says it "welcomes submissions by large labs working on language models."

WIRED reported that OpenAI had been directly encouraged to use tools like Hexagon, and quoted Kra: "they haven't changed their behavior." OpenAI's statement, as quoted by The Verge, leaves the door open: "We're continuing to explore other community-hosted alternatives for this release which meet the committee's guidelines."

AGMAI's own October 6 statement was careful to the point of coolness. Its advisory role "should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them." And: "This release is the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge."

Caveats

  • Every mathematical claim here is OpenAI's description of its own results. We found no independent verification of the headline results as of this writing.
  • Our grades read the repository and public statements. OpenAI may have disclosed more to AGMAI privately; the group says it has discussed its recommendations with the company.
  • The 162 count comes from formalization.yaml. Other Lean files in the library may cover lemmas or partial results that the catalog does not list as a formalized main result.

Our Take

This is the most complete math release we have seen from any lab. 722 papers instead of a blog post, Lean files with Comparator challenges, a problem count, and a promise of versioned corrections. Measured against "math by tweet," it is a different category.

Measured against the checklist, the pattern in the misses is hard to ignore. OpenAI delivered the items that make the results look more credible: formal proofs, a big denominator, 722 typeset papers. It withheld the items that would tell a competitor or a customer something: the model's name, the prompts, the dollar cost. And it kept the papers on its own GitHub, with the OpenAI name on every path, on the same day a community repository was announced and invited lab submissions.

AGMAI asked labs to "refrain from treating the release of mathematical results as marketing vehicles to promote their models." An unnamed model whose effort is measured in hours of ChatGPT Pro thinking is a strange way to honor that. Our read is that the math is probably the most important thing OpenAI published this year, and the release is still a product teaser with a bibliography.

Three signposts for the next 60 days:

  • Venue. Whether OpenAI deposits the collection in Hexagon or another community-run repository with persistent identifiers. That is the cheapest miss to fix.
  • First external verdict. Whether an independent group publicly reports running the Comparator challenges for flagship results such as the quasi-Riemann, Unique Games and free group factor files, or finds a gap in one of the roughly 560 papers with no formalized main result.
  • The name. Whether the model behind the collection gets a name, a price and a release, and whether OpenAI then publishes the per-result cost AGMAI asked for.

For background, see our coverage of OpenAI's Navier-Stokes result and the May unit-distance disproof.

Sources: openai/math on GitHub, AGMAI: On OpenAI's Release of Mathematical Results, AGMAI: Responsible Release of AI-Generated Mathematics, Terence Tao: Announcing AGMAI, Ben Antieau: Hexagon, The Verge, WIRED and Unite.AI.