OpenAI Released 722 AI-Written Math Papers on GitHub: What's Inside, What's Verified and Why Mathematicians Are Split
On October 6, 2026, OpenAI published 722 mathematical manuscripts in 372 result families, produced by an unreleased internal model, to a public GitHub repository with Lean formalizations for many of the proofs. Here is what was released, what has actually been checked, and why the math community is divided.

What OpenAI released
On October 6, 2026, OpenAI published Sharing AI progress in mathematics, announcing a broad batch of new mathematical results produced by an internal frontier model that has not been released. The results live in a public GitHub repository, openai/math, rather than in academic journals.
- 722 manuscripts, organized into 372 result families. A family groups related papers: a principal result, companion arguments, consequences or alternative proofs.
- Lean formalizations for many, but not all, of the proofs. Lean is a programming language that lets a computer check a proof step by step. OpenAI says it will add more as it obtains them.
- Protocols for revisions and citations: corrections are recorded as new versions, and earlier versions stay accessible.
- Process details: 10 abridged summaries of the model's reasoning, compute estimates and statistics about attempted problems.
This page is the long-form briefing behind the BSH Technologies Instagram carousel.
The release in numbers
- 722 manuscripts in 372 families.
- About 4,000 problems were posed to the model over the course of the evaluation, according to the repository README.
- About three hours of ChatGPT Pro thinking compute was used per result on average.
- 10 reasoning summaries were published, covering families such as the irrationality exponent of π, the symmetric and general Mahler conjectures, NP-hardness at the basic semidefinite threshold, Kaplansky's direct-finiteness conjecture in characteristic two, and spontaneous magnetization in the quantum Heisenberg ferromagnet.
The README says OpenAI expanded these evaluations after its models saturated its existing mathematical benchmarks. Most results came from the same fixed procedure. The exceptions it lists are work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties; the write-up for the Re(s) > 11/12 zero-free region was human-edited for readability. The Decoder reports that nearly every result came from a single prompt to a single agent, in contrast with the Navier-Stokes result OpenAI announced in September.
A problem one mathematician gave up on
New Scientist spoke to Francis Johnson at University College London, who worked for many years on Wall's D(2) problem, one of the puzzles addressed in the release. "I worked on this problem for 25 years. I produced two books on it. I, personally, gave up," he said, adding that he was surprised AI did it so quickly, but not that it did it.
Checked vs. claimed
OpenAI's own README is careful: the collection includes results at different stages of verification, not all have Lean formalizations, and some of the unformalized results could have issues. Kevin Buzzard at Imperial College London told New Scientist the release included 30 papers relevant to his field of number theory; only 7 seemed impressive to him, and only 1 was formally verified in Lean. "Acceptance of these results by the community will take time," he said. A Lean proof shows the logic is correct, but it cannot judge whether a result is new or important.
Why mathematicians are uneasy
OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI) at the Institute for Advanced Study. AGMAI's recommendations ask labs to disclose the model name, prompts and compute costs, formalize proofs where possible, report how many comparable problems failed, and publish through independent, versioned repositories. As The Verge notes, the group also urged labs to stop treating math results as marketing vehicles.
- Model: unnamed and still unreleased; OpenAI calls it an internal frontier model.
- Prompts: not published, and compute is reported only as an average rather than per problem.
- Venue: a GitHub repository controlled by OpenAI, not peer-reviewed journals. OpenAI says it is exploring community-hosted alternatives.
- Review load: hundreds of papers arrive at once, and human checking does not scale with them.
OpenAI says it will fund workshops, conferences and special programs around understanding AI-produced results, and is working to responsibly release the model.
What this means for teams
- Research and R&D leads: treat AI-generated proofs like any unreviewed preprint until they are formally verified or checked by experts.
- Engineering teams: the pattern to copy is the pairing of generation with machine-checkable verification, whether that is Lean for math or tests and type checks for code.
- Decision makers: capability claims are moving faster than review capacity; budget for verification, not just generation.
Primary sources
- OpenAI: Sharing AI progress in mathematics (Oct 6, 2026)
- GitHub: openai/math repository and README
- AGMAI: Advisory Group on Mathematics and Artificial Intelligence
- The Verge: OpenAI drops another batch of mathematical breakthroughs (Oct 6, 2026)
- New Scientist: OpenAI announces 722 mathematical discoveries in one go (Oct 7, 2026)
- The Decoder: OpenAI dumps 372 AI-generated math proofs on GitHub (Oct 7, 2026)
How BSH can help
At BSH Technologies we help teams put frontier AI to work with verification built in: automated checks on generated code, evaluation pipelines and human review where it matters. If you want AI output you can trust in production, our Thrissur engineers can help you design the checking layer first.
Frequently asked questions
What did OpenAI release on October 6, 2026?
722 mathematical manuscripts organized into 372 result families, produced by an unreleased internal OpenAI model and published in the public openai/math GitHub repository with Lean formalizations for many of the proofs.
Are all of the results verified?
No. OpenAI says the collection is at different stages of verification, not all results have Lean formalizations, and some unformalized results could have issues. Mathematicians are still reviewing them.
How much compute did each result use?
On average, about three hours of ChatGPT Pro thinking compute per result. The model was posed roughly 4,000 problems over the evaluation.
From the blog
View all posts
OpenAI Is Watermarking ChatGPT Text in the EU: How textGrain Works and What It Can't Prove
On October 5, 2026, OpenAI said it will add an invisible text watermark called textGrain to eligible ChatGPT and Codex output in the European Union to meet the EU AI Act, let API customers worldwide opt in for select models, and open its detector to approved researchers only.

Google Pauses Its Open-Source Bug Bounty After a Flood of AI Bug Reports
On October 1, 2026, Google stopped accepting new product vulnerability reports to its Open Source Software Vulnerability Reward Program (OSS VRP), citing a significant rise in automated submissions, the vast majority of which are not valid. Supply chain reports and reports already filed are unaffected, and Google promises an update in Q1 2027.