In short
- OpenAI revealed 722 math manuscripts in 372 outcome households on GitHub from an unreleased inside mannequin.
- Solely 162 of the 722 papers have a Lean-formalized primary outcome, and OpenAI warns that some unformalized outcomes may have points.
- MIT’s Andrew Sutherland says the one-prompt, single-agent declare is unverified till the mannequin is launched, and the discharge omits the prompts that an advisory group on the Institute for Superior Research really helpful disclosing.
OpenAI revealed 722 math manuscripts on GitHub on Tuesday, all produced by an inside mannequin the corporate has not launched. An OpenAI spokesperson said virtually all the pieces got here from a single immediate handed to a single AI agent, although some might have taken a number of makes an attempt.
It is a daring declare and a probably important breakthrough within the area of arithmetic. However not everyone seems to be a fan, or shopping for the hype.
“Till and until they launch the mannequin and other people can replicate their outcomes, I feel you must deal with any claims about one-shotting issues with a single agent as unverified,” Andrew Sutherland, a mathematician at MIT, informed Scientific American. “We should always ask for receipts,” he mentioned.
The papers are grouped into 372 “households” of associated outcomes, and a household can bundle a primary theorem with companion arguments, penalties or different proofs. That makes 722 a depend of manuscripts, not of solved issues. OpenAI says it posed roughly 4,000 issues to the mannequin and saved the outputs it judged important sufficient to publish.
The typical outcome used the equal of roughly three hours of ChatGPT Professional considering compute, per OpenAI. The Navier-Stokes declare final month appeared very completely different, with 10,000 coordinating agents working for 88 hours.
OpenAI launched abridged reasoning summaries for 10 of the outcomes. That mentioned, solely 162 of the 722 papers include a computer-checked primary outcome, in line with a formalization catalog within the repository. That’s about 22% of the gathering, translated into Lean, software program that checks each logical step mechanically.
OpenAI itself says not all manuscripts have Lean formalizations and that “a number of the unformalized outcomes may have points.” In different phrases, lots of what they revealed might be improper.
A passing Lean test confirms solely that the proof follows from the assertion as written in Lean. It does not present that the assertion matches the unique drawback, or that the result’s new or vital, which is the half mathematicians now have to guage.
And that is the place researchers increase their eyebrows.
I used to be making an attempt to learn the OpenAI proof that chromatic variety of a aircraft is >= 6. However it’s completely unbelievable alien math?
One way or the other the mannequin discovered that any Okay-coloring <=> “weakly measurable” Okay-coloring, which appears out of nowhere
1/2 pic.twitter.com/W13ScT3nel
— Dmitry Rybin (@DmitryRybin1) October 7, 2026
The openai/math repo has Points turned off and has by no means accepted a pull request. That is disappointing. Should you publish 722 manuscripts and ask for Lean formalisations, you want someplace for folks to ship them.
I am formalising OpenAI’s Saxl’s Conjecture proof in Lean 4 in opposition to…
— Keith Adler (@keithadler) October 7, 2026
“It’s now the case that AI can output mathematical arguments in conditions with out the human who prompted it with the ability to perceive the arguments, confirm them, or take duty for them,” The Institute for Superior Research in Princeton, New Jersey, said in a statement. “We imagine that human understanding of arithmetic stays of paramount significance. How, on this new period, can we work in the direction of a brand new paradigm that features human understanding of arithmetic as a part of accountable scholarly output?”
Others, although, like Professor Abhishek Saha, are fairly excited. “It’s a very massive day for arithmetic,” he wrote, however famous that a lot of the issues match within the classes of “distinctive advances inside an present program” of “shocking breakthroughs.”
This implies a lot of the issues within the set are attention-grabbing, however not unattainable or recreation altering just like the millennium issues. That spot is reserved for precisely one drawback out of the 722: the Quasi-Riemann Speculation.
Some additional ideas on the 372 outcomes launched by OpenAI right now, throughout 722 manuscripts.
If I have been to categorise theorems that mathematicians show and publish in line with their groundbreaking nature, I might (very roughly) divide them into 4 classes:
A) Non-breakthrough… https://t.co/BA10SxBlx7
— Abhishek Saha (@ObhishekSaha) October 7, 2026
The discharge additionally falls in need of what an advisory group on the Institute for Superior Research really helpful on September 29: the mannequin title, the prompts, a summarized chain of thought, the time taken and the compute price for each outcome. OpenAI revealed common compute figures and 10 reasoning summaries however no prompts, and says it’s nonetheless working to launch the mannequin responsibly.
Daniel Litt, a mathematician on the College of Toronto, took the other view, arguing there is no such thing as a cause to ask the corporate to maintain the solutions to those math questions secret.
Anthropic took a special route with its Lean-checked Fermat’s Last Theorem proof final month, posting all 13 million strains publicly on GitHub. That proof formalized a theorem Andrew Wiles revealed in 1995, fairly than claiming new outcomes.
OpenAI says it can add Lean formalizations because it obtains them; for now, 162 of the 722 manuscripts have one.
Each day Debrief E-newsletter
Begin daily with the highest information tales proper now, plus unique options, a podcast, movies and extra.
You might also like
More from Web3
Hackers Used AI Agents to Raid a Megachurch’s Database, Exposing 850,000 Members
Briefly Yoido Full Gospel Church in Seoul stated Wednesday that knowledge tied to 850,000 members might have been compromised, together …
Europol Warns Crypto Wallets Are ‘Primary Risk’ for Quantum Attacks
In short Europol, the European Union's legislation enforcement company, revealed two experiences Wednesday urging the crypto business and policymakers to …
Google Launches Nano Banana 2.1: Better Than Its Predecessor at Half the Price
In short Google launched Nano Banana 2.1 on Oct. 6, rolling it out within the Gemini app, Search's AI Mode, …





