Picture a mathematician on October 6, 2026, opening a GitHub page with a coffee going cold beside the keyboard. Not a press release. Not a journal. A code repository, the kind of place programmers keep half-finished side projects. Inside it: 722 mathematical manuscripts, dropped at once, credited to a model OpenAI has not named and will not release.
That is a lot of reading. It is also, depending on who you ask, either the beginning of a golden age or a problem wearing the costume of a breakthrough.
What actually happened
A month before the GitHub drop, on September 8, 2026, OpenAI announced that an internal model had produced a proof resolving the Navier-Stokes existence and smoothness problem. If that name means nothing to you, the shorthand is this: it is one of the seven Millennium Prize Problems, a list of questions so hard that solving one has long been treated as a career-defining event.
The method is the part I keep returning to. The company said the proof came out of roughly 10,000 concurrent AI agents working at the same time.
Then came October, and the 722 manuscripts. OpenAI said its internal model had resolved hundreds more mathematical problems, including a solution to the four-dimensional Kakeya conjecture and improvements on some of the world’s most important computer algorithms.
What 10,000 agents means in plain terms
Since this site exists to explain agents, let me sit with that number for a second.
A single AI agent is a model that has been given a goal and the ability to take steps toward it on its own, checking its own work as it goes. One agent working on a math problem is a bit like one researcher at a whiteboard. Ten thousand agents running at once is less like a researcher and more like a weather system. They try paths. Most of those paths go nowhere. A few survive.
This is a different shape of work than most people imagine when they picture AI doing mathematics. There is no single flash of insight to point at. There is search, at a scale no human department could staff, filtered down to whatever held up.
Two things follow from that, and they pull in opposite directions:
- Scale can genuinely find things humans missed, because humans only get to try so many approaches in a lifetime.
- Scale makes the output very hard to check, because nobody was watching 9,997 of those agents.
Why mathematicians pushed back
The reception was rough. Twenty-five Fields Medal winners argued that the push by AI companies to solve famous problems as a benchmark harms the science of mathematics and the mathematical community.
The Fields Medal is roughly mathematics’ equivalent of a Nobel. Getting 25 of them to agree on anything is itself notable. And their objection is not the one you might expect. They are not saying the machines are too dumb. They are saying the framing is wrong.
Famous unsolved problems were never meant to be scoreboards. They became famous because working on them produced useful side effects: new techniques, new language, new connections between fields that nobody expected to be related. The problem was the excuse. The learning was the point.
Treat those problems as benchmarks to be cleared, and you optimize for the wrong thing. You get an answer and skip the understanding. For a field where the whole value is in understanding, that is a real loss, not a sentimental one.
The verification problem nobody can skip
Mathematics has a slow, unglamorous quality control system. A result gets written up, submitted, read closely by people who would quite enjoy finding a hole in it, and only then accepted. That process takes months. Sometimes years.
722 manuscripts arriving simultaneously, from a model nobody outside the company can inspect, puts enormous strain on that system. The reviewing capacity of the entire field is finite. The generating capacity of 10,000 agents is, for practical purposes, not.
And because the model has not been released, there is no way for an outsider to reproduce the work, probe where it is weak, or understand how it reached its conclusions. You are asked to evaluate the output without access to the process.
What I’d watch for
I am not in the camp that says none of this matters. Some commentators see a new golden age for mathematics coming out of AI combined with human ingenuity, and that is plausible. Stephen Wolfram has been writing about what pure math research even looks like from here, which tells you serious people consider the question open rather than settled.
But the useful signal will not be the announcement. It will be what happens over the following months, as mathematicians work through those 722 documents and report what holds. Results that survive scrutiny get absorbed into the field. Results that do not, quietly disappear.
For non-technical readers, the takeaway is smaller and more portable than the headlines suggest. A claim produced by a system you cannot examine is still just a claim. Scale is impressive. Verification is what makes something true.
🕒 Published: