A company can announce that its AI has solved a difficult problem long before the rest of the world can judge what it has gained. The missing ingredient might be a proof, an explanation, access to the system or simply the time to examine the work. Each absence gives the company a different kind of power.
That is the useful question behind the latest discussion of hidden AI breakthroughs. In a September 15 essay, computer scientist Scott Aaronson reported rumors that AI companies were holding back solutions to major theoretical-computer-science problems after the response to the Navier–Stokes announcement. He did not publish those solutions. His account establishes that he has heard the claim, not that the unnamed results or the reported motive have been verified. Aaronson’s essay
There is a more concrete development. The independent Advisory Group on Mathematics and Artificial Intelligence says its current task is advising OpenAI on releasing significant mathematical results the company reports having produced internally. That confirms a release-coordination process around company-reported work. It does not certify every result awaiting release. The advisory group’s statement
For readers, the distinction prevents an exciting possibility from becoming a fictional inventory of solved problems. For science, it raises a larger issue: if discovery becomes concentrated inside private systems, how much independence will everyone else retain?
The big change
- What changed: Published AI-generated proof artifacts now sit alongside company-reported results awaiting coordinated release. Claims about additional hidden breakthroughs remain unverified.
- Why it matters: Laboratories can influence which discoveries become visible, when specialists can examine them and who gets to build on the methods.
- What to watch: Whether release arrangements give outsiders usable evidence and freedom to challenge it, while supporting the human expertise needed to choose research questions, understand results and share their benefits.
A breakthrough can be public while the capability remains private
OpenAI’s September 8 Navier–Stokes announcement offers an example of selective openness. The company released a written argument and formal proof files while describing the discovering system as an internal model more capable than GPT-6 Astra. Publishing an output and providing access to the system that produced it are different decisions. OpenAI’s announcement
The mathematical scope also needs care. The announced construction uses smooth external forcing. The official Clay problem statement explicitly permits forcing in its breakdown alternatives C and D; the unforced global-smoothness alternatives ask a different question. A valid forced construction can address the stated prize problem while leaving the unforced question unresolved. The official problem formulation
On September 11, Clay welcomed the apparent resolution while preserving an unhurried process for evaluation and credit. Its statement was an institutional response to an announcement, not a prize award. Clay’s response
We examined the wider mathematical sequence in our guide from unit distances to Navier–Stokes. Here, the significant distinction is who can do what after a release. An outside mathematician might be able to study the proof but lack the resources or model access to investigate a nearby problem in the same way. The discovery is shared; the means of producing the next discovery may remain concentrated.
That arrangement can still be valuable. Requiring every researcher to recreate the original search would make many results harder to use. What matters for checking a particular proof is access to the relevant evidence, not necessarily possession of every tool used to find it. But the ability to verify a published answer does not give outsiders equal influence over which questions get answered next.
The public needs a way to disagree
OpenAI’s public repository identifies its Navier–Stokes and Euler claims separately and supplies build instructions and a route to additional proof checking with Comparator. Those are inspectable artifacts rather than a demonstration video alone. We have read the documentation; we have not rebuilt or independently certified the proofs. The proof repository
Lean’s validation guide explains why a successful check has a defined scope. Reviewers must inspect the theorem’s meaning and assumptions, including dependencies. Stronger checks can compare a submitted proof with a separately specified statement and use independent checking software. That reduces reliance on the producer, while leaving stated assumptions about the tools and the specification. Lean’s validation guide
The hopeful possibility is considerable: a private system could perform an expensive search, then release evidence others can check without reproducing that entire search. AI-generated discoveries could become shared resources even when only a few institutions can afford to generate them at that scale.
The weaker version is a laboratory offering selected experts a private demonstration and asking everyone else to trust the endorsement. That may be a reasonable preliminary step. It gives the public much less ability to locate a disputed assumption, inspect a correction or pursue an unexpected application. Access to a conclusion is different from access to the grounds for disagreement.
Consider an original hypothetical. A university group is studying a class of optimization problems. A company privately finds a method that improves one important case. It could release the method and a precise guarantee, or provide an online service that returns answers while keeping the method private.
Both options might help users. Only the first necessarily gives the university something it can inspect, teach and adapt without continuing access to the service. If the company changes its terms, withdraws the model or stops supporting that field, those differences become practical. A useful service can coexist with a scientific dependency.
This is why public benefit cannot be measured by a discovery count alone. A release can expand what people know, what they can do, or what they must ask permission to do. Those outcomes sometimes move together. They need not.
A pause can improve science, or give one institution a longer lead
There are defensible reasons to coordinate a major release. Specialists may need time to check the claim, identify earlier work and prepare an explanation. A correction before publication can spare many people from building on a mistake. A credit dispute also deserves evidence rather than a race between announcements.
But the same delay can have a different effect on outsiders. Imagine a research group choosing its next project without knowing that a company claims to have settled the central conjecture. It could spend months pursuing the old question. Giving only selected collaborators advance access could instead give them an advantage in developing follow-up results. These are possible consequences of unequal information, not allegations about a particular unreleased theorem.
The relevant question is what a delay accomplishes and who is accountable for it. A review period with an identified claim, independent readers and a plan for communicating the outcome differs from indefinite secrecy accompanied by hints of extraordinary capability. Institutions need room to check work; audiences need a reason to distinguish that checking from promotion.
The new advisory group offers one channel. It says members are unpaid, recommendations will be public and companies retain decision-making authority. Its independence is useful, but it does not turn advice into an obligation to release a result or give researchers access. Independence and accountability
Our view is that a productive arrangement should make the response to advice visible. A company could explain which recommendations it accepted, what remains under examination and how reviewers can raise unresolved concerns. That would let outsiders assess the process without requiring a premature announcement of a theorem as settled.
The strongest criticism accepts that AI can succeed
The September 11 declaration A Severe Misalignment of AI in Mathematics explicitly acknowledges substantial improvements in mathematical capabilities. Its objection concerns incentives: using famous problems as benchmarks can neglect the development of ideas, attribution and the training through which people learn to formulate new questions. It also recognizes AI’s potential to improve mathematical study. The declaration
That is a more demanding argument than saying a machine cannot discover anything. It asks whether a successful system is optimizing for the same result as the community receiving its output. A laboratory might reasonably want an unambiguous demonstration of capability. A researcher may care more about a method that opens several new directions, including directions too obscure to attract a launch announcement.
There is no automatic reason for those interests to coincide. Nor must they conflict. A laboratory can fund the exploration and explanation that make its results useful. Universities can recognize work that adapts a powerful new method to a field the laboratory did not prioritize. Publication venues can make room for clarification and synthesis alongside first discoveries.
Criticism should still be answerable to evidence. A correct and consequential result should change an assessment of what a system can do. Discomfort about its origin cannot invalidate it. Equally, an impressive result does not answer a question about access or credit merely by being impressive. Keeping those judgments separate makes it possible to welcome the mathematics and challenge the terms under which it reaches the world.
Aaronson’s essay captures this tension: he regards recent progress as transformative while endorsing the effort to preserve human understanding. His description of a singularity beginning is an interpretation of events, not a demonstrated scientific threshold. His stated position
Faster discovery and meaningful human work are both real interests
The broader argument predates September’s announcements. In a December 2025 essay, AI researcher Julian Togelius distinguished tools that assist scientists from automation that makes their participation redundant. He defended human agency and the opportunity to contribute, even against a hypothetical in which full automation accelerated medical progress. This is a moral position about a possible future, not evidence that scientific work has already become redundant. Togelius’s essay
There is a serious cost to defending an activity solely because its current practitioners enjoy it. If a better method can reduce suffering, the interests of people who benefit deserve weight alongside the interests of people whose profession changes. A commitment to meaningful work should not become an automatic veto over useful discoveries.
But the choice need not be between protecting every existing task and removing people from science. Who selects the questions, identifies a neglected need, judges whether evidence applies and decides what to do with it? Those decisions affect the public even if machines eventually become better at many of their technical components. Retaining influence over them is a matter of governance as well as professional identity.
Dario Amodei’s October 2024 essay offers an expansive optimistic vision of AI accelerating biology and improving lives. It also identifies limits involving experiments, data and the physical world. The essay is a forecast built on assumptions about powerful future systems, not a timetable established by a mathematical breakthrough. Machines of Loving Grace
A proof about equations does not itself demonstrate a working treatment or a cheaper energy system. Each application needs further evidence. Nevertheless, it is reasonable to value the possibility of faster progress: assistance that helps researchers reject an unproductive direction or design a more informative experiment could matter well before research becomes autonomous.
The distribution question then returns. An advance can be scientifically sound and still reach only those who can afford it. Research can become faster while concentrating on the problems its owners find most valuable. Whether acceleration benefits a broad public depends partly on choices about funding, access and deployment. The number of discoveries does not settle those choices for us.
A proof of correctness is not a promise of good judgment
Alignment adds another question. A system might be excellent at finding an argument while being unreliable about following its operator’s intentions across a long research project. A checked mathematical output would establish something about that output. It would not, by itself, establish that the system consistently respects permissions, reports failures candidly or pursues the right goals.
The distinction matters for how laboratories use increasingly capable tools. Delegating a bounded search and examining its result is different from giving a system broad authority over experiments, spending or publication. The scientific achievement can justify trying more ambitious research while the permissions still require separate justification.
Preserving human understanding is helpful here, but it is not a universal solution to alignment. Knowing the mathematics behind a discovery does not reveal every aspect of the model’s behaviour. Conversely, a person need not reconstruct every intermediate calculation to exercise useful oversight. The practical aim is to preserve enough independent competence and evidence to notice when a system’s work no longer serves the intended purpose.
The question of meaning also deserves more than a promise that leisure will compensate for every lost role. Amodei argues that relationships and activities can remain meaningful even when AI outperforms people, while treating the economic arrangements of such a future as unsettled. That is one plausible outlook, not a guarantee of a painless transition. Work and meaning
A student can value understanding a result without being first to discover it. A community can value teaching, debate and inquiry without requiring its members to outperform every machine. Maintaining the institutions that make those activities possible, however, requires resources and choices. Personal fulfilment cannot pay for a laboratory, guarantee access to a model or preserve a department on its own.
Judge the release by the independence it creates
Can someone challenge the claim without depending entirely on the company’s account? Can another group use the result without reproducing an unaffordable search? Can researchers choose follow-up questions the producer has little interest in? Are there people and institutions able to explain the work, correct it and keep it available after attention moves elsewhere?
Those questions connect the optimistic and pessimistic futures. In one, powerful private tools contribute to a larger shared body of knowledge, and more people can investigate questions that were previously out of reach. In the other, outsiders receive impressive answers while becoming increasingly dependent on a few organizations to decide what is worth asking.
We do not yet know what the rumored hidden results contain. We can judge the arrangements being built around the results that do emerge. A major AI breakthrough should leave the rest of science with greater ability to act on its own.



