Article

OpenAI's 100-Plus Math Claim Needs a Review System as Scalable as the Search

OpenAI says an internal model resolved over 100 open math problems. The hard part now is scaling proof review, attribution, disclosure, and control.

A large field of mathematical proof paths converging on separate review, attribution, disclosure, and human-control gates

OpenAI now says an unnamed internal model has resolved more than 100 long-standing open problems across most areas of mathematics. The company announced that count alongside a new independent advisory group and a separate proposal for international standards around automated AI research. The obvious reading is that OpenAI has built an extraordinarily prolific mathematician. The more consequential reading is that mathematical search may be scaling faster than the institutions that can review, explain, attribute, and safely release its output.

The public evidence surface has not grown in proportion to the headline. OpenAI's announcement provides no batch catalog, item-level proofs, review ledger, model identifier, API endpoint, price, or release date. That absence does not show the claim is false. It means the public currently has an aggregate company statement where earlier releases offered inspectable artifacts. A counter can rise instantly; disciplinary confidence cannot.

This is a direct update to RohitAI's September 8 analysis, OpenAI's Navier–Stokes Claim Is a Proof-Factory Stress Test. Two facts in that article's snapshot have since changed. Clay now labels the problem Active and described the result as apparently settled while keeping evaluation and credit assignment deliberately unhurried. OpenAI also added a September 10 company finding that its investigation excluded influence from Tristan Buckmaster's recent Codex prompts, including through training. The first updates institutional reception; the second narrows a provenance dispute according to OpenAI. Neither is a prize decision or an external audit of every claim.

RohitAI's read: the new bottleneck is an operating model. Correctness, novelty, attribution, intelligibility, disclosure readiness, and permission to continue automated research are different decisions. If they share one label called verified, the label will fail exactly when the volume becomes important.

When discovery runs at fleet scale, review cannot remain an artisanal queue with one green check mark at the end.

The evidence changed shape three times in seven weeks

The cleanest way to understand today's announcement is to compare disclosure formats, not multiply the number of claimed results. OpenAI's August package, the Navier–Stokes campaign, and the September 21 aggregate each expose a different amount of evidence.

Release

Public evidence

What it established

What remains open

August 1: ten advances

Named results, paper, reasoning walkthroughs, and Lean certificates

A small, inspectable set of resolved or advanced problems

External review depth and generalization beyond the selected set

September 8: Navier–Stokes

Paper, theorem scope, repository, formalization, and system-scale details

A concrete forced-flow blowup construction with unusually rich artifacts

Community acceptance, credit, and Clay's formal prize process

September 21: 100-plus claim

Aggregate count and disclosure advisory arrangement

A claim awaiting item-level inspection

Inventory, counting rule, proof states, novelty checks, and model details

That progression creates an inversion: the larger the claim becomes, the less item-level evidence is available at announcement time. This can be a sensible staged-release choice. Reviewing a hundred serious results takes time, and publishing them carelessly would create priority, attribution, and benchmark-contamination problems. But staged disclosure needs visible states. Otherwise readers are asked to treat a work queue as a completed audit.

The Navier–Stokes package also warns against transferring system details from one campaign to the whole batch. OpenAI reported roughly 10,000 concurrent agents, 88 hours to discovery, another 17 hours of Astra formalization, 2.7 million messages, and about 130 billion output tokens for that effort. Treat those as campaign figures, not a recipe or budget for the wider batch. There is no defensible cost per theorem, retail-price conversion, or tenfold capability claim to calculate from the available material.

What the result counter leaves out

“Open problem” is not a standard unit. One item can be a narrowly scoped construction with a fast computational verifier; another can require a long argument, new concepts, and years of absorption by a field. Epoch's FrontierMath: Open Problems methodology explicitly separates significance tiers and notes that some verifiers provide strong numerical evidence rather than a general proof. That selected corpus is not a denominator for OpenAI's claim, and OpenAI has not supplied one.

A useful result counter therefore needs versioning. Epoch's September 16 tracker changes added human-plus-AI attribution and revised earlier labels; the tracker also records removals. A historical curve can move because the taxonomy changed, not because a model improved. The same problem will appear inside labs unless every headline count can be decomposed.

  • New discoveries under a fixed rubric: results first accepted during the measurement window.

  • Reclassifications: old work whose AI or human contribution was reassessed.

  • Corpus changes: problems added, removed, merged, split, or found to have prior solutions.

  • Review-state changes: claims that moved from generated to formally checked, expert-reviewed, disputed, or accepted.

Capability counters are only as stable as their attribution rules. A lab can apparently gain or lose “solved problems” without changing its model. Builders should publish the taxonomy version, attempted-problem denominator, human contribution, and revision history next to the count. Without those fields, the number is marketing telemetry, not capability telemetry.

Publishing a solution also retires an evaluation asset. Once a result becomes searchable, reproducing it no longer tests the same frontier capability. Labs need dated exposure ledgers and replacement tasks; mathematical communities need enough transparency to establish priority and credit. This creates a real tension, but secrecy is not the solution by default. The answer is a controlled disclosure pipeline that records when a problem stopped being clean.

Six reviews can reach six different conclusions

A formal proof artifact can settle one question exceptionally well while leaving several others untouched. The Lean validation guide distinguishes a kernel-accepted proof from the meaning of the theorem statement and recommends stronger procedures for unreviewed AI-generated code, including axiom inspection, fresh rechecking, sandboxing, trusted challenge statements, and external checkers. A green build is evidence. It is not a certificate of novelty, attribution, explanation, or authorized conduct.

Review layer

Question

Useful evidence

Natural owner

Problem identity

What exact version was attempted?

Versioned statement, scope, assumptions, exposure date

Domain editor or benchmark curator

Formal validity

Does the term prove the encoded statement?

Axiom report, reproducible build, fresh and external checks

Formal-methods reviewer

Semantic correspondence

Does the encoding mean the intended theorem?

Trusted statement comparison and definition audit

Domain expert plus formalizer

Novelty and priority

Was it already known, and whose ideas were used?

Prior-art search, retrieval log, provenance, dispute record

Independent experts and editors

Understanding and reuse

Can other mathematicians explain and extend it?

Readable proof, seminars, derived lemmas, successful reuse

The research community

Recognition and release

Is it ready to announce, publish, award, or feed into training?

Named decision, evidence packet, objections, conditions

Institution, publisher, and accountable lab leader

OpenAI's Navier–Stokes repository metadata declares zero sorry placeholders, standard Lean axioms, and a full main-result formalization. That is meaningful technical evidence, although its review label is self-assessed and RohitAI has not independently rebuilt the proof or audited every definition. Clay's process operates on another clock: its official rules require a qualifying publication, at least two years, and general acceptance before prize consideration. The two-year rule is an institutional recognition condition, not a theorem that a proof becomes logically correct after 730 days.

Scientific trust and workflow authorization require separate approvals. One asks whether the result should be trusted. The other asks whether the workflow that produced it was authorized: retrieval, data use, compute spend, experiment scaling, human intervention, and publication. A correct theorem can emerge from an unauthorized process. Conversely, a perfectly authorized run can still produce bad mathematics. Combining those decisions invites both scientific and operational failure.

AGMAI can influence disclosure without controlling the search

The new Advisory Group on Mathematics and Artificial Intelligence says it takes no payment from companies, will publish its recommendations, may advise any AI company, and has no corporate decision-making authority. Its immediate task is coordinating disclosure of results reported by OpenAI. Those are useful independence properties. They do not make the group an auditor, regulator, release veto, or operator of OpenAI's internal research system.

OpenAI explicitly excludes advice on the pace of its internal mathematical progress from the group’s responsibilities. This distinction matters more than the prestige of the roster. Advisory independence and operational power are different axes. A group can publicly criticize disclosure while having no ability to stop another search run tomorrow.

Independence answers “whose judgment?” Authority answers “who can make the system stop?” A credible operating model must answer both.

The mathematical debate is also richer than “experts versus AI.” A community declaration criticized trophy-style problem solving that neglects understanding, attribution, and exposition. Martin Hairer signed it and also joined AGMAI. Timothy Gowers, another group member, argued in a guest essay hosted on Terence Tao's blog that the forced Navier–Stokes result should be accepted while still emphasizing comprehension. That is not Tao's endorsement or a line-by-line independent audit. It is evidence that correctness and mathematical value can be discussed separately without dismissing either.

The test for AGMAI is therefore not whether every recommendation is flattering. It is whether recommendations identify the evidence reviewed, publish unresolved objections, distinguish claim states, and remain visible when OpenAI disagrees. Independence becomes operational through access and publication rights, not adjectives.

The standards proposal still needs operating rules

OpenAI's companion standards proposal calls for shared measurement of research autonomy, immediate human-review triggers for consequential automated research processes, and common incident classification and reporting. It says the standards would not themselves require licenses or mandatory prerelease approval; governments would decide whether and how to put them into law. The proposal specifies no numerical RSI trigger, common severity rubric, or reporting deadline. Its coordination agenda should not be mistaken for an adopted operating standard.

That makes the document an agenda for technical coordination, not a new binding regime. There is precedent for the slower machinery: a ten-government network convened through NIST/CAISI published preliminary consensus and open questions on automated evaluation practices earlier this year. Preliminary consensus is useful groundwork. It is not an adopted standard for recursive self-improvement.

Layer

Status today

Power

Important limit

AGMAI

Independent advisory arrangement

Review and public recommendations about mathematical disclosure

No company decision rights; internal pacing excluded

OpenAI Preparedness Framework

Existing voluntary company policy

Defines internal capability thresholds, evaluations, and leadership responsibilities

Not international law and not run by AGMAI

September 21 standards agenda

Company proposal

Offers categories for measurement, human review, and incident reporting

No adopted thresholds, deadlines, or enforcement

Government incorporation

Future national choices

Could create binding obligations within jurisdictions

Not established by today's announcement

OpenAI also says fully autonomous RSI is not happening today. That caveat should stay attached to every dramatic interpretation of the math count. The announcement is a reason to build measurement and escalation systems before they are urgently needed. It is not evidence that an AI has autonomously designed, trained, evaluated, and authorized its successor.

Following a result into the next model generation

A system can be excellent at mathematics, run thousands of agents, and increase research output without closing the recursive loop. The critical question is whether validated AI work changes the next model generation, shortens the critical path, and repeats with declining human control.

human chooses objective
  -> agents propose experiments
  -> tools execute under policy
  -> independent checks validate results
  -> authorized humans choose successor changes
  -> training and evaluations run
  -> pause, deploy, or repeat

OpenAI's September 6 internal research telemetry reported 3.1 agent-workdays for every human workday while stating that people still set priorities, choose which results to pursue, and decide whether to scale, pause, or deploy. It also reported that more than half of successful four-to-eight-hour tasks required intervention. Runtime is not wall-clock acceleration, and more experiments can reflect more compute as well as better research.

Anthropic's separate R&D Automation Index said Claude led 26% of its weighted August task basket under supervision and measured no fully autonomous category. As RohitAI argued when that index launched, a task share is not a speedometer. Neither company's metric can be mechanically compared with OpenAI's hundred-plus math count.

An RSI dashboard needs four unlike quantities. Track activity volume, validated throughput, critical-path time, and control capacity separately. Activity asks how much agents did. Throughput asks what survived review. Critical-path time asks whether a model generation arrived sooner. Control capacity asks whether monitors and humans could still understand, interrupt, and reverse consequential steps. Compressing them into one scalar loses distinctions that matter for both capability and safety.

The practical operating model for builders

For builders, the actionable lesson is infrastructure. Teams building research agents can prepare now by making every accepted result a versioned object rather than a persuasive transcript.

result_id: stable-and-versioned
problem_statement: exact_hash_and_scope
exposure_cutoff: timestamped_search_snapshot
model_and_harness: immutable_versions
human_interventions: complete_event_log
compute_and_tools: budget_plus_receipts
proof_artifacts: hashes_and_checker_versions
review_state: generated|formal|expert|accepted|disputed
authorization_state: allowed|paused|released
objections: preserved_not_overwritten
  1. Split the queues. Discovery, formal checking, semantic review, prior-art search, explanation, and external release need named owners and separate service levels. A result should not jump from generated to public because one checker returned green.

  2. Make the statement trusted before the proof is checked. For high-stakes formal work, compare generated code against a reviewer-controlled challenge statement, inspect axioms, rebuild in isolation, and preserve checker versions and receipts. Treat imported code as executable risk, not inert prose.

  3. Keep provenance outside the generator's control. Record retrieval snapshots, source exposure, human prompts, copied lemmas, and contribution claims in an append-only ledger. The agent that benefits from a novelty claim should not be the only system certifying it.

  4. Measure validated output and elapsed time. Count failed branches, review labor, interventions, and time to reusable acceptance. A larger fleet can inflate activity while leaving the actual bottleneck untouched.

  5. Separate scientific acceptance from operational authorization. Require explicit approval for spend escalation, changing an evaluation, feeding results into training, widening tool access, and public release—even when the mathematics checks out.

  6. Pre-commit pause and restart conditions. Name who can stop the run, what evidence triggers review, which actions are frozen, how logs are retained, and who can authorize restart. A pause without scope and restart criteria is a press statement, not a control.

This design also improves ordinary product work. A code agent that discovers an optimization, a biology agent that proposes a candidate, or a security agent that finds a vulnerability faces the same split: validate the artifact, verify provenance, authorize the workflow, then decide disclosure. Mathematics makes that distinction unusually clear.

What would make the 100-plus claim substantially stronger

The next useful milestone is not 200. It is a manifest. A staged public inventory could preserve responsible disclosure while showing problem identity, significance, artifact availability, human contribution, review state, known objections, and revision history. Even a partial release would let outsiders distinguish a large queue from a fully reviewed body of work.

  • A batch catalog: stable identifiers, exact statements, fields, and a clear reconciliation with Navier–Stokes.

  • A claim-state matrix: generated, formally checked, independently reviewed, disputed, revised, and broadly accepted.

  • Reproducible receipts: artifact hashes, toolchains, axioms, checker outputs, and isolated rebuild instructions.

  • A provenance record: search cutoffs, retrieval sources, model and harness versions, compute, and human interventions.

  • An authority map: who recommends, who decides, who can pause, who can release, and how dissent is published.

  • RSI-relevant telemetry: which validated results changed successor development, how much critical-path time fell, and whether oversight kept pace.

My forecast is that organized disclosure will arrive in tranches, because release coordination is AGMAI's stated immediate work. Provenance and review-state conventions will probably converge before governments agree on binding research-pacing thresholds; documentation is easier to standardize than authority. Public access to the unnamed model may also lag the first proof releases. These are predictions, not announced schedules.

There is also a product opportunity hiding under the governance argument. The durable tools may be those that turn generated proofs into reusable mathematics: extracting definitions, comparing arguments, tracking dependencies, preserving credit, and helping domain experts review only the uncertain layers. The success metric should be whether another researcher can correctly reuse an idea—not how many PDFs a fleet can emit.

Questions the headline leaves open

Did OpenAI independently verify more than 100 solved problems?

The public record supports an attributed company claim, not independent certification of the full batch. Earlier individual releases have richer artifacts; keep those evidence states separate.

Is Navier–Stokes now solved?

Clay currently labels the problem Active and describes it as apparently settled, while reserving deliberate evaluation and credit assignment. OpenAI’s construction concerns smooth external forcing, initially stationary fluid, finite energy, and finite-time blowup—not an unforced Navier–Stokes result. Clay’s prize process has publication, time, and general-acceptance requirements. “Apparently settled,” “formally encoded,” and “prize awarded” are different states.

Does AGMAI control OpenAI's mathematics research?

No. The group’s charter says it has no decision-making authority inside companies. Advice about disclosure is not operational control over research.

Does this demonstrate fully autonomous recursive self-improvement?

No. It demonstrates a company-reported jump in mathematical output from an internal model. Fully autonomous RSI would require a closed successor-development loop with much less human selection, evaluation, authorization, and intervention than either OpenAI or Anthropic currently reports. OpenAI explicitly says that condition is not happening today.

The next breakthrough is the review layer

OpenAI's claim may mark an extraordinary expansion in AI-assisted mathematics. The responsible response is neither instant coronation nor reflexive dismissal. It is to demand an evidence system capable of preserving what is already strong—the formal artifacts, inspectable theorem scope, and serious expert engagement—while exposing what remains aggregate, self-reported, disputed, or simply unfinished.

The hundred-plus number tells us that generation may no longer be the scarce resource. Review bandwidth, explanation, attribution, benchmark hygiene, and decision authority are becoming the scarce resources. AGMAI can help with one part of that stack. Technical standards may eventually connect another. Existing company policies cover still another. None should be mistaken for the whole operating system.

If OpenAI publishes a versioned result manifest and lets claim states mature in public, the story will become much bigger than a one-day headline. It will show how machine-speed discovery can enter human knowledge without making trust itself a casualty of scale.