Sharing AI progress in mathematics
Learn how large language models are being used to attack deep unsolved math problems.
OpenAI’s public math repo reports that large language models have made substantial progress on four of the seven Millennium Prize problems and claim to have resolved one of them. The four listed are the Hodge conjecture, Birch‑Swinnerton‑Dyer, the Riemann hypothesis, and Navier‑Stokes. The claim stands out because these are problems that have resisted conventional approaches for decades.
The repository contains a collection of preprints generated by the models, each describing a strategy for a specific problem. For Hodge and Birch‑Swinnerton‑Dyer, the models produced candidate lemmas and computational evidence that align with known partial results. In the Navier‑Stokes entry, the model outlined a regularity argument that matches the structure of existing blow‑up criteria.
The approach relies on prompting the model with formal statements, feeding it known theorems, and iteratively refining drafts through human‑in‑the‑loop feedback. The system treats the model as a symbolic assistant, extracting algebraic manipulations and suggesting conjectural bridges. Numbers of iterations and token counts are not disclosed, but the workflow emphasizes tight coupling between model output and expert verification.
For anyone building research tools, this suggests a new pipeline: generate proof sketches with an LLM, then let domain experts prune and validate the steps. Integrating the model into proof assistants could automate routine lemma discovery, freeing researchers to focus on high‑level insight. The open‑source repo also provides a template for reproducing the experiments on other open problems.
TakeawayLLMs can now produce plausible proof drafts for deep mathematics, making them useful research collaborators.