Morgan Vale

← All notes

722 proofs in one evening, three withdrawn the next day

OpenAI released hundreds of mathematical results from one of its internal models in a single batch. The next day it withdrew three. The repository says more than the headlines did.

Illustration in two halves. On the left, “The Headline”: a hand presses a SINGLE PROMPT button in front of a screen marked AI and the headline AI Solves Hundreds of Open Math Problems. On the right, “The Document”: a man at a Lean proof checker showing 42 percent verified, papers stamped SIGN ERROR and WITHDRAWN, and Andrew Sutherland's line “We should ask for receipts.”

The headline that went around

OpenAI's AI solves hundreds of open math problems with a single prompt.

What the document says

On the evening of October 6, OpenAI, the company behind ChatGPT, published 722 mathematical manuscripts on GitHub, grouped into 372 families of results and produced by an internal model that is not available to the public. According to the repository, the model was posed about 4,000 problems, and on average each result used the compute equivalent of about three hours of ChatGPT Pro reasoning. The post that accompanies the release is short: it announces the results, the proofs formalized in Lean, a language in which a computer checks every step of an argument, and ten summaries of the model's reasoning.

The repository is more careful than the headline that went around. It says the collection includes results at different stages of verification, and that some of the unformalized ones could have issues. Today 300 of the 719 main results, about 42 percent, are verified in Lean. Two of the most ambitious results, on the Riemann zeta function and on the Hodge conjecture, were not obtained with the standard procedure, and the write-up of one of them was edited by people for readability.

On October 7, the next day, the repository's history records the first withdrawals. A sign error invalidates an argument in one manuscript and the construction two others relied on. All three were withdrawn, including one on the rational Hodge conjecture for products of K3 surfaces. In the same update fourteen manuscripts were corrected, and thirteen updated their citations to related papers.

The single-prompt claim is not in the repository. A spokesperson made it to Scientific American: almost every result came from one request handed to one agent. The same spokesperson added that some results may have taken multiple attempts, and that many are not yet understood by the company's own mathematicians. Andrew Sutherland, a mathematician at MIT, said the claim should be treated as unverified until the model is released and the results can be replicated: “We should ask for receipts.”

What we don't know

How many of the results not yet verified in Lean, more than half, will survive mathematicians' scrutiny. It will take months, and Scientific American reports very different judgments on how readable the papers are.

Which prompts were used, and which of the roughly 4,000 problems the model failed to solve. The advisory group of mathematicians that OpenAI convened in September recommends disclosing the model, the exact prompt and the compute behind each result. OpenAI published only the average compute, and told Scientific American it is not bound by those recommendations.

Who chose the problems. Several mathematicians who spoke to Scientific American suspect the model proposed them itself; OpenAI did not answer the question.

Who found the error behind the three withdrawals. Scientific American says experts uncovered it; the repository does not say.

The chapter in the book

Welcome to the Era

Read the book

The method is always the same: the headline against the original document.