Whether frontier AI models can produce mathematics that mathematicians did not already have, or whether they are very fast literature searches, has been the argument of the summer. On Monday, Anthropic published a research note on Claude's mathematical capabilities saying an unreleased research version of the model improved a longstanding result tied to the Riemann hypothesis, raising a proven lower bound from 41.6% to 67.2%.
The Riemann hypothesis, posed in 1859 and one of seven problems carrying a $1 million Clay Mathematics Institute prize, holds that the zeros of a function governing the spacing of prime numbers all sit on one vertical line. Nobody has proved that for every zero, so mathematicians settle for proving it about a fraction of them. Decades of incremental work had lifted that fraction to 41.6%. The new figure is not a progress bar toward a proof, and Anthropic said in the same note that it does not expect Claude's technique to produce one.
For investors, the sharper detail is which model did this. Anthropic has not named the checkpoint, released the weights or made the system available. OpenAI did the same thing on August 1, when it published solutions to 10 open problems in mathematics and theoretical computer science and credited an internal version of Astra, a model with no release date or price. Both labs are now resting their strongest capability claims on systems nobody outside can test.
The note also lands three weeks after Anthropic's previous mathematical result. In July, Levent Alpöge, a number theorist at the company, posted a counterexample found with Claude Fable 5 that disproved the 87-year-old Jacobian conjecture in three dimensions and above, and Terence Tao worked through the counterexample publicly the following day. Alpöge is one of the two Anthropic mathematicians who went on to check the zeta result.
A non-mathematician ran the session
The part of the account drawing the most attention is who was at the keyboard. Jarred Sumner, the staff member who ran it, is not a mathematician, and Anthropic says he left the mathematical choices to the model. Results of this kind have so far come from mathematicians steering models with expert prompts. Anthropic writes that during the decisive run, Sumner's input was "mostly limited to sending Claude messages of encouragement," mostly variants of "keep going" or "believe in yourself." The company says that appears to have helped the model past its own early skepticism that it could get anywhere.
The computation was less delicate. Claude tried 650 ideas in the first pass and none worked. Told to try again, it spent about a day and a half coordinating roughly 60 copies of itself, which Anthropic calls subagents. Between them they ran 2,400 shell commands and wrote hundreds of Python scripts, testing candidate arguments against known zeros. Two of the 60 produced the key ideas, 30 produced nothing usable, and the two sessions consumed 31 million output tokens, the units AI labs bill by.
Claude then attacked its own result, having subagents hunt for counterexamples, pull 54 papers from arXiv to confirm the finding was new, and re-derive it from scratch. It proposed writing the work up as a paper and recommended that a human number theorist check it.
The result has not been peer reviewed
Alpöge and Ralph Furman validated the paper inside Anthropic. Brian Conrey and Dan Goldston, two specialists in this corner of number theory, read it on short notice at the company's request. Conrey set the 40% mark on this same bound in 1989. Expert eyes on short notice are not journal refereeing, which runs on months rather than news cycles.
Claude also produced a version of the proof a computer can check, written in the Lean proof assistant and published as a public repository. That rules out the quiet logical gaps that sink most claimed proofs of famous problems. It does not establish that the formalized statement is the one number theorists care about, and it says nothing about the model that generated it.
One figure is missing. Anthropic disclosed 31 million output tokens but no dollar cost, and with the model unreleased there is no published rate to multiply by. OpenAI estimated the tokens behind its 10 proofs at roughly $2,000. Anyone trying to judge whether this kind of research is cheap enough to run at industrial scale needs Anthropic's version of that number.
Until then the checkable claim is the narrow one. A model that had to be told to keep going moved a bound that human mathematicians had advanced by 1.6 percentage points in the 37 years since Conrey's result.
