Skip to content
← All news

News1 min read

670 of 672 on PutnamBencha Codex loop passes two purpose-built provers

One Ralph loop with Codex as both worker and reviewer closed 99.7% of PutnamBench and all of Putnam 2025 — past an EBM-based prover at 99.4% and an autoregressive one at 88.8%.

This has since moved

The solver later closed the last two: 672 of 672, joint first on the official leaderboard, every proof through Lean 4, Comparator and AXLE. The solver's README has the run.

PutnamBench is 672 formal statements. A Ralph loop running Codex as both worker and reviewer closed 670 of them — 99.7% — including every problem in Putnam 2025.

Solved
Humanfia — Ralph loop, Codex worker and reviewer670 / 672 · 99.7%
Logical Intelligence's Aleph, EBM-based99.4%
Goedel-Architect, autoregressive88.8%

The two that are not closed are not closed. They are drawn as such everywhere we show this.

Where the turns go ​

More than 90% of problems are solved within ten turns. The loop is not, for most of the benchmark, doing anything exotic — it writes a proof, the compiler objects, it fixes it, and it is done well inside the budget.

The tail is where the argument is. For the harder problems, the extra multi-turn trials raise the solve rate measurably: the same worker, given more rounds against the same feedback, closes problems it did not close in ten. That is the entire thesis of this site stated as an experiment — the model was capable of the proof the whole time, and what was missing was the arrangement that let it keep going without wandering.

Checking it ​

The official leaderboard is maintained by the PutnamBench team. Proof files are published as a dataset, and a preview set is open for review; the full solving pipeline is public.

humanfia/hoa-qed/putnambench · HOA

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it