Skip to content
← All news

News2 min read

672 of 672we closed PutnamBench's last two, and three others got there too

GPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.

In June this solver stood at 670 of 672. It has now closed the last two. Every one of PutnamBench's 672 formal statements has a Lean 4 proof that the PutnamBench team accepted, and the official leaderboard lists Humanfia at 672 (trishullab/PutnamBench#350, merged 1 September 2026).

We are not alone up there. Four entries now report 672 in the Lean column with answers given, in the leaderboard data:

EntryAddedReported cost
Aleph Prover (Logical Intelligence)2026-08-26average $74 per problem, maximum $1,468
Humanfia — GPT-5.6, xhigh reasoning2026-08-27average $44.50 per problem
NEAR AI, with DeepSeek V42026-09-04mean $0.17, median $0.04 per problem
Forall (Astrio), with Claude Opus 52026-09-24mean $3.52 actor-side per problem

The costs come from each team's own note on the leaderboard, and they are not measured the same way: one is total spend, one is actor-side only. Read them as orders of magnitude, not as a ranking. Read that way, the benchmark is saturated, and what separates the leaders now is the price of a proof, which is where we are not first. Two of the four got to 672 for a tenth of our cost or less.

The board also has a separate "no answer" column, where the prover is not given the answer to problems that ask for one. There, Midas Prover reports 672 at an average of $6.17 per problem. Our run is in the column with answers, so the two are not the same test.

What counts as a proof ​

A problem counts only when the Lean file passes every gate in the solver's verification method, and the model that wrote the proof never decides whether it is accepted:

  • the statement is byte-identical to the pinned benchmark statement;
  • no sorry, admit, axiom or native_decide;
  • it compiles against the pinned Lean 4.27.0 and Mathlib;
  • Comparator checks the statement and replays the proof through the Lean kernel; and
  • a separate reviewer, which cannot edit the file, gets okay: true back from the AXLE verification API.

The worker's shell has no network access, and no existing PutnamBench solution is ever mounted into either model's workspace.

Checking it ​

At the PutnamBench authors' request the full proof set is not public. The first 12 proofs are, along with the whole pipeline, and checking them through AXLE takes Python 3 and a network connection: the steps are in the repository.

humanfia/hoa-qed · HOA

Read next

NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it NewsFrom #2 to #1 on Lean-Eval in three weeks: 172 research-level proofsA fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level problems — every accepted proof sorry-free and independently re-verified.Read it