Skip to content
← All news

News1 min read

From #2 to #1 on Lean-Eval in three weeks: 172 research-level proofs

A fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level problems — every accepted proof sorry-free and independently re-verified.

Three weeks ago this flow was second on Lean-Eval, 149 of 219, behind a specialist prover. It is now first, with 172 research-level problems solved and machine-checked in Lean 4, using GPT-5.6 and a fully agentic workflow.

The model is one anybody can rent. What changed is the arrangement around it.

What the loop does ​

Argue first, formalize second. The flow writes a natural-language proof strategy before it writes any Lean, then converts that strategy into Lean statements. A formalization that starts from a plan fails in ways that are legible; one that starts from a blank file fails in ways that are not.

The compiler is the reviewer that never gets tired. Lean's error output is fed back into the run and the code is refined against it, round after round, by HOA and Humanize 1 agents. Nothing about this is clever — it is just that almost nobody does it for hundreds of rounds without a human losing patience.

Seed the hard ones with the easy ones. For the problems that resist, the argument is first turned into a precise, stepwise informal plan, and the run is seeded with reusable Lean proofs from related problems that already closed. The agent then spends its budget formalizing the ideas that are actually missing rather than re-deriving the ones that are not.

What counts as solved ​

Two independent gates, on two different agents, before anything is accepted:

  • the worker runs a comparator check on its own output; and
  • the reviewer re-verifies through the AXLE API, which has no access to how the proof was arrived at.

Nothing is accepted with a sorry placeholder or an unproved assumption standing in for a step. That bar is the whole reason a leaderboard position in formal mathematics means anything at all: a proof either has a term the Lean kernel accepts, or it does not exist.

HOA · the flows · the runtime

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it