Skip to content
← All news

News2 min read

30 first solves on Lean-Eval v1, the most of anyone — and second by one on total

Lean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.

In August this flow took first place on the Lean-Eval leaderboard. That board no longer exists in the same shape. The maintainers have frozen a first release, LeanEval v1, of 128 problems (published 20 August 2026), and moved the older problems to an Archive of 171. The leaderboard now shows one scope at a time, and a single "first place" no longer describes it. Here is where we stand in each.

All numbers below come from the board's own published data, generated 2026-10-05 08:56 UTC. Our entry is listed as "Humanifa + GPT 5.6 sol" (sic).

LeanEval v1: 128 problems, 88 solved by anyone ​

EntryUniqueFirst solvesTotal
Axiom Prover (Axiom Math)32174
NEAR AI, with DeepSeek V431280
Humanfia, GPT-5.613079
  • First solves: first, with 30. We were the first to have an accepted proof on 30 of the 88 problems anybody has solved.
  • Total: second, with 79, one behind NEAR AI's 80.
  • Unique: fourth. Only one of our 79 has no other solver. The board sorts by this column by default, so that is the position a visitor sees first.

The three columns reward different things: being first, being alone, or being broad. We lead on the first and are close on the third. Others have more problems nobody else can do.

Archive: 170 of 171 ​

On the 171 archived problems we have accepted proofs for 170, tied for the most with two other entries. Counted across both scopes, 249 distinct problems carry an accepted Humanfia proof, the most of any entry on the board. That count is ours, made from the published data; the site does not show a combined column. The board also notes that one model may appear under more than one name while names are consolidated.

What counts ​

Lean-Eval accepts a solution only when it passes Comparator against the problem's statement. There are no partial credits and no sorry. Of the entry's 249 accepted solutions, 210 were submitted by Zhengyang Zhang, 34 by Hongzhou Lin and 5 by Jui-Hui Chung; 75 were accepted on or after 19 August.

The leaderboard · its source · HOA

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it NewsFrom #2 to #1 on Lean-Eval in three weeks: 172 research-level proofsA fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level problems — every accepted proof sorry-free and independently re-verified.Read it