Skip to content
← All news

News1 min read

#2 on Lean-Evala general loop finishes 6 problems behind a prover

149 of 219 with a general model in a general agent loop — second on a leaderboard led by a purpose-built prover at 155, and close enough to be worth chasing.

This has since moved

Three weeks later the same loop took first place with 172 problems.

Running on gpt-5.6-sol at max effort, Humanize 2 solved 149 of 219 problems on Lean-Eval, which puts it second on the leaderboard.

First is ByteDance's Seedprover, at 155 of 219.

Six problems is not a rout in either direction, and the gap is worth naming precisely because of what is on each side of it. Seedprover is a system built for this: a prover, trained and tuned for formal mathematics. What is second is a general-purpose model driving a general-purpose agent loop, with no component anywhere in it that knows Lean specifically — the loop is the same one that runs kernel work and Kaggle competitions.

A specialist should beat a generalist on the specialist's benchmark. That it does so by six problems is the interesting number, and it is why we thought the gap was closable.

HOA · the flows

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it