Skip to content
← All news

News1 min read

6 of 6 at IMO 2026, Lean-checked and 3.2× faster than AxiomProver

A fully agentic run solved every problem of the 2026 International Mathematical Olympiad, each proof machine-checked in Lean 4 — on two different backends, and in 3.2× less API time than the previously reported agentic result.

All six problems of the 2026 International Mathematical Olympiad, solved by a fully agentic, YOLO-style run with no human in the turn, and every solution formally verified by Lean 4. There is no rubric here and no benefit of the doubt: the kernel accepts the proof or it does not.

Two workers were run independently. Both closed all six.

The times ​

API time — the time actually spent inside model calls, which is the only measure that does not reward a slower harness — against the times reported by AxiomProver on the same statements.

ProblemHumanfia (GPT-5.6)Humanfia (Kimi-K3)AxiomProver
Q133.6 min77.7 min24 min
Q296.2 min220.2 min360 min
Q3179.4 min338.4 min869 min
Q453.3 min65.6 min39 min
Q542.4 min86.7 min65 min
Q662.7 min209.0 min139 min
Total467.6 min · 3.2×997.6 min1,496 min

The shape of that table is the point. On Q1 and Q4 — the two easiest problems — we lose. On Q3, which is the hardest, we win by nearly five times. A loop earns its keep exactly where the work is long enough for the loop to matter, and nowhere else; on a problem a strong model closes in twenty minutes, the arrangement around it is overhead.

What is released ​

The formal statements, both sets of Lean solutions and the scripts that reproduce the solving process. Pinned to Lean 4.31.0 and Mathlib.

humanfia/hoa-qed/imo2026 · AxiomMath/IMO2026 · HOA

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it