Results
Numbers somebody else can check. Each was produced by a flow you can read, and scored by a verifier or a leaderboard we do not control. One page each, because a result and its caveats belong together rather than in a summary.
IMO 2026 — six of six
All six problems, machine-checked in Lean 4. Three times faster than the previously reported agentic result, at less than half the cost.
Formal mathematicsPutnamBench — 670 of 672
99.7% verified and first on the official leaderboard, with three independent checkers having to agree before a proof counted.
Systems performanceMLSys 2026 kernel contest
First, second and third on tracks of the FlashInfer contest, with MIT HAN Lab, on the contest's own hardware.
Live competitionKaggle — thirteen competitions
Sixteen official final ranks and thirty-three late-submission estimates, audited in public and never mixed together.
Check it yourself
We would rather be checked than believed, so every claim comes with the thing that produced it. Each page above ends with how to re-run it; this is the short version.
| Claim | How to check it |
|---|---|
| IMO 2026, 6/6 | Clone humanfia/imo2026 and run the AXLE verification script against the published Lean solutions. |
| PutnamBench, 99.7% | The official leaderboard, plus the AXLE API, which needs Python 3 and a network connection and nothing else. |
| Kernel contest placements | The contest repository has the evaluation and reproduction code. |
| Kaggle results | The audit states its snapshot times, its exclusions, and what each figure is not. |
On what these results do and do not show
None of this says agents are good at mathematics or at kernels. It says that the arrangement around an agent — who works, who reviews, what is remembered, what is deliberately forgotten, when to stop — is worth a large multiple on hard work, and that the multiple can be measured by people who are not us. That is the claim, and it is the reason FlowBench exists.