News
One result, one post. A number and its caveats belong in the same place, dated, with the people who produced it named at the top and a link to whatever it can be checked against. The longer arguments are on the blog. Subscribe by RSS.
672 of 672: we closed PutnamBench's last two, and three others got there too
GPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.
Ligeng Zhu · Zhengyang ZhangRead it
Full marks on five science exams, and the fine print on every one
Jing Xiong, Zhengyang Zhang, Ligeng ZhuRead it
30 first solves on Lean-Eval v1, the most of anyone — and second by one on total
Zhengyang Zhang, Hongzhou Lin, Jui-Hui ChungRead it
Six KDA pull requests land in SGLang, the best adding 8.7% end to end
Xiaoyu Zhang, Dongyun ZouRead it
Biohub closed at #188 of 3,947: top 5%, not the #8 the public board showed
In August an agent workflow sat 8th on the Biohub cell-tracking public leaderboard. The private board has now spoken: 188th of 3,947, top 4.8%. A real finish, a smaller one, and what the Kaggle tally looks like when only final ranks count.
Changye LiRead it
78.2% medal rate on MLE-bench in six hours, by making two agents take turns
Two coding agents alternating over one workspace, each starting fresh, beat both of them working alone on 75 MLE-bench tasks: 78.2% any-medal against 72.4% and 68.0%. Self-reported, and not on the official leaderboard.
Changye LiRead it
From #2 to #1 on Lean-Eval in three weeks: 172 research-level proofs
A fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level problems — every accepted proof sorry-free and independently re-verified.
Zhengyang Zhang, Hongzhou LinRead it
14 of 19 Kaggle competitions in the top 5%, five inside the top 1%
Ten agent workflows on live Kaggle competitions: of nineteen completed, fourteen finished in the top 5% — and our first read is that disagreement between agents, not a better model, did it.
Changye Li · Menghan Li · Yitong Liu · Zijian ZhangRead it
MSA indexer: 6.5× faster prefill in production, bitwise identical
Jiaming TangRead it
0% and 0.5% alone, 3.5% together: a review loop on ProgramBench
Zheng DuRead it
KDA 1.5 beats the best human kernels on every MLSys track, up to 1.39×
A rerun of the MLSys 2026 FlashInfer contest with CuteDSL, a profiling skill and a better loop: 1.25× to 1.39× past the best human entries — including a track where KDA 1.0 was slower than the baseline.
Dongyun ZouRead it
No certificate, no win: QCode Discovery's search now fails closed
17,520 definitions across 21 lattices, and no candidate counts as a win without an exact construction, an independent verification and a replay against the IBM baselines.
Jing XiongRead it
23/23 IPhO 2026 subproblems and 36/36 quantum tasks, proved in Lean
Jing Xiong, Zhengyang ZhangRead it
#2 on Lean-Eval: a general loop finishes 6 problems behind a prover
149 of 219 with a general model in a general agent loop — second on a leaderboard led by a purpose-built prover at 155, and close enough to be worth chasing.
Zhengyang ZhangRead it
6 of 6 at IMO 2026, Lean-checked and 3.2× faster than AxiomProver
A fully agentic run solved every problem of the 2026 International Mathematical Olympiad, each proof machine-checked in Lean 4 — on two different backends, and in 3.2× less API time than the previously reported agentic result.
Zhengyang ZhangRead it
First on SOLExec L1 — by 0.0024, on a track everyone had tuned
KDA-generated kernels took the L1 single-operation track of SOLExec Bench at 0.7608 against a previous best of 0.7584. A small margin, and why it is the better evidence.
Lesheng Jin, Yuchen JinRead it
670 of 672 on PutnamBench: a Codex loop passes two purpose-built provers
Zhengyang ZhangRead it
One node, one week, nobody watching: 53 first places on SOL Bench
A week of unattended kernel generation on a single 8×B200 node took 53 first-place rankings on SOL Bench — and it had not stopped finding them.
Lesheng Jin · Yuchen JinRead it
Thin docs, no corpus: KDA still finds speedups on AMD hardware
On ASM, HIP and ROCm — far less documentation, far fewer published kernels to copy from — KDA still lands the speedup. Two of them are merged in FlyDSL.
Jin PanRead it
SGLang merges 40+ agent-tuned operators, one lifting throughput 71%
Xiaoyu ZhangRead it
33× on DSA: KDA kernels place top three on every MLSys 2026 track
Dongyun ZouRead it
567 files, one engineer: agents port gem5's build from SCons to CMake
Sihao LiuRead it