Skip to content

News ​

One result, one post. A number and its caveats belong in the same place, dated, with the people who produced it named at the top and a link to whatever it can be checked against. The longer arguments are on the blog. Subscribe by RSS.

HOA

672 of 672: we closed PutnamBench's last two, and three others got there too

GPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.

Ligeng Zhu · Zhengyang ZhangRead it

HOA

Full marks on five science exams, and the fine print on every one

Jing Xiong, Zhengyang Zhang, Ligeng ZhuRead it

HOA

30 first solves on Lean-Eval v1, the most of anyone — and second by one on total

Zhengyang Zhang, Hongzhou Lin, Jui-Hui ChungRead it

KDA

Six KDA pull requests land in SGLang, the best adding 8.7% end to end

Xiaoyu Zhang, Dongyun ZouRead it

HMA

Biohub closed at #188 of 3,947: top 5%, not the #8 the public board showed

In August an agent workflow sat 8th on the Biohub cell-tracking public leaderboard. The private board has now spoken: 188th of 3,947, top 4.8%. A real finish, a smaller one, and what the Kaggle tally looks like when only final ranks count.

Changye LiRead it

HMA

78.2% medal rate on MLE-bench in six hours, by making two agents take turns

Two coding agents alternating over one workspace, each starting fresh, beat both of them working alone on 75 MLE-bench tasks: 78.2% any-medal against 72.4% and 68.0%. Self-reported, and not on the official leaderboard.

Changye LiRead it

HOA

From #2 to #1 on Lean-Eval in three weeks: 172 research-level proofs

A fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level problems — every accepted proof sorry-free and independently re-verified.

Zhengyang Zhang, Hongzhou LinRead it

HMA

14 of 19 Kaggle competitions in the top 5%, five inside the top 1%

Ten agent workflows on live Kaggle competitions: of nineteen completed, fourteen finished in the top 5% — and our first read is that disagreement between agents, not a better model, did it.

Changye Li · Menghan Li · Yitong Liu · Zijian ZhangRead it

KDA

MSA indexer: 6.5× faster prefill in production, bitwise identical

Jiaming TangRead it

Humanize 2

0% and 0.5% alone, 3.5% together: a review loop on ProgramBench

Zheng DuRead it

KDA

KDA 1.5 beats the best human kernels on every MLSys track, up to 1.39×

A rerun of the MLSys 2026 FlashInfer contest with CuteDSL, a profiling skill and a better loop: 1.25× to 1.39× past the best human entries — including a track where KDA 1.0 was slower than the baseline.

Dongyun ZouRead it

Humanize 2

No certificate, no win: QCode Discovery's search now fails closed

17,520 definitions across 21 lattices, and no candidate counts as a win without an exact construction, an independent verification and a replay against the IBM baselines.

Jing XiongRead it

HOA

23/23 IPhO 2026 subproblems and 36/36 quantum tasks, proved in Lean

Jing Xiong, Zhengyang ZhangRead it

HOA

#2 on Lean-Eval: a general loop finishes 6 problems behind a prover

149 of 219 with a general model in a general agent loop — second on a leaderboard led by a purpose-built prover at 155, and close enough to be worth chasing.

Zhengyang ZhangRead it

HOA

6 of 6 at IMO 2026, Lean-checked and 3.2× faster than AxiomProver

A fully agentic run solved every problem of the 2026 International Mathematical Olympiad, each proof machine-checked in Lean 4 — on two different backends, and in 3.2× less API time than the previously reported agentic result.

Zhengyang ZhangRead it

KDA

First on SOLExec L1 — by 0.0024, on a track everyone had tuned

KDA-generated kernels took the L1 single-operation track of SOLExec Bench at 0.7608 against a previous best of 0.7584. A small margin, and why it is the better evidence.

Lesheng Jin, Yuchen JinRead it

HOA

670 of 672 on PutnamBench: a Codex loop passes two purpose-built provers

Zhengyang ZhangRead it

KDA

One node, one week, nobody watching: 53 first places on SOL Bench

A week of unattended kernel generation on a single 8×B200 node took 53 first-place rankings on SOL Bench — and it had not stopped finding them.

Lesheng Jin · Yuchen JinRead it

KDA

Thin docs, no corpus: KDA still finds speedups on AMD hardware

On ASM, HIP and ROCm — far less documentation, far fewer published kernels to copy from — KDA still lands the speedup. Two of them are merged in FlyDSL.

Jin PanRead it

KDA

SGLang merges 40+ agent-tuned operators, one lifting throughput 71%

Xiaoyu ZhangRead it

KDA

33× on DSA: KDA kernels place top three on every MLSys 2026 track

Dongyun ZouRead it

Humanize 1

567 files, one engineer: agents port gem5's build from SCons to CMake

Sihao LiuRead it