Blog
One result, one post. There is no separate results page any more: a number and its caveats belong in the same place, dated, with the people who produced it named at the top. Subscribe by RSS.
Lean-Eval: first place, and 172 research-level proofs
A fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level mathematics problems — every accepted proof sorry-free and independently re-verified.
Zhengyang Zhang · Hongzhou LinRead it
The review is the next prompt
Read it
Four layers and a referee
Read it
Ten workflows, nineteen competitions
Read it
The MSA indexer, 3.3× faster in production
KDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, 3.3× on long-context decode, and bitwise-identical output throughout.
Jiaming TangRead it
ProgramBench: 3.5%, against 0.5% and 0%
The same two models that solve 0.5% and 0% of ProgramBench on their own solve 3.5% of it arranged as a builder and a reviewer in a loop — and the curve had not flattened when the clock ran out.
Read it
A quality tier, so the fast kernel can ship at all
Most profitable diffusion fusions change the output at the level of bf16 rounding order — enough to fail a bitwise CI, not enough for a human to see. So we gave the request a quality tier instead of arguing about the diff.
Read it
KDA 1.5 goes past human SOTA on every contest track
A rerun of the MLSys 2026 FlashInfer kernel contest with CuteDSL support, a profiling skill and a better loop — 1.25× to 1.39× past the best human entries on all three tracks.
Dongyun ZouRead it
QCode Discovery, made fail-closed
Read it
Physics and quantum, formalized end to end
Read it
Lean-Eval: 149 of 219, and second place
A general agent loop on a general model, second on a leaderboard led by a purpose-built prover — and six problems behind it.
Read it
Six of six at IMO 2026, on two different backends
A fully agentic, YOLO-style run solved every problem of the 2026 International Mathematical Olympiad, machine-checked in Lean 4 — 3.2× faster than the previously reported agentic result, with both backends closing all six.
Read it
Humanfia: from automated idea factory to realization
Read it
Model level, tool level, flow level
The same model, measured three ways — raw API, the vendor's CLI, and the CLI inside a flow — across PutnamBench, Physics Cup, SuperChem and an HLE subset. The flow is worth more than the gap between model generations.
Read it
First on SOLExec Bench, L1 single operations
KDA-generated kernels took the L1 single-operation track at 0.7608, confirmed as first place against a previous best of 0.7584.
Read it
Lossless is a choice: the numerics under SGLang-Omni
A competing stack bought its speed with approximate kernels and skipped steps. Reading their implementation carefully turned into three lossless PRs and 20%+ on our own LTX2.3 baseline.
Read it
670 of 672 on PutnamBench
Read it
Fifty-three firsts on SOL Bench, from one node
A week of unattended kernel generation on a single 8×B200 node produced 53 first-place rankings, and it had not stopped finding them.
Lesheng Jin · Yuchen JinRead it
KDA generalizes to hardware it was never tuned for
ASM, HIP and ROCm — the loop finds optimization opportunities on less popular, less documented hardware too, which is the property that matters for the parts that do not exist yet.
Jin PanRead it
Forty operators into SGLang
Read it
Top three on every track at the MLSys FlashInfer contest
Read it
One person, some agents, and gem5's build system
Read it