Skip to content
← All news

News1 min read

KDA 1.5 beats the best human kernels on every MLSys track, up to 1.39×

A rerun of the MLSys 2026 FlashInfer contest with CuteDSL, a profiling skill and a better loop: 1.25× to 1.39× past the best human entries — including a track where KDA 1.0 was slower than the baseline.

In May, KDA took top-three placements on every track of the MLSys 2026 FlashInfer kernel contest. That was a good result against the field. It was not a good result against the field's best entries, which were written by people.

KDA 1.5 changes three things and reruns the contest:

  1. CuteDSL support, so the agent can express the layouts it wants rather than the ones the previous toolchain made convenient;
  2. the IKET profiling skill, which turns a profiler report into something an agent can reason from instead of a wall of counters; and
  3. a better agent loop flow — the arrangement of who writes, who checks and what carries between attempts.

The rerun ​

All numbers are speedups over the FlashInfer baseline, on B300.

TrackKDA 1.0 · 2026-05Human SOTA · 2026-05KDA 1.5 · 2026-07KDA 1.5 over human
MoE FP8 block scale0.67×1.80×2.25×1.25×
DSA sparse attention11.91×22.99×29.95×1.30×
GDN prefill1.16×4.42×6.10×1.39×

The MoE row is the one worth sitting with. KDA 1.0 was slower than the baseline there — 0.67× — and the same application, with a profiling skill and a different loop, is now past the best human entry. Nothing about the model changed.

A workload nobody had tuned ​

Contest tracks have been optimized by many people, which makes them a fair test and a stale one. So we also pointed KDA 1.5 at Flash-KMeans, which nobody had worked over: balanced k-means for the Wan 2.2 workload on B200 came back 6.1× faster.

KDA · the contest kernels

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it