Skip to content
← All news

News1 min read

33× on DSAKDA kernels place top three on every MLSys 2026 track

KDA-designed kernels placed in the top three on all three tracks of the MLSys 2026 FlashInfer contest — 1.4× on FP8 MoE, 33.3× on DSA and 17.6× on GDN over the baseline.

Kernels designed with KDA placed in the top three on every track of the MLSys 2026 FlashInfer kernel contest. Speedups over the FlashInfer baseline:

TrackSpeedup
FP8 MoE1.4×
DSA33.3×
GDN17.6×

A kernel contest is the closest thing to an ideal test for an agent loop. The score is a wall-clock measurement, the machine belongs to the organizers, the reference implementation is theirs, and correctness is a gate rather than a matter of interpretation — a kernel that is faster because it has quietly become incorrect does not place, it fails.

The kernels are published.

The contest kernels · the contest · KDA

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it