KDA
Kernel Design Agents. An agent workflow for the one kind of programming where the score is never in doubt: making a kernel faster, on real hardware, without making it wrong.
mit-han-lab/kernel-design-agents · KDA-Pilot · built with MIT HAN Lab
What it has done
| Result | Written up | |
|---|---|---|
| MSA indexer, in production | 6.5× geomean on prefill, 3.3× on long-context decode, bitwise identical | 26-08-14 |
| MLSys 2026 FlashInfer | Past human SOTA on all three tracks, 1.25–1.39× | 26-08-02 · 26-05-15 |
| SGLang-Diffusion | A quality tier, so a bf16-order-changing fusion can ship at all | 26-08-06 |
| SOLExec Bench | First on the L1 single-operation track, and 53 firsts in batch | 26-07-02 · 26-06-22 |
| SGLang, upstream | More than forty operators merged, with the numbers attached | 26-06-05 · 26-07-02 |
| Other hardware | ASM, HIP and ROCm — it works where the corpus is thin | 26-06-15 |
The problem
Kernel work is the worst case for a coding agent and the best case for a good loop.
A change is a one-line edit and a three-hour investigation. The feedback is a number, but the number is noisy, hardware-specific and easy to fool: a kernel that is faster because it is now subtly incorrect will happily report a speedup. The search space is enormous and mostly bad. And what separates a good attempt from a hopeless one lives in profiler traces, architecture manuals and other people's kernels, not in a docstring.
An agent told to "optimise this kernel" and left alone will produce something plausible, report a win, and be wrong. The engineering is all in what happens around that.
What KDA does about it
Research before writing. The agent gets the reference material a human would want — a profiler skill that turns an ncu report into something readable, and a kernel wiki of techniques and prior art — and is expected to come back with a plan grounded in the repository and the hardware rather than a hunch.
A contract, before any code. The objective, the constraints, the validation command and the bar for promoting a candidate are written down first. Everything after is judged against that, which is what stops "it got faster" from quietly replacing "it got faster and is still correct".
Small iterations, each verified. Implement, validate, benchmark, profile, decide. A candidate is promoted only when it passes the check fixed in advance.
A record that outlives the run. Candidates, benchmark results, profiling evidence and promotion decisions are written down as the run goes, so another engineer can see what was tried, what passed, and why the winner won. On a week-long optimisation that matters more than any single trick.
Why it is here
KDA is one of the places our flows go to be found out. The loop is a flow like any other, the runtime under it is Humanize 2, and the score is a wall-clock measurement on somebody else's benchmark — or a pull request a maintainer who did not ask for it has to be willing to merge.
Try it
It is an early research prototype under active development, and the maintainers want feedback. It is deliberately independent of any one benchmark harness or hardware target: a downstream task brings its own evaluator, datasets, profiling tools and references.
git clone --recurse-submodules https://github.com/mit-han-lab/kernel-design-agents.gitThe repository has the agent flow, the prompt templates and the skills.