Skip to content
← All news

News1 min read

Thin docs, no corpusKDA still finds speedups on AMD hardware

On ASM, HIP and ROCm — far less documentation, far fewer published kernels to copy from — KDA still lands the speedup. Two of them are merged in FlyDSL.

Everything published about KDA so far has been on NVIDIA hardware with excellent documentation, a mature profiler and a large public corpus of prior kernels. That is a reasonable place to start and a bad place to stop, because it leaves open the obvious objection: perhaps the agent is retrieving, not reasoning.

So we pointed it at hardware where there is much less to retrieve. On ASM, HIP and ROCm — including AMD parts with substantially thinner documentation and far fewer published kernels — KDA still finds the optimization opportunity and still lands the speedup. Two of them are in FlyDSL, as pull requests #711 and #685.

This is the property that decides whether any of this is durable. New accelerators arrive with no corpus at all: NVIDIA's Rubin parts and Groq's hardware will both, at some point, be something an agent has read nothing about. A loop that works only where the answers already exist on the internet is a search engine with extra steps. A loop that works on the thin documentation is a method.

KDA · NVlabs/kda

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it