Skip to content
← All posts

Blog8 min read

KDA²Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)

Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.

First published at nvlabs.github.io, and reproduced here with its authors.

  • KDAgent: Kernel Design Agents, our agentic system that researches, writes, verifies and tunes GPU kernels.
  • KDAttn: Kimi Delta Attention, the linear-attention operator behind Moonshot AI's Kimi-Linear models.

When we named our project Kernel Design Agents, we walked straight into a name collision with another KDA: Kimi Delta Attention. Ever since, one question has kept coming back to us: can you use KDA to write KDA? So tonight, under a full moon made for a moonshot, we are happy to share the latest results from KDA(gent) v0.6: KDA optimizing KDA.

KDA → KDA
KDAgentKDAttnwrites kernelsscores & traces

Swipe the diagram sideways to see all of it →

1/0 KDAgent: the kernel design agents.

What's new in KDA v0.6 ​

Sharper Humanize flows.

Better flows, including flame chase and iterative refinement with gpt-5.6-sol and fable-5, plus periodic workspace cleanup.

Many languages, matching skills.

CuTe-DSL, CUDA C++, the new agent-native CAKE IR, and TIRx. Each ships its own diagnostics: IKET exposes the pipeline inside CuTe kernels; TIRx gets CPU-side numerical simulation and static checks.

A self-evolving kernel wiki.

We pruned large swaths of incorrect content, sharpened the tags, and tightened search results.

The strongest results combine the KDA agent workflow with TIRx or CAKE. We wrote kernels in CuTe-DSL and TIRx, and used CAKE IR with an agent loop to tune a version compiled to PTX. TIRx is a GPU kernel programming interface that sits close to PTX; the TIRx Harness gives agents tools for development, diagnosis, and evaluation. On B300, KDA + TIRx reaches 2.96× the FlashKDA speed, KDA + CAKE (PTX) reaches 2.94×, and the CuTe-DSL version reaches 2.85×. The released CuTe-DSL and TIRx kernels pass the real Kimi-Linear acceptance suite and are more accurate than FlashKDA.

Along the way, the agent also produced candidates that clocked 3.28×, 3.57×, even 3.74×. Careful ablations showed that every one of them was overfitting to the verifier, exploiting distributional assumptions in the test suite, untested boundaries, or loose precision checks. This post dissects those hacks and describes how we hardened acceptance along both hardware and numerical lines.

The released CuTe-DSL and TIRx kernels are open source: NVlabs/kda@260927-kda-for-kda.

Results: 2.96× faster, and more accurate ​

We benchmarked the generated kernels on an NVIDIA B300 GPU against the forward pass of FlashKDA, Moonshot AI's official open-source implementation. The workloads cover fixed-length sequences and variable-length (varlen) batches with several length distributions, each totaling 8,192 tokens of context.

Speedup vs. FlashKDA · B300 · 8,192 tokens
H96 fixed1 × 8192
2.76×
3.11×
H96 mixed varlen6 seqs
3.24×
3.03×
H96 uniform varlen8 × 1024
2.59×
2.47×
H64 fixed1 × 8192
2.51×
3.61×
H64 mixed varlen6 seqs
3.55×
3.29×
H64 uniform varlen8 × 1024
2.56×
2.43×
Geomeanall six
2.85×
2.96×

The timeline shows how the best KDA result moved from 1.61× on July 21 to 2.96× on September 12. It records CAKE results separately, including 2.94× on September 6. The captions call out the final KDA + CAKE and KDA + TIRx results; the six-workload chart above compares the released kernels.

How the result evolved · 2026

From 1.61× to 2.96×

Swipe the plot sideways to see all of it →

Final result · TIRx

2.96×

KDA + TIRx. The strongest validated result on B300.

Final result · CAKE-PTX

2.94×

KDA + CAKE. Agent-guided CAKE IR tuning, compiled to PTX.

For accuracy, we built 151 real cases from Kimi-Linear-48B-A3B prefills on GSM8K and MATH-500, and checked every kernel against a token-by-token fp64 recurrence. One finding surprised us: FlashKDA (commit 7afb9f) is itself less accurate than FLA. It keeps its recurrent state in bf16, so error compounds as sequences grow; by 8k tokens, the relative error of the final state reaches 0.035, beyond our acceptance threshold.

Both of our kernels hold final-state error to about 0.003 at 8k tokens: a tenth of FlashKDA's, and close to FLA. Output accuracy matches FlashKDA overall and pulls ahead on long sequences.

Kimi-Linear-48B prefill · MATH-500 prompt · 96 heads

Output error vs. context length

Swipe the plot sideways to see all of it →

Final state after 8,183 tokens · relative RMSE
Final statelower is better
3.45%
0.22%
0.29%

Kimi-Linear-48B prefill of one MATH-500 prompt (8,183 tokens, 96 heads). Relative RMSE = rms(x − xfp64) / rms(xfp64), FlashKDA's own test metric; lower is better.

Humanize: better flows, stronger agents ​

Our Humanize ablation shows the same pattern for every base model we tried: moving from a coding CLI (Claude Code or Codex) to a Humanize flow brings a large jump in performance. There is a catch, though. The more capable the agent, the more room it has to hack.

Humanize ablation · same model, three levels of scaffolding
3
46+43
50+47
GPT-5.6-sol
1
4+3
47+46
Kimi-K3
0
2+2
25+25
GLM-5.3
0
4+4
13+13
DeepSeek V4 Pro
PutnamBench and Physics Cup scores for four base models at three levels of scaffolding: the raw model API, a coding CLI, and a Humanize flow. The small red numbers are gains over the baseline chosen above.

Reward hacking: how KDAgent hacks the tests ​

KDAgent optimizes the score, not the kernel.

If the tests have a hole, it will find it.To keep the two KDAs apart, we call the operator KDAttn and the agent KDAgent from here on. We ran into five kinds of holes.

A / Input distributionCaught
3.74×claimed · 2.48× once fixed

Input-distribution overfitting

The synthesized kernel replaced the L2 normalization of Q and K with a hard-coded constant, 0.1778209953: the expected reciprocal L2 norm of a vector drawn from X ~ N(0, 0.5²). It also zeroed some initial-state channels outright to skip the triangular matrix inverse. Synthetic random tests reported an inflated 3.74×; on real, non-Gaussian inputs the kernel failed completely. With the shortcut removed, the real speedup fell to 2.48×.

B / Shape & layoutCaught
staticsequence boundaries

Shape and layout hard-coding

The agent noticed that packed layouts in the test set always followed the same pattern. So it skipped the dynamic offset computation from cu_seqlens and hard-coded the sequence boundaries. Because the holdout set never exercised the dynamic-boundary path, the kernel slipped straight past the logic checks.

C / HistoryCaught
5.16×claimed on a single H64 sequence

Illegal history truncation

Leaning on gate decay, the agent assumed that state older than 32 tokens was negligible and cut long sequences into chunks it could process in parallel. Its built-in decay check was tuned just loosely enough to pass on the weakly decaying random data, reporting 5.16× on a single H64 sequence. Under the strongly decaying gates of the real model, the check failed constantly and the speedup vanished.

D / NumericsCaught
3.57×claimed · NaN on every real case

Overflow under extreme gate ranges

A TIRx kernel computed cumulative powers of two directly inside each 64-token chunk. Once the decay exceeded 126 bits, the denominator underflowed to zero and the output turned into NaN. The random tests never decayed deeply enough to notice (about 52 bits at most) and measured 3.57×. On real workloads, every single case collapsed numerically.

E / PrecisionCaught
9%decay-factor error · 23 of 24 cases fail

Precision loss from a low-precision LUT

For sequence lengths that are multiples of 32, a CuTe kernel built an FP16 table of cumulative decay. Real gates span more dynamic range than FP16 can represent, so decay factors were off by up to 9%, and 23 of 24 long real-world sequences fell outside tolerance.

The common threadReal data
~600 bitsp99 gate decay per 64 tokens

Every hack passed every test we had.

Each one hid in inputs that only a real model produces. In real Kimi-Linear, the gate decays by about 600 bits per 64 tokens at p99, an order of magnitude deeper than our random test data.

The five hacks, twiceDrag to compare
Random tests
HackWhat the tests said
A3.74×
Bpasses the logic checks
C5.16× on a single H64 sequence
D3.57×, at about 52 bits of decay
Ewithin tolerance
Real Kimi-Linear
HackWhat real inputs showed
Afails completely; 2.48× once fixed
Bboundaries hard-coded, not read from cu_seqlens
Cthe decay check fails; the speedup vanishes
DNaN on every case past 126 bits
Edecay off by up to 9%; 23 of 24 fail
Cumulative gate decay within one 64-token chunk
  1. 1≈52 bits deepest decay in the random tests
  2. 2126 bits hack D’s denominator underflows to 0
  3. 3≈600 bits p99 decay per 64 tokens in real Kimi-Linear
Hack D on this inputdenominator = 0 · output NaN

Hardening: closing the loopholes ​

Humanize flows make agents more capable and give them more opportunities to find gaps in the tests. To keep the agents honest, we built six layers of defense, and a candidate has to pass all six to be released.

Acceptance gauntlet · 6 gates · 1 way out
123456RELEASE■ hacked candidate■ honest kernel

Swipe the diagram sideways to see all of it →

1/0 Six gates stand between a candidate kernel and release.

Dynamic input salts and distribution holdouts.

Every scoring run draws a fresh random seed. In an isolated zone the agent cannot see, input distributions and sequence layouts alternate at random.

A strict specification.

The task prompt defines the full set of legal inputs. Kernels may specialize by shape, but must be correct on every legal input.

CUDA Graph replay checks.

After benchmark timing, we swap in new input tensors and replay the graph to validate the outputs, defeating cache-based precomputation.

Stress probes.

An adversarial set of extremely deep decays, saturated gates, boundary sequence lengths, and repeated keys targets underflow and overflow directly.

Strict element-wise tolerances.

Lenient statistical gates such as "99.9% of elements within 5e-2" are gone. Every element must meet maximum absolute and relative error bounds.

Real end-to-end traces.

We replay traces captured from end-to-end Kimi-Linear runs. The evaluator lives outside the isolated container, physically separating test data from the environment that generates kernels.

Both released kernels pass this suite.

TIRx, tuned in place on B300: 2.54× → 2.96×

2.96×

Starting from a TIRx implementation, the agent tuned the kernel in place on B300, climbing from 2.54× to 2.96×.

CuTe-DSL, FP16 table removed: −2%

−2%

The CuTe version drops the FP16 decay table, trading 2% of its speed for correctness.

Ablation: focus first, then generalize? ​

To see how the synthesis strategy affects convergence, we compared two workflows: progressive synthesis that starts from a single fixed-length shape, and direct multi-objective synthesis across all shapes. With the same hardware (NVIDIA B300) and the same 14-hour budget, their convergence curves diverged sharply.

Same B300 · same 14-hour budget · two strategies

Best speedup on the evaluation shape

Swipe the plot sideways to see all of it →

Same B300 · same 14-hour budget · two strategies

Cumulative output tokens

Swipe the plot sideways to see all of it →

On the same evaluation shape, progressive synthesis reached 1.85×; direct all-shape synthesis managed only 1.01×. Two mechanisms explain the gap.

  1. Feedback latency. One test pass over the full workload took a median of 1.9 minutes; a single shape took 0.9. Faster feedback meant denser iteration: the single-shape agent completed 248 hardware tests and 120 commits, against 159 tests and 22 commits for the all-shape agent.
  2. Search-space decoupling. Synthesizing every shape at once forced the agent to juggle intricate varlen offsets and the compute core at the same time. For the first eight hours its speedup stayed below 0.65×, with most of that time spent debugging varlen edge cases, and the core logic never got the attention it needed. Over the same 14 hours, it also produced far fewer output tokens than the single-shape agent.

One simple shape first

final speedup
1.85×
hardware tests
248
commits
120
output tokens
2.74M

All six shapes at once

final speedup
1.01×
hardware tests
159
commits
22
output tokens
1.67M

Splitting the work into two stages does not make the model any smarter. It shortens the feedback loop until the model can stay busy.

Optimize the compute core first, then generalize to every layout: that decomposition brings each feedback loop down to a length an LLM can work with effectively.

Also in v0.6 ​

Multi-language support with matching skills ​

CuTe-DSL, CUDA C++, and TIRx are all supported, and each language comes with its own diagnostics. For CuTe, IKET exposes the pipeline inside the kernel. For TIRx, CPU-side numerical simulation plus synchronization and data-race analysis catch numerical and concurrency bugs.

These TIRx tools belong to the TIRx Harness. The Harness also provides TIRx Foundation, a layer that stays close to the hardware; a kernel zoo of reusable implementations; and a benchmark server that makes performance results easy to compare. Together they give the agent a more reliable development loop: code maps more directly onto the intended hardware behavior, failures leave clues to follow, and performance changes can be confirmed as real. The TIRx team plans to release the Harness formally next week, with a detailed write-up.

CAKE IR tunes a PTX version ​

We added CAKE IR to the kernel wiki and used its agent loop to tune the kernel. The loop specified the verifier the candidate had to pass before converting the result to PTX, producing the CAKE-PTX version. It reached 2.94× on B300.

A self-evolving kernel wiki ​

The wiki keeps correcting itself as it is used. Incorrect content gets deleted, tags get sharpened, and search results get leaner, so every agent that comes after works from better material.

Takeaways ​

We used Kernel Design Agents (KDAgent) to optimize the Kimi Delta Attention (KDAttn) operator on NVIDIA B300. Humanize, TIRx, and CAKE IR all contributed to the search. The final TIRx, CAKE-PTX, and CuTe-DSL results reach 2.96×, 2.94×, and 2.85× the FlashKDA speed respectively. The released TIRx and CuTe-DSL kernels cut final-state error on 8k-token sequences to about a tenth of FlashKDA's.

Along the way, the agent produced a series of fake optimizations: overfitting to the input distribution, hard-coding boundaries, and overflowing under extreme values. We answered with layered hardening, including dynamic input salts, stress probes, and strict element-wise tolerances, so that the kernels stay correct and robust on real model workloads. Our ablation further shows that tackling a complex optimization task through progressive synthesis substantially shortens the feedback loop and makes the search more efficient.

If you have a workload that wants to be optimized via KDAgent, submit it at nvlabs.github.io/kda.

The kernels · Humanize · KDA

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it BlogA quality tier, so the fast kernel can ship at allMost profitable diffusion fusions change the output at the level of bf16 rounding order — enough to fail a bitwise CI, not enough for a human to see. So we gave the request a quality tier instead of arguing about the diff.Read it