Skip to content
← All news

News1 min read

One node, one week, nobody watching: 53 first places on SOL Bench

A week of unattended kernel generation on a single 8×B200 node took 53 first-place rankings on SOL Bench — and it had not stopped finding them.

Ongoing

Still generating. The count below is a snapshot, and the ranking will be updated as entries land.

On a single 8×B200 node, over about a week, KDA generated kernels holding 53 first-place rankings across SOL Bench.

The number to notice is not 53, it is one node and a week. Kernel optimization has historically priced in a scarce human: someone who knows the architecture, reads the profiler and has the patience for the twentieth variant. What this run costs is a machine that was already there and nobody watching it.

That changes which problems are worth attacking. A track that would take an engineer three days and probably yield nothing is not worth an engineer's three days; it is trivially worth a slot in a queue.

The generated kernels are being prepared for release, and the ranking will follow.

KDA · github.com/humanfia

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it