Skip to content
← All news

News1 min read

SGLang merges 40+ agent-tuned operators, one lifting throughput 71%

More than forty agent-optimized operators are upstream in SGLang. Six of them, with the numbers: −41% TTFT, +71% throughput, 2.32× denoising and a 1.41× VAE decode.

More than 40 operators optimized and merged into SGLang, with another batch of twenty running for GLM5.2. These are the six worth showing the numbers for.

Router long-context tokenization dedup — #28744. Idle TTFT: −29% at 60k, −41% at 125k. Under load at 60k: −34% to −49%.

Qwen3-Next FlashInfer allreduce fusion — #22664. Throughput +71.4%; mean TTFT 456.24 ms → 167.54 ms, about −63.3%.

Cohere2Moe NVFP4 fused-MoE — #27401. Throughput +26% on chat, +21% on summarization; against vLLM, +4.1% and +6.8%.

Kimi Delta Attention CuteDSL prefill — #27488. 1.08× to 1.52× over Triton.

Spectral Progressive Diffusion — #27524. Denoising: FLUX.1 1.63×, FLUX.2 1.77×, Z-Image 2.07×, Wan 2.32×, Qwen-Image 1.6×.

LTX-2 VAE decode, channels-last-3d — #27431. Decode 5.41 s → 3.84 s, about 1.41×; latency about −29.0%.

Why an upstream PR is the test we like most ​

A contest track is scored by a harness we can read. A production serving framework is scored by maintainers who did not ask for the change, do not care that an agent wrote it, and will reject it if it is slower on a shape we did not think about, harder to maintain, or wrong. Every number above survived that.

KDA · KDA-Pilot

Read next

NewsSix KDA pull requests land in SGLang, the best adding 8.7% end to endSince August, SGLang has merged six pull requests carrying kernels our agents wrote, and the MLSys contest kernels are now public with a benchmark anyone can rerun. One upstream merge was reverted the next day.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it NewsMSA indexer: 6.5× faster prefill in production, bitwise identicalKDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, up to 3.3× on long-context decode, and bitwise-identical output throughout.Read it