系列:vLLM & SGLang Serving Notes

(3) SGLang Internals: RadixAttention and the Program View of Inference

1. A Different Starting Point

vLLM asks “how do we manage memory without waste?” SGLang asks a different question:

What do requests in a real LLM application actually look like relative to each other?

The answer: highly repetitive.

If every request recomputes prefill from scratch, these identical prefixes get recomputed thousands of times.

The name states the stance: Structured Generation Language — it treats an LLM call as a structured program, not a stream of unrelated requests.

2. RadixAttention: Prefix KV Reuse via a Radix Tree

2.1 The Mechanism

vLLM also has prefix caching, but early on it was mostly exact-hash matching. SGLang goes further: it organizes the KV of all active requests into a radix tree.

system prompt + tools 1800 tok, KV stored once Chat A, history 620 Chat B, history 340 Eval batch, stem 900 turn 5 q retry branch turn 3 q option A option B option C No reuse (independent prefill) 6 requests x (1800 + history + q) ~16,000 tokens recomputed RadixAttention (tree sharing) shared spans computed once ~2,000 tokens, lower TTFT and memory

Fig 1: the radix tree folds repeated prefixes into shared paths. Longer shared prefixes and more branches mean bigger gains.

2.2 The Boundary of the Gain

RadixAttention's benefit depends entirely on the prefix-sharing rate of your workload. Multi-turn agents, batch evaluation, and tree-of-thought can cut over 70% of prefill; but if your requests share no common prefix (e.g. summarizing distinct documents), you only pay a little tree-maintenance overhead. Look at your traffic shape first.

SGLang v0.5.16 (July 25, 2026) pushed this further: UnifiedRadixTree became the default prefix cache, unifying previously separate cache paths (plain prefix cache, tiered cache, PD-scenario cache) into a single tree.

3. Structured Output: A Compressed State Machine

The second differentiator is constrained decoding.

When you require strict JSON, a regex match, or an EBNF grammar, the naive approach checks the constraint after each token and masks illegal logits to -inf — on the CPU, yet another per-step sync point.

SGLang uses a compressed finite state machine:

For agent tool calls that return fixed schemas, this is very effective — many structural characters are free.

This is also a bug-prone area. SGLang PR #30747 fixed a crash when PP (pipeline parallelism) and structured output are enabled together (Issue #28424). Root cause: scheduling and constraint-validation logic racing on shared state; the fix aligns constraint validation with micro-batch boundaries.
vLLM addressed the sibling issue too: 0.26 makes a grammar failure a per-request error instead of crashing the engine.

4. Zero-Overhead Spec V2

This is SGLang’s most important 2026 performance work, default since v0.5.15 (July 10), for about +11% end-to-end TPS.

4.1 The Hidden Cost of Speculative Decoding

The algorithm itself (draft model proposes, target model verifies in one parallel pass) is covered in the Speculative Decoding Notes. Here we focus on the engineering waste.

A traditional implementation, every step:

GPU: draft model proposes 4 candidates
GPU: target model verifies in parallel
GPU -> CPU (D2H): copy back "how many accepted"      <- sync point!
CPU: decide next sequence length and KV layout
CPU -> GPU (H2D): copy new metadata back             <- sync point!

Between those copies the GPU is fully idle. And because the accept count is a runtime-dependent value, the whole flow cannot be captured by a CUDA Graph — every step relaunches many small kernels.

4.2 The Fix

Traditional: two syncs per step, no full-graph capture draft x4 verify D2H CPU len/KV H2D draft x4 GPU idle bubble (every step)

Spec V2: draft-extend captured in a CUDA Graph, branching on GPU one graph: draft + verify + metadata next step graph next step graph CPU assembles the next step ahead, GPU has no bubbles, +about 11% TPS

Follow-up PRs: #31468 DFlash removes per-step host sync (CPU leads by a full step) - #31487 fewer prefill CUDA graph pads #31986 stack DSpark dense-draft per-layer ctx KV projections into one GEMM - #31985 fold draft embedding into the draft graph via forward_embed

Three moves:

  1. Make draft-extend CUDA-graph capturable — fixed upper-bound tensor shapes plus GPU masking replace runtime-variable lengths;
  2. Cut D2H / H2D — keep the accept count on the GPU, rewrite branching into GPU-executable form;
  3. Fuse metadata computation — page-table and sequence-length updates enter the graph too.

What is saved is accelerator idle time, not compute — so low-concurrency, latency-sensitive agent workloads benefit most.

4.3 IndexShare MTP

The GLM-5.2 optimizations add another highlight: IndexShare MTP.

MTP (Multi-Token Prediction) is speculative decoding with the model’s own draft head. On DSA (sparse attention) models, the draft step would normally recompute its own top-k sparse indices. IndexShare lets the draft step reuse the top-k indices the target model already computed, cutting draft cost by about 1.9x on long context.

Add TopK-V2 (Lightning-TopK) — a selection algorithm replacing full sorting, tuned for 80k-scale inputs.

Combined result: GLM-5.2 NVFP4 hits 500+ tok/s/user on Blackwell (lmsys blog, July 13).

5. Other Capabilities Worth Knowing

FeatureNoteVersion
Breakable CUDA GraphGraphs can be interrupted; DP attention defaults to breakable prefill graph (#31682)since v0.5.15
MLA context parallel decodingContext parallelism for MLA models (DeepSeek family)v0.5.15
FlashInfer all-to-all MoE routingMoE routing via FlashInfer all-to-allv0.5.15
HPC-Ops attention backendMore attention operator options (#30540)July 22 main
Native web searchBuilt-in retrieval tool callingv0.5.15
Cross-request ViT batchingBatch vision encoding under multimodal concurrency (#24013)July 18 main
Large-scale EPExpert parallelism for MoE modelsongoing
Security note: SGLang 0.5.5–0.5.12 had a multimodal path-traversal vulnerability (GHSA-qwrp-wghp-94q2), fixed in 0.5.15+. Upgrade any older deployment.

6. The Two Personalities

Across many releases, the personalities are clear:

vLLMSGLang
OriginMemory managementProgram structure / prefix reuse
StrengthBreadth of models + hardware + quantDepth in prefix sharing, structured output, frontier throughput
Release styleSteady, dense bug fixes after big changesAggressive, frontier work lands fast on main
Sweet spotGeneral base, multi-hardware, multi-modelAgent / multi-turn / batch eval / large-scale EP
Quant aggressivenessFull-format coverageFaster NVFP4 / MXFP4 expansion

A telling observation from the July 22 window: SGLang’s main branch was all speculative-decoding polish (#31986/#31985), while vLLM was all stability fixes (#48524/#49302/#48843/#49306) — the same theme in different phases.

7. Summary

Next: the shared frontier — why “killing the sync stall” became the mid-2026 storyline, and the emerging PD disaggregation paradigm.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。