1. A Different Starting Point
vLLM asks “how do we manage memory without waste?” SGLang asks a different question:
What do requests in a real LLM application actually look like relative to each other?
The answer: highly repetitive.
- Multi-turn chat: turn 5’s prefix is all of turn 4
- Agent / ReAct: the same system prompt plus tool definitions resent every step
- Few-shot batch jobs: thousands of requests sharing one multi-thousand-token exemplar block
- Tree-of-thought / self-consistency: N branches forked from one intermediate node
- Batch evaluation: one question template, only the final options differ
If every request recomputes prefill from scratch, these identical prefixes get recomputed thousands of times.
The name states the stance: Structured Generation Language — it treats an LLM call as a structured program, not a stream of unrelated requests.
2. RadixAttention: Prefix KV Reuse via a Radix Tree
2.1 The Mechanism
vLLM also has prefix caching, but early on it was mostly exact-hash matching. SGLang goes further: it organizes the KV of all active requests into a radix tree.
- each node stores the KV blocks for a token span;
- a new request does longest-prefix matching down the tree, reuses the match, and only prefills the remainder;
- LRU eviction manages capacity, so hot prefixes stay resident;
- branches share naturally — tree-of-thought branches are just sibling subtrees.
2.2 The Boundary of the Gain
SGLang v0.5.16 (July 25, 2026) pushed this further: UnifiedRadixTree became the default prefix cache, unifying previously separate cache paths (plain prefix cache, tiered cache, PD-scenario cache) into a single tree.
3. Structured Output: A Compressed State Machine
The second differentiator is constrained decoding.
When you require strict JSON, a regex match, or an EBNF grammar, the naive approach checks the constraint after each token and masks illegal logits to -inf — on the CPU, yet another per-step sync point.
SGLang uses a compressed finite state machine:
- the grammar is compiled into an FSM ahead of time;
- FSM paths with a unique legal successor (e.g.
{"name"in JSON is necessarily followed by:) can emit multiple tokens at once without going through the model; - mask computation is pushed onto the GPU and overlapped with the forward pass.
For agent tool calls that return fixed schemas, this is very effective — many structural characters are free.
vLLM addressed the sibling issue too: 0.26 makes a grammar failure a per-request error instead of crashing the engine.
4. Zero-Overhead Spec V2
This is SGLang’s most important 2026 performance work, default since v0.5.15 (July 10), for about +11% end-to-end TPS.
4.1 The Hidden Cost of Speculative Decoding
The algorithm itself (draft model proposes, target model verifies in one parallel pass) is covered in the Speculative Decoding Notes. Here we focus on the engineering waste.
A traditional implementation, every step:
GPU: draft model proposes 4 candidates
GPU: target model verifies in parallel
GPU -> CPU (D2H): copy back "how many accepted" <- sync point!
CPU: decide next sequence length and KV layout
CPU -> GPU (H2D): copy new metadata back <- sync point!
Between those copies the GPU is fully idle. And because the accept count is a runtime-dependent value, the whole flow cannot be captured by a CUDA Graph — every step relaunches many small kernels.
4.2 The Fix
Three moves:
- Make draft-extend CUDA-graph capturable — fixed upper-bound tensor shapes plus GPU masking replace runtime-variable lengths;
- Cut D2H / H2D — keep the accept count on the GPU, rewrite branching into GPU-executable form;
- Fuse metadata computation — page-table and sequence-length updates enter the graph too.
What is saved is accelerator idle time, not compute — so low-concurrency, latency-sensitive agent workloads benefit most.
4.3 IndexShare MTP
The GLM-5.2 optimizations add another highlight: IndexShare MTP.
MTP (Multi-Token Prediction) is speculative decoding with the model’s own draft head. On DSA (sparse attention) models, the draft step would normally recompute its own top-k sparse indices. IndexShare lets the draft step reuse the top-k indices the target model already computed, cutting draft cost by about 1.9x on long context.
Add TopK-V2 (Lightning-TopK) — a selection algorithm replacing full sorting, tuned for 80k-scale inputs.
Combined result: GLM-5.2 NVFP4 hits 500+ tok/s/user on Blackwell (lmsys blog, July 13).
5. Other Capabilities Worth Knowing
| Feature | Note | Version |
|---|---|---|
| Breakable CUDA Graph | Graphs can be interrupted; DP attention defaults to breakable prefill graph (#31682) | since v0.5.15 |
| MLA context parallel decoding | Context parallelism for MLA models (DeepSeek family) | v0.5.15 |
| FlashInfer all-to-all MoE routing | MoE routing via FlashInfer all-to-all | v0.5.15 |
| HPC-Ops attention backend | More attention operator options (#30540) | July 22 main |
| Native web search | Built-in retrieval tool calling | v0.5.15 |
| Cross-request ViT batching | Batch vision encoding under multimodal concurrency (#24013) | July 18 main |
| Large-scale EP | Expert parallelism for MoE models | ongoing |
6. The Two Personalities
Across many releases, the personalities are clear:
| vLLM | SGLang | |
|---|---|---|
| Origin | Memory management | Program structure / prefix reuse |
| Strength | Breadth of models + hardware + quant | Depth in prefix sharing, structured output, frontier throughput |
| Release style | Steady, dense bug fixes after big changes | Aggressive, frontier work lands fast on main |
| Sweet spot | General base, multi-hardware, multi-model | Agent / multi-turn / batch eval / large-scale EP |
| Quant aggressiveness | Full-format coverage | Faster NVFP4 / MXFP4 expansion |
A telling observation from the July 22 window: SGLang’s main branch was all speculative-decoding polish (#31986/#31985), while vLLM was all stability fixes (#48524/#49302/#48843/#49306) — the same theme in different phases.
7. Summary
- SGLang’s core insight is that an LLM application is a structured program, from which RadixAttention and compressed-FSM constrained decoding grow;
- zero-overhead Spec V2 drives speculative-decoding scheduling overhead near zero, the headline of v0.5.15 (+11% TPS);
- UnifiedRadixTree (v0.5.16) unifies prefix-cache paths;
- the payoff is workload-dependent: high prefix-sharing loads favor SGLang.
Next: the shared frontier — why “killing the sync stall” became the mid-2026 storyline, and the emerging PD disaggregation paradigm.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。