0. What This Series Covers
This series reorganizes months of daily tracking of the vLLM / SGLang communities into one progressive storyline:
- Why these two frameworks exist (this post)
- vLLM internals: PagedAttention, continuous batching, the V1 architecture, Model Runner V2
- SGLang internals: RadixAttention, structured output, zero-overhead Spec V2
- Shared frontier battles: speculative decoding, killing sync stalls, PD disaggregation, low-bit quantization
- Version timeline: what the 0.25 / 0.5.15 generation change actually did
- Model support and selection: which engine for which workload
No prior reading required, but every post grounds the concrete facts from those daily reports (version numbers, PR IDs, performance figures) into the right conceptual slot.
1. Code That Runs But Cannot Serve
Almost everyone’s first LLM inference code looks like this:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", device_map="cuda")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
inputs = tok("Explain speculative decoding", return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0]))
This is functionally correct. One user, one request — it gives you the right answer.
But wrap it in an HTTP service and let 50 people hit it at once, and it collapses visibly: OOM, latency spiking into tens of seconds, and GPU utilization stuck somewhere in the teens.
The problem is not the model. It is how you serve it.
2. Three Fatal Bottlenecks
2.1 KV Cache Fragmentation
During autoregressive generation, every new token attends to the Key/Value vectors of all previous tokens. Caching them is the KV cache.
How big? Roughly:
KV memory = 2(K,V) x layers x kv_heads x head_dim x seq_len x batch x dtype_bytes
For a 32-layer, 8-KV-head (GQA), head-dim-128, FP16 8B model, that is about 128 KB per token. An 8K-context request costs ~1 GB. Forty concurrent requests need 40 GB of KV alone.
The naive implementation preallocates a contiguous block sized by max_new_tokens. Set it to 2048 and it reserves 2048 tokens of space — even if the request stops after 30 tokens.
The result is memory full of reserved-but-unused holes. In practice, naive serving achieves only 20%–40% effective KV utilization. Over 60% is pure waste.
2.2 Static Batching: Head-of-Line Blocking
Naive services usually do static batching: collect N requests, run them together, wait for the longest one, return the whole batch, repeat.
Some requests finish in 20 tokens, others need 2000. The short ones do not free their slots — those slots idle until the longest finishes. And new arrivals must wait for the entire batch to complete.
2.3 GPU Waiting on CPU: The Invisible Killer
Within one decode step, the GPU may spend only a few hundred microseconds on matmuls, while the CPU schedules the next batch, prepares tensors, copies metadata, decides sampling results, checks stop conditions — often requiring device-to-host (D2H) and host-to-device (H2D) copies.
At every such synchronization point, the GPU simply waits. With a few hundred microseconds of compute and comparable launch plus sync overhead, roughly half the wall time is idle.
3. What Each Engine Solves
vLLM starts from memory: manage the KV cache like OS virtual memory (paging plus a block table), combine it with continuous batching, and push throughput up. It then grew into a general production base — the widest model coverage, the most hardware backends (CUDA / ROCm / XPU / TPU), the most quantization formats.
SGLang starts from program structure. It observed that real LLM applications — agents, multi-turn chat, batch evaluation, tree-of-thought — have many requests sharing identical prefixes. So it organizes prefix KV into a radix tree for cross-request reuse: RadixAttention. It also pushes aggressively on structured output and frontier throughput work.
4. The Metrics That Matter
Before choosing anything, be clear about the three numbers that matter:
| Metric | Meaning | Who cares |
|---|---|---|
| TTFT (Time To First Token) | Time to the first output token, dominated by prefill | Chat feel, agent first response |
| TPOT / ITL | Interval between subsequent tokens, dominated by decode | Streaming smoothness |
| Throughput | Total tokens per second across the server | Cost per million tokens |
These trade off against each other. Larger batches raise throughput and worsen per-request TPOT; PD disaggregation cuts TTFT but adds KV transfer. Every design decision picks a point inside this triangle.
5. Prefill vs Decode
These two words recur throughout the series. One request has two phases with opposite compute characteristics:
| Prefill | Decode | |
|---|---|---|
| Work | Process the whole prompt at once, compute all KV | Process one new token |
| Parallelism | High (thousands of tokens at once) | Very low (one at a time) |
| Bottleneck | Compute-bound | Memory-bound |
| Determines | TTFT | TPOT |
This difference explains every optimization that follows:
- Why decode needs large batches? It is memory-bound; weights fetched once serve many requests, so bigger batches amortize better.
- Why speculative decoding? Decode emits one token per forward while compute units idle — so verify a few extra candidates for free.
- Why PD disaggregation? Opposite characteristics interfere on the same GPU; a prefill burst stretches decode latency. Separate them and give each its optimal parallelism and hardware.
6. Summary
| Naive problem | Engine solution | Origin |
|---|---|---|
| KV fragmentation | Paged KV cache + block table | vLLM PagedAttention |
| Head-of-line blocking | Continuous batching (iteration-level scheduling) | Both |
| Recomputing identical prefixes | Radix-tree prefix reuse | SGLang RadixAttention |
| GPU waiting on CPU | Zero-sync model runner / full CUDA Graph | vLLM MRv2 / SGLang Spec V2 |
| Idle decode compute | Speculative decoding (MTP / EAGLE / DFlash / DSpark) | Both |
| P/D interference | Prefill-decode disaggregation | Dynamo / Mooncake / NIXL |
Next: inside vLLM — from the virtual-memory analogy of PagedAttention to the seemingly contradictory act of removing PagedAttention in 0.25.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。