系列:vLLM & SGLang Serving Notes

(1) Why You Need vLLM / SGLang: What Naive Inference Gets Wrong

0. What This Series Covers

This series reorganizes months of daily tracking of the vLLM / SGLang communities into one progressive storyline:

  1. Why these two frameworks exist (this post)
  2. vLLM internals: PagedAttention, continuous batching, the V1 architecture, Model Runner V2
  3. SGLang internals: RadixAttention, structured output, zero-overhead Spec V2
  4. Shared frontier battles: speculative decoding, killing sync stalls, PD disaggregation, low-bit quantization
  5. Version timeline: what the 0.25 / 0.5.15 generation change actually did
  6. Model support and selection: which engine for which workload

No prior reading required, but every post grounds the concrete facts from those daily reports (version numbers, PR IDs, performance figures) into the right conceptual slot.

1. Code That Runs But Cannot Serve

Almost everyone’s first LLM inference code looks like this:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", device_map="cuda")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")

inputs = tok("Explain speculative decoding", return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0]))

This is functionally correct. One user, one request — it gives you the right answer.

But wrap it in an HTTP service and let 50 people hit it at once, and it collapses visibly: OOM, latency spiking into tens of seconds, and GPU utilization stuck somewhere in the teens.

The problem is not the model. It is how you serve it.

2. Three Fatal Bottlenecks

2.1 KV Cache Fragmentation

During autoregressive generation, every new token attends to the Key/Value vectors of all previous tokens. Caching them is the KV cache.

How big? Roughly:

KV memory = 2(K,V) x layers x kv_heads x head_dim x seq_len x batch x dtype_bytes

For a 32-layer, 8-KV-head (GQA), head-dim-128, FP16 8B model, that is about 128 KB per token. An 8K-context request costs ~1 GB. Forty concurrent requests need 40 GB of KV alone.

The naive implementation preallocates a contiguous block sized by max_new_tokens. Set it to 2048 and it reserves 2048 tokens of space — even if the request stops after 30 tokens.

The result is memory full of reserved-but-unused holes. In practice, naive serving achieves only 20%–40% effective KV utilization. Over 60% is pure waste.

Naive: contiguous preallocation used 30 tok wasted (2048 reserved) used 400 tok wasted Request 3: out of memory, queued Effective utilization ~20%-40%

Paged: on-demand blocks Req A: 2 blocks

Req B: 5 blocks Req C: admitted now Utilization > 90%, several times more concurrency

Fig 1: contiguous reservation vs paged allocation. Each small square is a fixed-size block (e.g. 16 tokens).

2.2 Static Batching: Head-of-Line Blocking

Naive services usually do static batching: collect N requests, run them together, wait for the longest one, return the whole batch, repeat.

Some requests finish in 20 tokens, others need 2000. The short ones do not free their slots — those slots idle until the longest finishes. And new arrivals must wait for the entire batch to complete.

2.3 GPU Waiting on CPU: The Invisible Killer

Within one decode step, the GPU may spend only a few hundred microseconds on matmuls, while the CPU schedules the next batch, prepares tensors, copies metadata, decides sampling results, checks stop conditions — often requiring device-to-host (D2H) and host-to-device (H2D) copies.

At every such synchronization point, the GPU simply waits. With a few hundred microseconds of compute and comparable launch plus sync overhead, roughly half the wall time is idle.

In one line: naive inference wastes memory holes, idle slots, and idle GPU cycles. Inference engines exist to fill all three.

3. What Each Engine Solves

Naive HF generate x Contiguous KV, fragmented x Static batch, HOL blocking x Per-step sync, GPU idle x No prefix reuse x No quant / parallelism ~15% utilization vLLM + Paged KV / block table + Continuous batching + MRv2 zero sync + Widest HW / model reach + Full-step CUDA Graph General production base SGLang + RadixAttention prefix tree + Fast structured output + Zero-overhead Spec V2 + Large-scale EP / PD + Fast frontier adoption High-sharing / agent loads

Fig 2: vLLM optimizes for breadth (models and hardware); SGLang for depth (prefix sharing and frontier throughput).

vLLM starts from memory: manage the KV cache like OS virtual memory (paging plus a block table), combine it with continuous batching, and push throughput up. It then grew into a general production base — the widest model coverage, the most hardware backends (CUDA / ROCm / XPU / TPU), the most quantization formats.

SGLang starts from program structure. It observed that real LLM applications — agents, multi-turn chat, batch evaluation, tree-of-thought — have many requests sharing identical prefixes. So it organizes prefix KV into a radix tree for cross-request reuse: RadixAttention. It also pushes aggressively on structured output and frontier throughput work.

4. The Metrics That Matter

Before choosing anything, be clear about the three numbers that matter:

MetricMeaningWho cares
TTFT (Time To First Token)Time to the first output token, dominated by prefillChat feel, agent first response
TPOT / ITLInterval between subsequent tokens, dominated by decodeStreaming smoothness
ThroughputTotal tokens per second across the serverCost per million tokens

These trade off against each other. Larger batches raise throughput and worsen per-request TPOT; PD disaggregation cuts TTFT but adds KV transfer. Every design decision picks a point inside this triangle.

Common mistake: selecting an engine on aggregate throughput alone. Offline batch jobs should indeed maximize total throughput, but online agent workloads are gated by TTFT and P99 TPOT — which live in the small-to-medium batch regime, exactly where the 2026 generation change (vLLM MRv2 / SGLang Spec V2) pays off most.

5. Prefill vs Decode

These two words recur throughout the series. One request has two phases with opposite compute characteristics:

PrefillDecode
WorkProcess the whole prompt at once, compute all KVProcess one new token
ParallelismHigh (thousands of tokens at once)Very low (one at a time)
BottleneckCompute-boundMemory-bound
DeterminesTTFTTPOT

This difference explains every optimization that follows:

6. Summary

Naive problemEngine solutionOrigin
KV fragmentationPaged KV cache + block tablevLLM PagedAttention
Head-of-line blockingContinuous batching (iteration-level scheduling)Both
Recomputing identical prefixesRadix-tree prefix reuseSGLang RadixAttention
GPU waiting on CPUZero-sync model runner / full CUDA GraphvLLM MRv2 / SGLang Spec V2
Idle decode computeSpeculative decoding (MTP / EAGLE / DFlash / DSpark)Both
P/D interferencePrefill-decode disaggregationDynamo / Mooncake / NIXL

Next: inside vLLM — from the virtual-memory analogy of PagedAttention to the seemingly contradictory act of removing PagedAttention in 0.25.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。