1. PagedAttention: OS Virtual Memory for the KV Cache
Naive KV caching preallocates by maximum length and reaches only 20%–40% utilization. vLLM’s first big idea borrows directly from operating systems.
1.1 The Analogy
| Operating system | vLLM |
|---|---|
| Process virtual address space | Logical KV sequence of a request |
| Physical page (4 KB) | KV block (typically 16 tokens) |
| Page table | Block table |
| Demand paging, copy-on-write | On-demand blocks, prefix sharing |
A request’s KV is logically contiguous (token 0, 1, 2, …) but physically scattered anywhere in GPU memory, with a block table doing the mapping.
1.2 Three Direct Wins
- Fragmentation nearly vanishes: waste is capped by the unfilled part of the last block — at most 15 tokens, not thousands.
- Free prefix sharing: system prompts, few-shot examples, conversation history — identical prefixes point to the same physical blocks with reference counting and copy-on-write.
- Cheap parallel sampling: best-of-n from one prompt stores the prompt KV once.
2. Continuous Batching
Paging solves memory; head-of-line blocking still needs fixing.
vLLM schedules at iteration level, not batch level:
After every decode step:
- finished requests return immediately and free their blocks
- queued requests are admitted into the freed slots
- if memory runs short, preempt (swap out or recompute a request's KV)
Batch composition flows dynamically, so the GPU always runs active work instead of waiting for the slowest member. With chunked prefill, long prompts are split into chunks interleaved with decode steps, so a long prompt no longer monopolizes a step.
3. The V1 Architecture
Key design: the API server and EngineCore run in separate processes. CPU-heavy frontend work (tokenization, HTTP parsing, JSON serialization) no longer blocks the GPU scheduling loop; the two sides talk over ZMQ. This is also the foundation that made PD disaggregation natural later.
4. The Turn: Why 0.25 Removed PagedAttention
vLLM v0.25.0 (July 11–12, 2026) did something counterintuitive: it removed PagedAttention (PR #47361) while making Model Runner V2 the default for all dense models (#39337).
It looks like self-sabotage. It is actually an abstraction expiring.
4.1 Why It Could Go
PagedAttention was originally an intermediate layer, because attention kernels of that era only read contiguous KV, so vLLM needed logical-to-physical translation above them.
By 2026, modern attention backends (FlashAttention 3/4, FlashInfer) natively read block tables and gather paged KV inside the kernel. Paging had sunk into the kernel itself.
That made vLLM’s PagedAttention layer pure redundant overhead — an extra Python-side dispatch, an extra reshape, unnecessary synchronization. Removing it keeps paging semantics intact (block tables and KVCacheManager remain) and straightens the execution path.
PagedAttention operator abstraction. The paged KV cache mechanism still exists, now handled directly by attention backend kernels. This is abstraction sinking, not feature regression.
4.2 MRv2: Recording a Whole Step as One Graph
The same release promoted Model Runner V2, targeting the third bottleneck: the GPU waiting on the CPU.
MRv2 is async-first with zero CPU-GPU synchronization:
- every “copy back to CPU and branch” point on the forward path is eliminated, handled on-GPU instead;
- with no host sync points, the entire decode step (including speculative draft and verify) can be captured as one complete CUDA Graph;
- step N and N+1 overlap — the CPU prepares a full step ahead.
The effect: per-token kernel launch overhead collapses from ~300 µs to ~5 µs.
4.3 Costs and Pitfalls
- custom operators and older GPUs fall back to the MRv1 path and see no gain;
- Transformers v4 is deprecated (#40389) — migrate to v5;
- the build now requires C++20;
- legacy partial-prefill flags were removed (#49244).
v0.25.1 (July 14) is a mandatory two-commit patch:
- #48330 mixed-dtype quant fusion guard — fixes NVFP4 models emitting silent garbage. Root cause: FlashInfer’s fused
allreduce + RMSNorm + static-quantkernel assumed matching dtypes; with BF16 activations and FP32 RMSNorm weights it read 4-bit NVFP4 with the wrong bit pattern, corrupting hidden states and producing repeated!!!!!. The fix adds a dtype sentinel: mismatched dtypes take a safe path, matched ones keep the fusion. - #47888 — TorchCodec no longer blocks startup when FFmpeg is missing.
5. Where vLLM Differentiates
Across many release cycles, vLLM’s real moat is breadth:
- Model breadth:
Transformers backend parity(since 0.25 the Transformers backend matches native speed) means any new architecture with an HF implementation can be served at full speed on day zero. - Hardware breadth: CUDA / ROCm (AITER) / Intel XPU (DeepSeek-V4
fuse_index_qSYCL path) / TPU / CPU — the widest of any engine. - Quantization breadth: FP8 / INT4 / AWQ / GPTQ / NVFP4 / MXFP4 / compressed-tensors.
- Ecosystem breadth: the KV Connector interface (Mooncake / NIXL / LMCache), a Rust router, and the two-tier KV TieringManager (PR #42285).
6. Summary
| Mechanism | Solves | Status |
|---|---|---|
| Paged KV + block table | Fragmentation, prefix sharing | Retained; operator layer sunk into attention backends |
| Continuous batching + chunked prefill | Head-of-line blocking | Stable foundation |
| API/Engine process split (V1) | CPU frontend blocking the GPU loop | Stable; basis for PD disaggregation |
| Model Runner V2 | GPU idling on the CPU | Default for all dense models since 0.25 |
| Full CUDA Graph capture | Kernel launch overhead | 300 µs to 5 µs |
| Transformers backend parity | Time-to-serve for new models | Day-zero at full speed |
Next: SGLang, which starts from a completely different question — not “how do we manage memory” but “what does an LLM program actually look like”.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。