系列:vLLM & SGLang Serving Notes

(2) vLLM Internals: From PagedAttention to Removing PagedAttention

1. PagedAttention: OS Virtual Memory for the KV Cache

Naive KV caching preallocates by maximum length and reaches only 20%–40% utilization. vLLM’s first big idea borrows directly from operating systems.

1.1 The Analogy

Operating systemvLLM
Process virtual address spaceLogical KV sequence of a request
Physical page (4 KB)KV block (typically 16 tokens)
Page tableBlock table
Demand paging, copy-on-writeOn-demand blocks, prefix sharing

A request’s KV is logically contiguous (token 0, 1, 2, …) but physically scattered anywhere in GPU memory, with a block table doing the mapping.

Logical view (what the request sees) tok 0-15 tok 16-31 tok 32-47 tok 48-63 Request A (64 tokens) Block Table [0]->#7 [1]->#2 [2]->#9 [3]->#4

Physical GPU memory (block pool, scattered) #0 #1 B #2 A1 #3 #4 A3 #5 B #6 #7 A0 #8 #9 A2 #10 B #11

Shared prefix: copy-on-write If A and B share a prompt prefix, both point to the same block

Fig 1: logical-to-physical mapping. Waste is bounded by one partially filled block, and identical prefixes share physical blocks. Result: effective KV utilization rises from ~20-40% to >90%, multiplying concurrency at the same memory budget.

1.2 Three Direct Wins

  1. Fragmentation nearly vanishes: waste is capped by the unfilled part of the last block — at most 15 tokens, not thousands.
  2. Free prefix sharing: system prompts, few-shot examples, conversation history — identical prefixes point to the same physical blocks with reference counting and copy-on-write.
  3. Cheap parallel sampling: best-of-n from one prompt stores the prompt KV once.

2. Continuous Batching

Paging solves memory; head-of-line blocking still needs fixing.

vLLM schedules at iteration level, not batch level:

After every decode step:
  - finished requests return immediately and free their blocks
  - queued requests are admitted into the freed slots
  - if memory runs short, preempt (swap out or recompute a request's KV)

Batch composition flows dynamically, so the GPU always runs active work instead of waiting for the slowest member. With chunked prefill, long prompts are split into chunks interleaved with decode steps, so a long prompt no longer monopolizes a step.

3. The V1 Architecture

API Server (OpenAI-compatible frontend process) ZMQ / IPC EngineCore (separate process, busy loop) Scheduler iteration-level, preemption chunked prefill, priority KVCacheManager block pool, block table prefix cache, offload tiers Structured Output grammar / JSON schema streaming parser engine Model Runner V2 (MRv2) async-first, zero CPU-GPU sync, full-step CUDA Graph capture Attention backends FlashAttn / FlashInfer Quant kernels TP / PP / EP

Key design: the API server and EngineCore run in separate processes. CPU-heavy frontend work (tokenization, HTTP parsing, JSON serialization) no longer blocks the GPU scheduling loop; the two sides talk over ZMQ. This is also the foundation that made PD disaggregation natural later.

4. The Turn: Why 0.25 Removed PagedAttention

vLLM v0.25.0 (July 11–12, 2026) did something counterintuitive: it removed PagedAttention (PR #47361) while making Model Runner V2 the default for all dense models (#39337).

It looks like self-sabotage. It is actually an abstraction expiring.

4.1 Why It Could Go

PagedAttention was originally an intermediate layer, because attention kernels of that era only read contiguous KV, so vLLM needed logical-to-physical translation above them.

By 2026, modern attention backends (FlashAttention 3/4, FlashInfer) natively read block tables and gather paged KV inside the kernel. Paging had sunk into the kernel itself.

That made vLLM’s PagedAttention layer pure redundant overhead — an extra Python-side dispatch, an extra reshape, unnecessary synchronization. Removing it keeps paging semantics intact (block tables and KVCacheManager remain) and straightens the execution path.

Two different things: what was removed is the internal PagedAttention operator abstraction. The paged KV cache mechanism still exists, now handled directly by attention backend kernels. This is abstraction sinking, not feature regression.

4.2 MRv2: Recording a Whole Step as One Graph

The same release promoted Model Runner V2, targeting the third bottleneck: the GPU waiting on the CPU.

MRv2 is async-first with zero CPU-GPU synchronization:

The effect: per-token kernel launch overhead collapses from ~300 µs to ~5 µs.

MRv1: a sync point every step CPU prep GPU forward D2H CPU branch H2D CPU prep GPU forward GPU idle window ~ half of each step

MRv2: whole step in one CUDA Graph, CPU runs a step ahead prep N prep N+1 prep N+2

graph N graph N+1 graph N+2 no bubbles

Fig 3: launch overhead 300 µs to 5 µs. Gains are largest at small and medium batch sizes.

4.3 Costs and Pitfalls

0.25 is a wide-reaching change; full regression testing is mandatory:
  • custom operators and older GPUs fall back to the MRv1 path and see no gain;
  • Transformers v4 is deprecated (#40389) — migrate to v5;
  • the build now requires C++20;
  • legacy partial-prefill flags were removed (#49244).

v0.25.1 (July 14) is a mandatory two-commit patch:

The lesson from #48330: fused kernels genuinely save HBM round trips, but the implicit "dtypes always match" assumption can break. A dtype sentinel buys both speed and correctness — the classic tradeoff for aggressive kernel fusion.

5. Where vLLM Differentiates

Across many release cycles, vLLM’s real moat is breadth:

6. Summary

MechanismSolvesStatus
Paged KV + block tableFragmentation, prefix sharingRetained; operator layer sunk into attention backends
Continuous batching + chunked prefillHead-of-line blockingStable foundation
API/Engine process split (V1)CPU frontend blocking the GPU loopStable; basis for PD disaggregation
Model Runner V2GPU idling on the CPUDefault for all dense models since 0.25
Full CUDA Graph captureKernel launch overhead300 µs to 5 µs
Transformers backend parityTime-to-serve for new modelsDay-zero at full speed

Next: SGLang, which starts from a completely different question — not “how do we manage memory” but “what does an LLM program actually look like”.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。