1. The One-Line Summary of This Era
In mid-2026, the throughput frontier is no longer “faster kernels” but “stop letting the GPU wait for the host to feed it data”.
Modern matmul is near saturation; paging, continuous batching, and tree speculative decoding are all mature. Where is the remaining throughput? In the gaps where the accelerator idles waiting for the host.
vLLM attacks from the model-runner side (MRv2), SGLang from the speculative-decoding scheduler side (Spec V2). Both dismantle the same synchronization bottleneck from opposite ends.
2. Speculative Decoding: Squeezing Idle Decode Compute
2.1 Why Decode Must Speculate
Decode is memory-bound: a forward fetches the whole model from HBM to compute just one token, leaving compute units idle.
Speculative decoding’s essence: since you fetch the weights anyway, verify a few extra candidate tokens for free. Same fetch cost, several times the output.
2.2 Draft Methods Compared
| Method | Draft source | Traits | Adoption |
|---|---|---|---|
| Standalone small model | A smaller same-family model | General but needs extra memory and matching vocab | Both; vLLM #38174 supports heterogeneous vocab |
| EAGLE / EAGLE-3 | Lightweight draft head reusing target hidden states | High accept rate, low overhead, mainstream | First-class in both |
| MTP | The model’s own multi-token head | Learned at training time, ~4 tokens per step | Built into DeepSeek / GLM / Step |
| DFlash | Block-diffusion drafting + KV injection | One parallel pass proposes a whole block | SGLang #31468 removes host sync; vLLM #48524 fixes layer sizing |
| DSpark | Semi-autoregressive drafting + Markov head + confidence scheduling | Draft length adapts to confidence | vLLM native for DeepSeek V4 Pro; SGLang #31986/#31985 optimize |
For the difference between DFlash and DSpark, see the sibling series: DFlash vs DSpark head to head.
2.3 Details Worth Noting
- Thinking-budget-aware speculation (vLLM #34668): reasoning models behave differently in the thinking phase vs the answer phase; this PR makes speculation budget-aware, safely speeding up reasoning models.
- Heterogeneous-vocab speculation (vLLM #38174): draft and target models with different vocabularies can be paired, greatly widening usable draft models.
- Quant + speculation pitfalls: vLLM #48816 fixes a GPTQ Qwen3.5 MTP weight-loading error when speculation is on. Teams combining quantization and spec decode should watch for these interaction bugs.
3. PD Disaggregation: An Emerging Paradigm
3.1 Why Split
Prefill and decode have opposite characteristics; on the same GPU they hurt each other:
- a long prompt’s prefill monopolizes a step, stalling all streaming outputs (TPOT spikes);
- decode wants large batches for throughput, prefill wants to finish quickly for low TTFT;
- their optimal parallelism differs (prefill likes TP + context parallel, decode likes large EP + large batch).
PD disaggregation deploys the two phases on different GPU pools; prefill computes and ships the KV cache to decode nodes.
3.2 The Players
| Project | Role | Trait |
|---|---|---|
| NVIDIA Dynamo 1.0 | Orchestration on top of vLLM/SGLang | Planner + 4-tier KV manager; DeepSeek-R1 671B compound 30x |
| llm-d (CNCF) | K8s-native PD orchestration | Standard cloud-native path |
| vLLM V1 Rust router | Built-in router | 25% higher RPS than llm-d, no K8s dependency |
| SGLang PD router | Built-in, multiple routing policies | Cancels the paired decode on prefill failure |
| Mooncake | KV storage and transfer backend | Multi-tier, covers 75% of real requests |
| NIXL | Unified transport layer | Both converge on it |
3.3 Interesting Variants
- PPD (ICML 2026): a critique of classic PD — keep KV on the decode node and only append new tokens. TTFT drops 68% from turn 2 onward with no TPOT regression. Big for agents.
- Together CPD: split prefill again into pre-prefill + prefill, +35–40% on long context on B200.
- vLLM TieringManager (PR #42285): two-tier P/D cache across HBM / DRAM / SSD.
- SGLang PrefillDelayer (#31835): delay prefill negotiation until KV-budget admission, avoiding “computed but decode can’t hold it”.
4. Low-Bit Quantization
Decode is memory-bound, so shrinking the weights is the most direct speedup.
| Format | Bits | Note | Status |
|---|---|---|---|
| FP8 | 8 | Native since Hopper, tiny precision loss | A production default; vLLM #42569 adds FP8 KV cache on SM100 via FA4 |
| INT4 / AWQ / GPTQ | 4 | Classic weight quant | Mature; Step-series Int4 weights not yet supported in vLLM |
| NVFP4 | 4 | Blackwell-native 4-bit float, better than INT4 | 2026 headline; needs FlashInfer (since vLLM 0.26) |
| MXFP4 | 4 | Micro-scaled 4-bit, AMD-side push (SGLang #28291) | Expanding |
| NVFP4_AWQ | 4 | NVFP4 + AWQ calibration (SGLang #31825) | Frontier combo |
Observation: SGLang expands more aggressively on FP4 (NVFP4_AWQ, marlin_nvfp4 fixes, unified per-token-group quant kernel #30924); vLLM wins on format coverage.
- vLLM #48330: a mixed-dtype fusion kernel reads NVFP4 with the wrong bit pattern, producing silent
!!!!! (fixed in 0.25.1)- vLLM #48816: GPTQ + speculation MTP weight-loading error
- SGLang #31762: wrong routed_scaling_factor in marlin_nvfp4
Conclusion: when quant + speculation + new hardware stack up, always run full correctness regression, not just throughput.
5. A Side Debate: Should You Leave transformers
A persistent disagreement in the community.
- vLLM’s choice: dual track. Native for performance, the Transformers backend for coverage — and since 0.25 the Transformers backend matches native speed, so new models are day-zero full-speed. The cost is following transformers v5 (v4 deprecated #40389).
- Model authors’ choice: more Chinese models ship
trust_remote_codecustom modeling (e.g. Step 3.7 Flash, where transformers 5.0+ is only for debug). - Tradeoff: custom pure-PyTorch modeling gives controllable operators, low overhead, and easy attention/sampling customization; the cost is implementing your own tokenizer, weight loading, and parallelism.
An observation: in the July 22 window, vLLM had zero “leave transformers” PRs — instead it was converging the two backends (#49292 M-RoPE for Qwen3-VL on the Transformers backend, #47298 Ovis2_5 transformers-v5 special tokens). The direction is “native MRv2 + transformers v5 in parallel”, not either/or.
6. Summary: How the Four Battle Lines Relate
Sources of throughput / latency
|
+----------+-----------+-----------+----------+
| | | | |
kill sync speculative PD disagg low-bit
decoding quant
| | | |
GPU never one forward phases fewer weight
waits many tokens do not bytes moved
interfere
| | | |
MRv2 EAGLE/MTP Dynamo FP8/NVFP4
Spec V2 DFlash NIXL MXFP4
DSpark Mooncake AWQ
They stack: MRv2’s zero-sync lets speculation be fully graph-captured; PD disaggregation lets decode pools run bigger batches, which makes speculative verification more worthwhile; quant frees memory that holds more KV blocks.
Next: assemble these into a timeline and see exactly what the two July 2026 release waves changed, and in what order.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。