系列:vLLM & SGLang Serving Notes

(4) The Shared Frontier: Sync Stalls, Speculative Decoding, PD Disaggregation, Low-Bit Quant

1. The One-Line Summary of This Era

In mid-2026, the throughput frontier is no longer “faster kernels” but “stop letting the GPU wait for the host to feed it data”.

Modern matmul is near saturation; paging, continuous batching, and tree speculative decoding are all mature. Where is the remaining throughput? In the gaps where the accelerator idles waiting for the host.

vLLM attacks from the model-runner side (MRv2), SGLang from the speculative-decoding scheduler side (Spec V2). Both dismantle the same synchronization bottleneck from opposite ends.

One bottleneck, attacked from both ends vLLM - MRv2 from the model-runner side async-first, zero CPU-GPU sync overlap step N and N+1 SGLang - Spec V2 from the spec-decode scheduler draft-extend full-graph capture cut D2H / H2D sync Sync Stall time the GPU waits on the host

saves accelerator idle time, not compute; biggest win for low-concurrency latency-sensitive agents

vLLM: MRv2 default (#39337) - PagedAttention removed (#47361) - full-step CUDA Graph (300us to 5us) SGLang: Spec V2 default (+11% TPS) - #31468 DFlash removes per-step host sync - #31487 fewer prefill graph pads - #31986/#31985 DSpark GEMM stacking

2. Speculative Decoding: Squeezing Idle Decode Compute

2.1 Why Decode Must Speculate

Decode is memory-bound: a forward fetches the whole model from HBM to compute just one token, leaving compute units idle.

Speculative decoding’s essence: since you fetch the weights anyway, verify a few extra candidate tokens for free. Same fetch cost, several times the output.

2.2 Draft Methods Compared

MethodDraft sourceTraitsAdoption
Standalone small modelA smaller same-family modelGeneral but needs extra memory and matching vocabBoth; vLLM #38174 supports heterogeneous vocab
EAGLE / EAGLE-3Lightweight draft head reusing target hidden statesHigh accept rate, low overhead, mainstreamFirst-class in both
MTPThe model’s own multi-token headLearned at training time, ~4 tokens per stepBuilt into DeepSeek / GLM / Step
DFlashBlock-diffusion drafting + KV injectionOne parallel pass proposes a whole blockSGLang #31468 removes host sync; vLLM #48524 fixes layer sizing
DSparkSemi-autoregressive drafting + Markov head + confidence schedulingDraft length adapts to confidencevLLM native for DeepSeek V4 Pro; SGLang #31986/#31985 optimize
Real numbers: DeepSeek V4 Pro + DSpark hits ~250 tok/s on 8xB300, 12-42% above MTP (native in vLLM 0.25).
For the difference between DFlash and DSpark, see the sibling series: DFlash vs DSpark head to head.

2.3 Details Worth Noting

3. PD Disaggregation: An Emerging Paradigm

3.1 Why Split

Prefill and decode have opposite characteristics; on the same GPU they hurt each other:

PD disaggregation deploys the two phases on different GPU pools; prefill computes and ships the KV cache to decode nodes.

Router Rust / K8s / Planner Prefill pool compute-bound, sets TTFT GPU GPU GPU TP + context parallel (PCP) higher-compute cards KV transfer NIXL / Mooncake NVLink / RDMA Decode pool memory-bound, sets TPOT GPU GPU GPU GPU large EP + large batch bandwidth and capacity bound Tiered KV cache (4 layers) GPU HBM (fastest) NVLink pool Host DRAM SSD / object (Mooncake)

Fig 2: PD disaggregation. Each phase uses its optimal parallelism and hardware, KV shipped via NIXL / Mooncake. Representative numbers: - NVIDIA Dynamo 1.0 (Planner + 4-tier KV manager): DeepSeek-R1 671B compound 30x, Llama-70B Hopper clean 2x - DeepSeek V3 full disagg stack ~545 tok/s/GPU (north star) - Together CPD long-context B200 +35-40% - Mooncake covers 75% of real requests

3.2 The Players

ProjectRoleTrait
NVIDIA Dynamo 1.0Orchestration on top of vLLM/SGLangPlanner + 4-tier KV manager; DeepSeek-R1 671B compound 30x
llm-d (CNCF)K8s-native PD orchestrationStandard cloud-native path
vLLM V1 Rust routerBuilt-in router25% higher RPS than llm-d, no K8s dependency
SGLang PD routerBuilt-in, multiple routing policiesCancels the paired decode on prefill failure
MooncakeKV storage and transfer backendMulti-tier, covers 75% of real requests
NIXLUnified transport layerBoth converge on it

3.3 Interesting Variants

PD disaggregation is not free: it adds KV network transfer (possibly gigabytes on long context), cross-node scheduling complexity, and a wider failure domain. It pays off only at sufficient scale, with a high share of long prompts and fast interconnect (NVLink / RDMA). For small models on a single 8-GPU box, colocated deployment is fine.

4. Low-Bit Quantization

Decode is memory-bound, so shrinking the weights is the most direct speedup.

FormatBitsNoteStatus
FP88Native since Hopper, tiny precision lossA production default; vLLM #42569 adds FP8 KV cache on SM100 via FA4
INT4 / AWQ / GPTQ4Classic weight quantMature; Step-series Int4 weights not yet supported in vLLM
NVFP44Blackwell-native 4-bit float, better than INT42026 headline; needs FlashInfer (since vLLM 0.26)
MXFP44Micro-scaled 4-bit, AMD-side push (SGLang #28291)Expanding
NVFP4_AWQ4NVFP4 + AWQ calibration (SGLang #31825)Frontier combo

Observation: SGLang expands more aggressively on FP4 (NVFP4_AWQ, marlin_nvfp4 fixes, unified per-token-group quant kernel #30924); vLLM wins on format coverage.

Quantization is bug-prone; real cases:
- vLLM #48330: a mixed-dtype fusion kernel reads NVFP4 with the wrong bit pattern, producing silent !!!!! (fixed in 0.25.1)
- vLLM #48816: GPTQ + speculation MTP weight-loading error
- SGLang #31762: wrong routed_scaling_factor in marlin_nvfp4
Conclusion: when quant + speculation + new hardware stack up, always run full correctness regression, not just throughput.

5. A Side Debate: Should You Leave transformers

A persistent disagreement in the community.

An observation: in the July 22 window, vLLM had zero “leave transformers” PRs — instead it was converging the two backends (#49292 M-RoPE for Qwen3-VL on the Transformers backend, #47298 Ovis2_5 transformers-v5 special tokens). The direction is “native MRv2 + transformers v5 in parallel”, not either/or.

6. Summary: How the Four Battle Lines Relate

                 Sources of throughput / latency
                          |
   +----------+-----------+-----------+----------+
   |          |           |           |          |
kill sync  speculative  PD disagg   low-bit
           decoding                 quant
   |          |           |           |
GPU never  one forward  phases       fewer weight
waits      many tokens  do not       bytes moved
                        interfere
   |          |           |           |
 MRv2      EAGLE/MTP   Dynamo      FP8/NVFP4
 Spec V2   DFlash      NIXL         MXFP4
           DSpark      Mooncake     AWQ

They stack: MRv2’s zero-sync lets speculation be fully graph-captured; PD disaggregation lets decode pools run bigger batches, which makes speculative verification more worthwhile; quant frees memory that holds more KV blocks.

Next: assemble these into a timeline and see exactly what the two July 2026 release waves changed, and in what order.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。