系列:Community Tracker Notes

vLLM & SGLang Community Tracker · 2026-08-07 — KV Reuse Enters "Combinatorial Correctness" (HiCache / Mooncake Principles & Architecture)

★ Most Worth Watching Today

SGLang #30393 “HiCache supports packed / sidecar draft cache” — not a perf optimization, but the disclosure of a silent correctness defect. Any production environment running tiered KV cache (HiCache L2/L3) together with speculative decoding (MTP / EAGLE / DSpark) may, with perfectly normal cache-hit rates, have been silently losing accept length because the draft-side state was never restored — no alert, no error log, just flat throughput. Do today: compare accept_length before/after upgrading to main.

Runners-up: SGLang #28836 (torch 2.13 major bump, schedule ahead) and vLLM #49206 (PRIORITY scheduler silently dropping requests — a correctness fix).

1. Baseline (no new tag from any of the three)

ObjectLatest stableReleasedPrevious
vLLMv0.26.02026-07-27v0.25.1 (07-14)
SGLangv0.5.162026-07-25v0.5.15.post1 (07-14)
StepFunStep-3.7-Flash / 3.5-Flashpush frozen 06-01 / 04-03—

All dynamics this period are from main: vLLM 100+ commits, SGLang 100+ commits (48h hit the single-page cap).

2. The Theme: KV Reuse Enters “Combinatorial Correctness”

For a year, the prefix-cache / KV-offload race was about hit rate — how much recompute saved. This window both frameworks exposed and fixed the same class of problem: cache reuse silently breaks when combined orthogonally with other features.

One line: cache is no longer “a buffer in VRAM” but a storage system that must coexist correctly with speculative decoding, PD disaggregation, and multi-tenancy.

3. 🔬 Deep Dive: What HiCache and Mooncake Are, and How They Evolved

3.1 HiCache — SGLang’s Tiered KV Cache

Positioning: extends KV cache from “a buffer in VRAM” into cross-tier storage — L1 = GPU VRAM, L2 = host memory, L3 = remote storage (object store / distributed KV pool). Lower tiers save more recompute; cross-node prefix reuse means multi-replica deployments need not each recompute long system prompts.

Core structure: prefix matching on a radix tree by token sequence — the判定 is “is the target model’s KV available”. Cache moves between L2/L3 as a unit of “host pool”.

Key evolution this period (#30393): speculative decoding introduces a second state machine; HiCache used to move only the target pool, leaving draft state missing. Two integration paths:

PathTopologyHost-pool layout
PackedStandard NextN MTP/EAGLE: DeepSeek-V3.2/V4, GLM-5.x, MiMo-V2.5; incl. DeepSeek-V4 DSparkDraft KV / indexer / SWA buffers appended as a tail layer of the target host pool, same slots, one HiCache operation
SidecarStandalone EAGLE/EAGLE3, DFlash, non-V4 DSparkBuild standalone DRAFT / DRAFT_INDEXER / DRAFT_SWA pools only for non-empty draft state; index derived from target KV/SWA; attached to the same L2/L3 op

Constraint: the draft cache’s index space must be derivable from the target’s, or one move cannot align both. Packed = “draft layers are just extra tail layers of the target”; Sidecar = structurally independent draft models. Sidecar built only for non-empty draft state — zero overhead when spec-decoding is off.

3.2 Mooncake — Distributed KV Store (vLLM’s main push)

Positioning: Mooncake is the distributed KV store standard in vLLM’s KV connector ecosystem — it hosts KV cache on a separate cluster so multiple vLLM deployments share one prefix cache, enabling PD disaggregation and cross-deployment reuse.

Key evolution this period (#48069 / #44956): multi-tenancy and group semantics.

Deployment impact: running multiple vLLM deployments on one Mooncake cluster (multi-line / multi-env) finally has official isolation — previously only key-prefix hacks. #51067 further switches Docker to Mooncake’s official wheel instead of self-building. Mooncake has gone from “one connector impl” to a standard part both frameworks depend on — a sign of ecosystem convergence.

3.3 Why “Tiered Cache × Speculative Decoding” Silently Loses Performance

A “prefix hit” is boolean in metrics. But in a spec-decoding system, a hit has two qualities:

A partial hit errors nowhere, drops no hit-rate metric — it just silently shortens accept length, ending as “high cache-hit rate, flat throughput”.

What speculative decoding actually caches (the second state machine):

These live in device pools separate from the target KV. When HiCache tiers L2/L3 and moves only the target pool, the draft pool is either unbacked or restored with mismatched slot mapping. Prefix matching judges “computed” by target KV → skip prefill → draft model proposes with empty/mismatched draft KV → mass rejection → accept_length 45 → 12, cache_hit_rate pretty but TPS flat. No log tells you what broke.

Two more versions of the same bug, same day:

Transferable judgment: the feature matrix is too large to exhaustively combo-test (prefix cache × tiered store × spec decode × PD disagg × hybrid arch × heterogeneous TP). These bugs share the shape “module A’s implicit assumption broken by module B, with no assertion to catch it”. For users: don’t assume two features both marked stable are safe together; for every new feature, diff accept_length, cached_tokens, TTFT — not just whether it errors.

4. 🟢 vLLM Highlights (main 08-05 ~ 08-07)

5. 🟠 SGLang Highlights (main 08-05 ~ 08-07)

6. 🇨🇳 StepFun: MTP Still Stuck on an Open PR

vLLM #49642 (still open, +18/−5, two weeks): Step3p5AMultiTokenPredictor builds draft blocks over range(num_hidden_layers, num_hidden_layers + num_nextn_predict_layers), i.e. MTP layer layer_idx == num_hidden_layers, one beyond base decode layers; several layer-indexed config tables in step3p5.py (layer_types / rope_theta / use_rope_layers / partial_rotary_factors / swiglu_limits) are length num_hidden_layers with no bounds check → adding --speculative-config '{"method":"mtp"}' on Step-3.7-Flash raises IndexError at load. Fix: bounds-check every layer-indexed lookup; out-of-range draft layers fall back to standard defaults.

Why it matters: Step-3.7-Flash’s three selling points are “mixed sliding-window attention + MTP self-draft + sparse MoE (196B total / 11B active)” — MTP is exactly the official throughput pitch, yet won’t launch on vLLM main. Same window, vLLM merged Ling 3.0 Flash’s “BF16 + MTP + parser” whole stack to main. The gap isn’t model capability, it’s upstream maintenance investment.

Action: ① run Step-3.7-Flash without MTP for now, or cherry-pick #49642; ② for MTP gains short-term, SGLang (--reasoning-parser step3p5 --tool-call-parser step3p5 + NVFP4) is steadier.

7. ⚖️ Side-by-Side

DimensionvLLMSGLang
KV reuse focusExternal storage ecosystem: Mooncake tenant/group/official wheel, intra-block tail reuseInternal tier coupling: HiCache draft-state + observability
Multimodal speedupSplit physical topology: E/P/D separation kills dup transform, preprocess to GPU (1.9~8.6×)Change execution language: whole vision pipeline pure Rust, worker pool + rid sidecar zero-copy
Attitude to PythonKeep fallback + transformers compat, breadth-firstUnsupported = fail to start, no Python fallback, determinism-first
Dep cadenceFlashAttention → torch stable ABI, gradual decoupletorch 2.11→2.13 one-shot big bump, aggressive
Product boundaryLLM/VLM serving, front-end Rust + scheduling governanceBig expansion into diffusion (FLUX/ERNIE-Image/Ideogram/Z-Image + dp-size)
Step supportOfficial recipe complete, but MTP won’t launch (#49642 open)NVFP4 + trtllm_mha + dual parser usable, community activity frozen 07-06

One-line contrast: both solve the same problem’s two halves — “the path before a request hits the GPU and after cache leaves VRAM”. vLLM re-partitions physical resources (encoder as separate node, preprocess to GPU, KV to external multi-tenant store); SGLang rewrites the execution carrier (pipeline into Rust, draft state into tiered cache, PD transport aligned to grid). Choose: breadth + fast new-model onboarding → vLLM; tail-latency determinism + can absorb torch 2.13 cadence → SGLang.

Sources: GitHub API direct queries of vLLM / SGLang releases, main commits, and named PRs (#30393 / #30545 / #28836 / #48069 / #44956 / #50507 / #49206 / #51045 / #49642, etc.). Perf numbers and line deltas are from official repos / PR descriptions; judgments and selection advice are analysis — test against your own workload.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。