系列:每日AI热点

Daily AI Hotspot · 2026-09-15: vLLM v0.29.1rc0 Ships Speculative-Decode Dual-Key Watermark + EPD Three-Segment Split, SGLang main Adds Diffusion & XPU DFlash

★ Most Worth Your Attention Today

vLLM has swung its rhythm back from “releases” to “patches + mainline engineering”: the headline of v0.29.1rc0 (09-12) is the speculative-decode dual-key Gumbel-max watermark (#56122) — production can now stamp draft/verify outputs with an anti-forgery watermark; main also lands EPD three-segment split (Encoder–Prefill–Decode, #56657/#56786) so multimodal vision encoding scales independently. SGLang holds v0.5.19 (09-05) with no new tag, but main is busy: mixed INT8 diffusion embeddings + Comfy NVFP4 encoder (MiniMax-H3 video/audio diffusion smoother), DFLASH on XPU (domestic cards can run DFlash too). All three have no formal release; increments are all on main. The real main line today is NVIDIA’s official co-design guide pinning down the “acceptance length vs draft overhead” hardware constraint.

Three layers of fact:

  1. vLLM v0.29.1rc0 headline = speculative-decode dual-key watermark (#56122): draft and verify tokens use two independent Gumbel-max sampling keys, adding a verifiable watermark to speculative-decode output without hurting quality → a direct win for production compliance, anti-forgery, and traceability (finance/regulated scenarios can use it immediately).
  2. EPD three-segment split lands on main (#56657/#56786 + #56546): splits inference into Encoder / Prefill / Decode; vision encoding (multimodal VL) can scale independently; #56546 fixes the encoder media option → the multimodal request’s encoding stage is no longer bound to decode.
  3. SGLang main two-way reinforcement (diffusion + XPU): [diffusion] mixed INT8 embeddings + Comfy NVFP4 encoder (09-09) make MiniMax-H3 video+audio diffusion land smoother; DFLASH for XPU (09-11) + Rust TreeCore external cache linker (09-10) let domestic XPUs run DFlash drafts too.

Actionable conclusion: the dual-key watermark upgrades speculative decoding from a “speed trick” to an “auditable production capability” — but the speedup itself is still bounded by NVIDIA’s guide: Speedup = E[L]×Ttarget/(Tdraft+Tverify), and acceptance length E[L] is the ceiling. For my own OpenInfer / Qwen3-4B DFlash: parallel drafts save serial latency, but draft quality (E[L]) is the hard ceiling — prioritize benchmarking E[L] at different D rather than blindly widening the draft tree.

Worth saying separately: the lead swings back from “VRAM pooling/tiering” (0913) to “speculative decoding + multimodal / domestic-hardware adaptation” — no new tag, but three things push the inference stack toward production + domestic substitutability: making speculative decode auditable, porting DFlash onto XPU, and splitting multimodal encoding into an independent segment.

2. vLLM & SGLang Community Tracking

Version status: vLLM v0.29.1rc0 (09-12) cut the next patch, stable still v0.29.0 (09-08/09-09); SGLang v0.5.19 (09-05) remains the latest stable with no new tag; Step frozen. Almost all increments this round come from main, with no formal release.

vLLM (main · auditable speculative decode + multimodal split main line)

New features / architecture evolution:

Breaking changes:

ChangeImpact
Dual-key watermark added to speculative-decode output (#56122)Draft/verify sampling path changes; downstream hash/verification consistency must be re-verified
EPD three-segment split (#56657/#56786)Deployment topology change; re-verify parity after prefill/decode/encoder separation
nvfp4_ds_mla path extraction (#55538)AMD DSA path change; re-verify NVFP4 MLA parity on ROCm

SGLang (main · diffusion + domestic XPU main line)

New features / major adaptation:

Production practice: diffusion + NVFP4 encoder fills SGLang’s gap in “multimodal generation”; DFLASH for XPU extends speculative decode’s parallel-draft capability from NVIDIA to domestic cards, opening another segment of the domestic-substitution chain (main-branch increment, not yet in the v0.5.19 tag).

Under the Hood: Speculative Decode “Acceptance Length vs Draft Overhead” Hardware Constraint (NVIDIA co-design guide 09-02)

NVIDIA pins the speculative-decode speedup to one formula: Speedup = E[L]×Ttarget / (Tdraft + Tverify), with verify token count = 1 + D.

Insight: for my own OpenInfer / Qwen3-4B DFlash, parallel drafts save serial latency, but E[L] is the hard ceiling — prioritize benchmarking E[L] at different D, pick the D that maximizes “E[L]×Ttarget/(Tdraft+Tverify)”, not blind widening. This constraint shares roots with 0913’s “cache-boundary Pareto tradeoff”: both find the optimal operating point on a tradeoff curve.

My read: today’s main line is “auditable speculative decoding + draft overhead nailed by the formula.” The dual-key watermark solves “dare we ship to production,” NVIDIA’s formula solves “how wide is worth it” — together, speculative decoding finally moves from trick to engineering.

Standing Topics

PD disaggregation: vLLM main has EPD three-segment split + 9 KV connectors (NIXL/Mooncake/LMCache) + llm-d (CNCF); SGLang #37506 unified-pool unifies MHA/MLA/SWA/full+SWA+Mamba into Mooncake PD transfer (rejects speculative PD, needs Mooncake + equal TP + PP=1 + lazy compaction). Horizontal: vLLM leans connector pluginization, SGLang leans routing + native unified tree cache. Mooncake daily trillion tokens, KV hit >90%, batch read <50ms (benchmark).

Architecture evolution: speculative-decode four routes (EAGLE-3 / MTP / DFlash / DSpark) systematized by NVIDIA’s guide; sparse attention HiSparse (vLLM #53781/#56629, SGLang #39337 code owners); FP8/NVFP4 quantization (NVFP4 lm_head still a pit); multimodal diffusion (SGLang MiniMax-H3 / Step Video-Audio).

PyTorch vs transformers: no new “leave transformers” PR this cycle. Stance holds: vLLM v0.29.0 reversed FlexOlmo/Olmo3/Hunyuan back to the transformers backend (#53615), pure PyTorch self-dev only pays off on head models; Step-3.7-Flash debugging needs transformers≥5.0. Horizontal conclusion: self-dev inference stack fits high-frequency / head / extreme optimization, long-tail leverages transformers.

Step adaptation: vLLM stepfun37 (FP8/BF16 MTP k=3; NVFP4 4-card TP4 needs modelopt + FP8 KV cache + async-scheduling); SGLang dev-step-3.7-flash EAGLE multi-layer draft. Pricing ¥1.35 input (miss) / ¥0.27 cache / ¥8.10 output; already on OpenRouter + NVIDIA NIM + local. IPO: completed nearly $2.5B Pre-IPO in 2026-5, HKEX filing still “sprinting.” Horizontal: vLLM first-class citizen > SGLang dev image.

3. AI Papers & Industry Hotspots

Highlight (one line)

UC Berkeley’s DayDreamer puts the world model directly onto a real robot, no simulator — a quadruped learns to walk from scratch in ~1h, adapts ~10min after being pushed, opening a third path of “real-world world-model RL”; while the EAGLE feature-level speculative-decode head (~3×) and S-LoRA multi-LoRA concurrent serving extend today’s “inference speedup” main line from the framework layer down to the operator layer — on the industry side Unitree’s G1 big upgrade, Skild AI’s ARR breaking $100M, and harmonic-reducer bottleneck = moat each tag the embodied “see + act” thesis with a price.

Paper core (DayDreamer · UC Berkeley, CoRL 2022, arXiv:2206.14176)

Operator explainer: EAGLE (feature-level speculative-decode head)

Performance optimization: S-LoRA (multi-LoRA concurrent serving)

Industry hotspots (embodied-intelligence companies / chain speed · pinned)

⚠️ Industry dynamics do not constitute investment advice.

4. The One-Line Takeaway

vLLM v0.29.1rc0 (09-12) headline = speculative-decode dual-key Gumbel-max watermark (#56122, production compliance/anti-forgery), and main lands EPD three-segment split (Encoder–Prefill–Decode, #56657/#56786) so multimodal vision encoding scales independently; SGLang holds v0.5.19 but main is busy — mixed INT8 diffusion embeddings + Comfy NVFP4 encoder (MiniMax-H3 video/audio diffusion), DFLASH for XPU (domestic cards can run DFlash too). All three have no new tag; the main line is “auditable speculative decoding + draft overhead nailed by NVIDIA’s formula (Speedup = E[L]×Ttarget/(Tdraft+Tverify), E[L] saturates with D),” with the takeaway for my OpenInfer/Qwen3-4B DFlash being to prioritize benchmarking E[L] at different D. On the research side DayDreamer puts the world model on the real robot (quadruped walks in 1h, adapts in 10min after a push), opening a third path of real-world world-model RL; EAGLE (~3×) and S-LoRA extend the speedup main line down to the operator layer; while Unitree’s G1 big upgrade, Skild’s ARR breaking $100M, and harmonic-reducer bottleneck = moat each tag the embodied “see + act” thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。