★ Most Worth Your Attention Today
vLLM has swung its rhythm back from “releases” to “patches + mainline engineering”: the headline of v0.29.1rc0 (09-12) is the speculative-decode dual-key Gumbel-max watermark (#56122) — production can now stamp draft/verify outputs with an anti-forgery watermark; main also lands EPD three-segment split (Encoder–Prefill–Decode, #56657/#56786) so multimodal vision encoding scales independently. SGLang holds v0.5.19 (09-05) with no new tag, but main is busy: mixed INT8 diffusion embeddings + Comfy NVFP4 encoder (MiniMax-H3 video/audio diffusion smoother), DFLASH on XPU (domestic cards can run DFlash too). All three have no formal release; increments are all on main. The real main line today is NVIDIA’s official co-design guide pinning down the “acceptance length vs draft overhead” hardware constraint.
Three layers of fact:
- vLLM v0.29.1rc0 headline = speculative-decode dual-key watermark (#56122): draft and verify tokens use two independent Gumbel-max sampling keys, adding a verifiable watermark to speculative-decode output without hurting quality → a direct win for production compliance, anti-forgery, and traceability (finance/regulated scenarios can use it immediately).
- EPD three-segment split lands on main (#56657/#56786 + #56546): splits inference into Encoder / Prefill / Decode; vision encoding (multimodal VL) can scale independently; #56546 fixes the encoder media option → the multimodal request’s encoding stage is no longer bound to decode.
- SGLang main two-way reinforcement (diffusion + XPU): [diffusion] mixed INT8 embeddings + Comfy NVFP4 encoder (09-09) make MiniMax-H3 video+audio diffusion land smoother; DFLASH for XPU (09-11) + Rust TreeCore external cache linker (09-10) let domestic XPUs run DFlash drafts too.
Actionable conclusion: the dual-key watermark upgrades speculative decoding from a “speed trick” to an “auditable production capability” — but the speedup itself is still bounded by NVIDIA’s guide: Speedup = E[L]×Ttarget/(Tdraft+Tverify), and acceptance length E[L] is the ceiling. For my own OpenInfer / Qwen3-4B DFlash: parallel drafts save serial latency, but draft quality (E[L]) is the hard ceiling — prioritize benchmarking E[L] at different D rather than blindly widening the draft tree.
Worth saying separately: the lead swings back from “VRAM pooling/tiering” (0913) to “speculative decoding + multimodal / domestic-hardware adaptation” — no new tag, but three things push the inference stack toward production + domestic substitutability: making speculative decode auditable, porting DFlash onto XPU, and splitting multimodal encoding into an independent segment.
2. vLLM & SGLang Community Tracking
Version status: vLLM v0.29.1rc0 (09-12) cut the next patch, stable still v0.29.0 (09-08/09-09); SGLang v0.5.19 (09-05) remains the latest stable with no new tag; Step frozen. Almost all increments this round come from main, with no formal release.
vLLM (main · auditable speculative decode + multimodal split main line)
New features / architecture evolution:
- Speculative-decode dual-key watermark (#56122, v0.29.1rc0 headline): draft/verify take two independent Gumbel-max keys, output gets a verifiable watermark, quality-neutral → production compliance / anti-forgery / traceability ready out of the box.
- AMD/ROCm perf (main 09-14): DeepSeek-V4.1 mHC post-processing folded into the latency pre-projection, DeepSelect TopK into the DSA sparse indexer (#56464), nvfp4_ds_mla extracted from the fused norm+rope (#55538) → DSA cheaper on AMD, cleaner NVFP4 MLA path.
- EPD three-segment split (#56657/#56786): Encoder–Prefill–Decode split, #56546 fixes the encoder media option → multimodal vision encoding can scale independently, no longer bound to decode.
- stable v0.29.0 already has: /v1/messages/render, OffloadingConnector fix (no zero-hit under MTP/EAGLE), FP8 QSA indexer (GB300 decode up to 1.3×).
Breaking changes:
| Change | Impact |
|---|---|
| Dual-key watermark added to speculative-decode output (#56122) | Draft/verify sampling path changes; downstream hash/verification consistency must be re-verified |
| EPD three-segment split (#56657/#56786) | Deployment topology change; re-verify parity after prefill/decode/encoder separation |
| nvfp4_ds_mla path extraction (#55538) | AMD DSA path change; re-verify NVFP4 MLA parity on ROCm |
SGLang (main · diffusion + domestic XPU main line)
New features / major adaptation:
- diffusion reinforcement (09-09): mixed INT8 embeddings + Comfy NVFP4 encoder → MiniMax-H3 video+audio diffusion lands smoother, expanding multimodal generation from “text” to “video+audio”.
- Routing/cache: sgl-router bucket-aware policy domains + native cache indexing (09-06), Rust TreeCore external cache linker (09-10), DFLASH for XPU (09-11) → domestic XPUs can also run DFlash parallel drafts.
- Docs: the site adds four day-0 cookbooks — Qwen3.8-Flash-Next / MiniMax-H3 / Kimi-K3 / Inkling.
Production practice: diffusion + NVFP4 encoder fills SGLang’s gap in “multimodal generation”; DFLASH for XPU extends speculative decode’s parallel-draft capability from NVIDIA to domestic cards, opening another segment of the domestic-substitution chain (main-branch increment, not yet in the v0.5.19 tag).
Under the Hood: Speculative Decode “Acceptance Length vs Draft Overhead” Hardware Constraint (NVIDIA co-design guide 09-02)
NVIDIA pins the speculative-decode speedup to one formula: Speedup = E[L]×Ttarget / (Tdraft + Tverify), with verify token count = 1 + D.
- MTP: needs D serial steps, Tdraft ≈ D×Ltarget (draft cost grows linearly with D).
- DFlash / DSpark: one parallel forward generates D drafts, ≈ 5×Ltarget, the large target model’s overhead negligible.
- Key trap: acceptance length E[L] saturates as D grows (SPEED-Bench proven) — widening the draft tree past a point stops raising (even lowers) speedup, because more draft tokens get rejected.
Insight: for my own OpenInfer / Qwen3-4B DFlash, parallel drafts save serial latency, but E[L] is the hard ceiling — prioritize benchmarking E[L] at different D, pick the D that maximizes “E[L]×Ttarget/(Tdraft+Tverify)”, not blind widening. This constraint shares roots with 0913’s “cache-boundary Pareto tradeoff”: both find the optimal operating point on a tradeoff curve.
My read: today’s main line is “auditable speculative decoding + draft overhead nailed by the formula.” The dual-key watermark solves “dare we ship to production,” NVIDIA’s formula solves “how wide is worth it” — together, speculative decoding finally moves from trick to engineering.
Standing Topics
PD disaggregation: vLLM main has EPD three-segment split + 9 KV connectors (NIXL/Mooncake/LMCache) + llm-d (CNCF); SGLang #37506 unified-pool unifies MHA/MLA/SWA/full+SWA+Mamba into Mooncake PD transfer (rejects speculative PD, needs Mooncake + equal TP + PP=1 + lazy compaction). Horizontal: vLLM leans connector pluginization, SGLang leans routing + native unified tree cache. Mooncake daily trillion tokens, KV hit >90%, batch read <50ms (benchmark).
Architecture evolution: speculative-decode four routes (EAGLE-3 / MTP / DFlash / DSpark) systematized by NVIDIA’s guide; sparse attention HiSparse (vLLM #53781/#56629, SGLang #39337 code owners); FP8/NVFP4 quantization (NVFP4 lm_head still a pit); multimodal diffusion (SGLang MiniMax-H3 / Step Video-Audio).
PyTorch vs transformers: no new “leave transformers” PR this cycle. Stance holds: vLLM v0.29.0 reversed FlexOlmo/Olmo3/Hunyuan back to the transformers backend (#53615), pure PyTorch self-dev only pays off on head models; Step-3.7-Flash debugging needs transformers≥5.0. Horizontal conclusion: self-dev inference stack fits high-frequency / head / extreme optimization, long-tail leverages transformers.
Step adaptation: vLLM stepfun37 (FP8/BF16 MTP k=3; NVFP4 4-card TP4 needs modelopt + FP8 KV cache + async-scheduling); SGLang dev-step-3.7-flash EAGLE multi-layer draft. Pricing ¥1.35 input (miss) / ¥0.27 cache / ¥8.10 output; already on OpenRouter + NVIDIA NIM + local. IPO: completed nearly $2.5B Pre-IPO in 2026-5, HKEX filing still “sprinting.” Horizontal: vLLM first-class citizen > SGLang dev image.
3. AI Papers & Industry Hotspots
Highlight (one line)
UC Berkeley’s DayDreamer puts the world model directly onto a real robot, no simulator — a quadruped learns to walk from scratch in ~1h, adapts ~10min after being pushed, opening a third path of “real-world world-model RL”; while the EAGLE feature-level speculative-decode head (~3×) and S-LoRA multi-LoRA concurrent serving extend today’s “inference speedup” main line from the framework layer down to the operator layer — on the industry side Unitree’s G1 big upgrade, Skild AI’s ARR breaking $100M, and harmonic-reducer bottleneck = moat each tag the embodied “see + act” thesis with a price.
Paper core (DayDreamer · UC Berkeley, CoRL 2022, arXiv:2206.14176)
- One-line positioning: traditional world models depend on simulators and struggle to transfer to real hardware; DayDreamer does world-model RL directly on the real robot, no simulator — a quadruped learns to walk from scratch in ~1 hour, adapts ~10 minutes after being pushed, proving “real-world world-model RL” is viable.
- Core idea: apply the world-model objective (“predict future observations”) directly to online real-robot learning, bypassing the sim-to-real gap; run the “model-predictive-control + RL” loop on the real sensor stream.
- Significance: alongside “train in sim then transfer” and “pure offline policy,” it becomes a third path for real-world embodied learning — critical for embodiments with scarce data and inaccurate sims (legged / dexterous hands). Already written into the AI knowledge base [Embodied Intelligence] and turned into a Zhihu explainer.
Operator explainer: EAGLE (feature-level speculative-decode head)
- Positioning: autoregressive drafting on the LLM’s last-layer hidden states (not token-level), with tree attention verifying multiple branches at once → ~ 3× speedup.
- Why fast: drafts are generated in representation space, reusing the target model’s deep features, more accurate (higher E[L]) than an independent small draft model — corroborating today’s NVIDIA formula that “E[L] is the ceiling.”
- Status: EAGLE-3 is one of NVIDIA’s systematized four speculative-decode routes, alongside MTP / DFlash / DSpark.
Performance optimization: S-LoRA (multi-LoRA concurrent serving)
- Problem: serving thousands of adapters on one base model explodes VRAM from per-request LoRA loading and fragments batching.
- Solution: unified paging + heterogeneous batching — page-manage different LoRA delta weights, schedule uniformly across requests, serving thousands of adapters on one base model with near-zero extra overhead.
- Value: turns “one big model + a pile of small adapters” into a scalable multi-tenant service — the bedrock of LoRA inference industrialization (A/B testing, personalization, multi-task).
Industry hotspots (embodied-intelligence companies / chain speed · pinned)
- [Embodied] Unitree (688836·A-share): G1 big upgrade released (from ¥95k / 25DoF), lowering the consumer-humanoid bar again with maxed-out DOF; as regulators tighten the humanoid listing bar, first-mover positioning value stands out.
- [Embodied] Skild AI: discloses ARR breaking $100M, 60+ paying customers — a “monetize-first” sample of the embodied base model, validating the “robot brain” subscription model.
- [Embodied] Boston Dynamics: delays 2027 IPO — even the leader chooses to defer securitization, the sector entering a valuation cool-down.
- [Regulation]: China tightens humanoid IPO regulation, raising the bar for listing, favoring already-listed / mass-production-leading players (Unitree, Zhiyuan, UBTECH).
- [Chain] XPeng IRON: mass production advancing, world’s first high-end general-humanoid automated line enabled, core-process automation rate >80%.
- [Chain]: harmonic reducers and other core-component bottlenecks = moat, mapping to Leader 688017 / Sanhua 002050 / Top 601689 / robot ETF Huaxia 562500 — reducers, screws, and joint encoders remain the chokepoints of domestication and capacity.
⚠️ Industry dynamics do not constitute investment advice.
4. The One-Line Takeaway
vLLM v0.29.1rc0 (09-12) headline = speculative-decode dual-key Gumbel-max watermark (#56122, production compliance/anti-forgery), and main lands EPD three-segment split (Encoder–Prefill–Decode, #56657/#56786) so multimodal vision encoding scales independently; SGLang holds v0.5.19 but main is busy — mixed INT8 diffusion embeddings + Comfy NVFP4 encoder (MiniMax-H3 video/audio diffusion), DFLASH for XPU (domestic cards can run DFlash too). All three have no new tag; the main line is “auditable speculative decoding + draft overhead nailed by NVIDIA’s formula (Speedup = E[L]×Ttarget/(Tdraft+Tverify), E[L] saturates with D),” with the takeaway for my OpenInfer/Qwen3-4B DFlash being to prioritize benchmarking E[L] at different D. On the research side DayDreamer puts the world model on the real robot (quadruped walks in 1h, adapts in 10min after a push), opening a third path of real-world world-model RL; EAGLE (~3×) and S-LoRA extend the speedup main line down to the operator layer; while Unitree’s G1 big upgrade, Skild’s ARR breaking $100M, and harmonic-reducer bottleneck = moat each tag the embodied “see + act” thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。