系列:每日AI热点

Daily AI Hotspot · 2026-09-07: SGLang Ships v0.5.19 (786 PRs — W4A8 MoE +12% / DeepEP v2 Decode into CUDA Graph), vLLM Stuck in 0.29.0 Candidate with Two Security Advisories on Stable

★ Most Worth Your Attention Today

SGLang ships v0.5.19 (786 PRs / 214 contributors — the biggest feature drop this week), while vLLM is stuck in the 0.29.0 candidate and its stable line carries two security advisories — who is “more solid” has a clear answer this week.

Three layers of fact:

  1. SGLang v0.5.19 (09-05) is officially out: edge/Hopper quantization (W4A8 MoE +12%), DeepEP v2 brings decode into CUDA graph (including cross-node), LayerNorm sequence-parallelism trims prefill, and speculative kernels speed up 1.3–1.8× (DSA) / MTP verify-and-commit −45%~63%.
  2. vLLM is still in the 0.29.0 candidate: rc1 (09-02, CUTLASS MoE padding-routing out-of-bounds fix) → rc4 (TRT-LLM ragged-prefill sync bug), never a formal tag; the stable line is still 08-26’s v0.28.0.
  3. And vLLM’s stable line carries two security advisories: a video-decoder VRAM-exhaustion bug (PyNvVideoCodec GPU backend eats VRAM even when the server starts in soft-decode with no reservation; unauthenticated calls can crash workers) plus a chat-audio decompression bomb (affects >= 0.24.0).

Actionable conclusion: if you run multimodal vLLM, patch security first (upgrade to the fixed build + enforce auth/quota + reject caller-controlled decode backends + rate-limit multimodal routes separately); whether to jump to SGLang v0.5.19’s new features can wait until 0.29.0 ships — right now vLLM’s “new stuff” is all in the candidate, and should not be on production.

Worth saying separately: this week’s contrast is textbook — SGLang leads on release cadence and feature density, but vLLM’s landmine sits inside the already-shipped stable line. The former is a “wait-and-see,” the latter is a “live mine on your cluster.” Defuse first.

2. vLLM & SGLang Community Tracking

Version status: SGLang v0.5.19 (09-05) officially released; vLLM stable still v0.28.0 (08-26), 0.29.0 candidate (rc1→rc4, no formal tag).

SGLang v0.5.19

New features:

vLLM (0.29.0 candidate + security advisories)

0.29.0 candidate progress: rc1 (09-02, CUTLASS MoE padding-routing out-of-bounds fix) → rc4 (TRT-LLM ragged-prefill sync bug). No formal tag; stable remains v0.28.0.

Security advisories (ops must-read):

Under the Hood: Lean Attention (AMD)

During long decode batches, put the otherwise-idle CUs to work → on MI355X, GLM-5.2 disaggregated decode drops from 23 ms to 8 ms/output token.

Insight: decode is memory-bound, compute is under-fed; “raise utilization” pays more than “reduce FLOPs.” First judge compute-bound vs memory-bound, then pick your optimization direction.

My read: this week’s main threads (Lean attention / LayerNorm SP / Unified Radix Tree) are essentially all about reclaiming wasted compute and redundant calculation — not computing faster, but not leaving idle hardware idle. This “utilization” line is worth far more long-term attention than any single-operator FLOPs tweak.

Standing Topics

PD disaggregation: DCP’s default MLA backend on Blackwell (trtllm_mla); at 128K, 8×B200 pure TP hits a ~680 tok/s wall, while DCP keeps scaling with concurrency.

Speculative decoding: DSA v2 / KDA fused speedups; beam search not yet compatible.

Quantization: W4A8 MoE on Hopper, DSv4-Flash +12%.

Hardware: AMD is the biggest winner this week (Lean attention + LayerNorm SP) — MI355X decode utilization is finally unlocked.

Step adaptation: no new adaptation PR; step-3.7-flash still served direct via Alibaba Bailian.

3. AI Papers & Industry Hotspots

Highlight (one line)

Google DeepMind combines a “thinking embodied-reasoning model” with a “general VLA” into an agentic system (Gemini Robotics 1.5), advancing the general-robot paradigm a step further; in parallel we dig into DeepSeek V4’s two engineering workhorses — MoE gating and lossless Lookahead decoding.

Paper core (Gemini Robotics 1.5 / GR-ER 1.5)

Operator explainer: MoE Gating Affinity Sigmoid → Sqrt(Softplus) + Aux-Loss-Free Expert Bias (DeepSeek V4)

Performance optimization: Lookahead Decoding (ICML 2024)

Industry hotspots (embodied-intelligence companies / chain speed)

4. The One-Line Takeaway

SGLang v0.5.19 delivers, with 786 PRs of density, edge quantization (+12%), DeepEP v2 decode into CUDA graph, and faster speculative kernels into your hands, while vLLM idles in the 0.29.0 candidate and its stable line carries two security advisories — defuse the mines before talking upgrades; on the research side Gemini Robotics 1.5 nails the “thinking VLM + general VLA” agentic paradigm and DeepSeek V4’s Sqrt(Softplus) gating with Lookahead decoding hand you copy-ready engineering homework, while UBTech’s 10k-unit delivery, AgiBot’s HK-IPO sprint, and Unitree’s valuation regression are each tagging the embodied thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。