★ Most Worth Your Attention Today
SGLang ships v0.5.19 (786 PRs / 214 contributors — the biggest feature drop this week), while vLLM is stuck in the 0.29.0 candidate and its stable line carries two security advisories — who is “more solid” has a clear answer this week.
Three layers of fact:
- SGLang v0.5.19 (09-05) is officially out: edge/Hopper quantization (W4A8 MoE +12%), DeepEP v2 brings decode into CUDA graph (including cross-node), LayerNorm sequence-parallelism trims prefill, and speculative kernels speed up 1.3–1.8× (DSA) / MTP verify-and-commit −45%~63%.
- vLLM is still in the 0.29.0 candidate: rc1 (09-02, CUTLASS MoE padding-routing out-of-bounds fix) → rc4 (TRT-LLM ragged-prefill sync bug), never a formal tag; the stable line is still 08-26’s v0.28.0.
- And vLLM’s stable line carries two security advisories: a video-decoder VRAM-exhaustion bug (PyNvVideoCodec GPU backend eats VRAM even when the server starts in soft-decode with no reservation; unauthenticated calls can crash workers) plus a chat-audio decompression bomb (affects >= 0.24.0).
Actionable conclusion: if you run multimodal vLLM, patch security first (upgrade to the fixed build + enforce auth/quota + reject caller-controlled decode backends + rate-limit multimodal routes separately); whether to jump to SGLang v0.5.19’s new features can wait until 0.29.0 ships — right now vLLM’s “new stuff” is all in the candidate, and should not be on production.
Worth saying separately: this week’s contrast is textbook — SGLang leads on release cadence and feature density, but vLLM’s landmine sits inside the already-shipped stable line. The former is a “wait-and-see,” the latter is a “live mine on your cluster.” Defuse first.
2. vLLM & SGLang Community Tracking
Version status: SGLang v0.5.19 (09-05) officially released; vLLM stable still v0.28.0 (08-26), 0.29.0 candidate (rc1→rc4, no formal tag).
SGLang v0.5.19
New features:
- W4A8 MoE on Hopper: MXFP4 experts + FP8 activations, DeepSeek-V4-Flash output throughput +12%, GSM8K accuracy unchanged (needs FlashInfer 0.6.18).
- DeepEP v2:
--moe-a2a-backend deepep_v2, ElasticBuffer fixed buffer → decode runs CUDA graph (including cross-node). - LayerNorm sequence parallelism:
--enable-layernorm-sp, Qwen3-8B prefill H100 −3.5% / B200 −5.6%, bigger TP saves more. - Speculative kernel speedups: DSA prefill top-k v2, B200 1.3–1.8×; KDA fused accept
SGLANG_OPT_KDA_FUSED_ACCEPT_STATE=1, MTP verify-and-commit −45%~63%, bit-exact. - Unified Radix Tree defaulted (breaking change #35081).
- Beam search: pass
beam_widthto get back n best sequences; not yet compatible with speculative decoding / PD disaggregation / DP attention / HiCache (#31626). - New models: Qwen3.8 (2.4T-A95B), Qwen3.8-27B, dots3.note, Ling-3.0-flash/tiny, Spark2.5, MiniCPM-SALA, Granite 4.2, LongCat-Image-Edit.
vLLM (0.29.0 candidate + security advisories)
0.29.0 candidate progress: rc1 (09-02, CUTLASS MoE padding-routing out-of-bounds fix) → rc4 (TRT-LLM ragged-prefill sync bug). No formal tag; stable remains v0.28.0.
Security advisories (ops must-read):
- Video-decoder VRAM exhaustion: requests routed via
media_io_kwargs.video.video_backendselect the PyNvVideoCodec GPU backend even when the server starts in soft-decode with no VRAM reserved; unauthenticated calls can crash the worker. - chat-audio decompression bomb: affects >= 0.24.0.
- Mitigation: upgrade to the fixed build; enforce auth + quota, reject caller-controlled decode backends, limit request bodies at the reverse proxy, rate-limit multimodal routes separately.
Under the Hood: Lean Attention (AMD)
During long decode batches, put the otherwise-idle CUs to work → on MI355X, GLM-5.2 disaggregated decode drops from 23 ms to 8 ms/output token.
Insight: decode is memory-bound, compute is under-fed; “raise utilization” pays more than “reduce FLOPs.” First judge compute-bound vs memory-bound, then pick your optimization direction.
My read: this week’s main threads (Lean attention / LayerNorm SP / Unified Radix Tree) are essentially all about reclaiming wasted compute and redundant calculation — not computing faster, but not leaving idle hardware idle. This “utilization” line is worth far more long-term attention than any single-operator FLOPs tweak.
Standing Topics
PD disaggregation: DCP’s default MLA backend on Blackwell (trtllm_mla); at 128K, 8×B200 pure TP hits a ~680 tok/s wall, while DCP keeps scaling with concurrency.
Speculative decoding: DSA v2 / KDA fused speedups; beam search not yet compatible.
Quantization: W4A8 MoE on Hopper, DSv4-Flash +12%.
Hardware: AMD is the biggest winner this week (Lean attention + LayerNorm SP) — MI355X decode utilization is finally unlocked.
Step adaptation: no new adaptation PR; step-3.7-flash still served direct via Alibaba Bailian.
3. AI Papers & Industry Hotspots
Highlight (one line)
Google DeepMind combines a “thinking embodied-reasoning model” with a “general VLA” into an agentic system (Gemini Robotics 1.5), advancing the general-robot paradigm a step further; in parallel we dig into DeepSeek V4’s two engineering workhorses — MoE gating and lossless Lookahead decoding.
Paper core (Gemini Robotics 1.5 / GR-ER 1.5)
- One-line positioning: combines the “embodied-reasoning thinking model GR-ER 1.5 (VLM)” with the “general VLA GR 1.5,” letting the robot plan/think first, then act, and zero-shot transfer skills across embodiments.
- Core idea (3 lines): ① Motion Transfer joint-trains heterogeneous data from ALOHA / dual-arm Franka / Apollo humanoid to unify motion representation → zero-shot cross-embodiment transfer; ② Embodied Thinking generates a natural-language thought trajectory h_{1:t-1} before emitting actions, spliced back into context (a_t = f_VLA(o_t, i, h, θ)); ③ GR-ER 1.5 tops SOTA on 15 embodied-reasoning benchmarks.
- Impact: the agentic system’s long-horizon task progress score approaches 80% (Thinking VLA alone only 44%), reinforcing the “brain = VLM reasoning + VLA action” paradigm — bullish for the embodied-intelligence software / data / model layers.
Operator explainer: MoE Gating Affinity Sigmoid → Sqrt(Softplus) + Aux-Loss-Free Expert Bias (DeepSeek V4)
- Positioning: solves routing score and load balance at once.
- Formula: V4 uses s_i = √(softplus(u·e_i)) instead of V3’s σ(u·e_i) (prevents large-logit saturation → vanishing gradients); the aux-loss-free bias b_i is added only to the routing score, with a per-step overload −= γ / underload += γ closed loop (zero gradient); shared experts absorb general common sense.
- In production: DeepSeek V4 / V4-Pro (384 routed + 1 shared, activate 6), Qwen3-MoE, MiniMax; pitfall: the bias buffer must be
requires_grad=Falseand decoupled from the optimizer; Top-K uses the biased score, the gate uses the raw score (keeping the two scores separate is the key).
Performance optimization: Lookahead Decoding (ICML 2024)
- Positioning: resolves the autoregressive serial-latency bottleneck without a draft model or data store.
- Mechanism: treats decoding as a nonlinear equation system solved by Jacobi fixed-point iteration for parallel prediction; a Lookahead branch (2D window W×N collecting n-grams) plus a Verification branch (parallel verify accepting the longest prefix); lossless, decode steps ∝ log(FLOPs/step).
- Gains: single card 1.5–2.3× (MT-bench 1.8×), multi-card Lookahead Parallelism reaches 4×, and it is compatible with FlashAttention for another +20%; embodied link: shortens VLA action-decode latency and needs no draft model, making edge deployment easy.
Industry hotspots (embodied-intelligence companies / chain speed)
- UBTech (09880.HK) | launches U1 10k-unit mass-production delivery on 9/16 (order 13361, full-year >10k units — the industry’s single largest batch) | bullish for whole-machine + actuators/reducers (Sanhua 002050 / Top Group 601689 / Lead 688017).
- AgiBot / Zhiyuan (rushing a HK IPO at 400–500B HKD) | H1 shipments 9,700, ranked #1 globally, industrial orders locked in (Longcheer / Fulin) | maps to Shangwei New Material (688585 · AgiBot-controlled).
- Unitree (688836) | market cap evaporated by over 210B post-listing, the halving is a valuation regression (2025 revenue +335%, deducted non-net profit 600M, not deteriorating) | monitor name flagged for high volatility.
- Mass-production first-year quality: 2026H1 global humanoid shipments 22k (+300%) / China 97%, but real “working” < 20% — valuation shifting from tech Beta to delivery Alpha.
- Model / capital: Gemini Robotics 1.5 released; DeepSeek V4 becomes the inference benchmark; the humanoid track’s H1 financing broke 100B.
4. The One-Line Takeaway
SGLang v0.5.19 delivers, with 786 PRs of density, edge quantization (+12%), DeepEP v2 decode into CUDA graph, and faster speculative kernels into your hands, while vLLM idles in the 0.29.0 candidate and its stable line carries two security advisories — defuse the mines before talking upgrades; on the research side Gemini Robotics 1.5 nails the “thinking VLM + general VLA” agentic paradigm and DeepSeek V4’s Sqrt(Softplus) gating with Lookahead decoding hand you copy-ready engineering homework, while UBTech’s 10k-unit delivery, AgiBot’s HK-IPO sprint, and Unitree’s valuation regression are each tagging the embodied thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。