系列:每日AI热点

Daily AI Hotspot · 2026-09-10: vLLM Ships v0.29.0 (MRV2 Default + Mamba Prefix Cache −9~25% TTFT), SGLang Holds at v0.5.19

★ Most Worth Your Attention Today

vLLM ships v0.29.0 (594 commits / 277 contributors — MRV2 default + Mamba prefix cache −9~25% TTFT + per-request speculative-decode acceptance metrics) — last week it idled in the candidate, this week it actually puts production-ready new stuff in your hands; SGLang holds at v0.5.19 with a maturing Step-3.7-Flash ecosystem.

Three layers of fact:

  1. vLLM v0.29.0 ships as a formal release (594 commits / 277 contributors), replacing 8/24’s v0.28.0. Last week (the 0907 entry) it was still in the rc1→rc4 candidate with no formal tag; this week it finally tags — the “candidate idling, stable line carrying advisories” worry from last week is replaced by a formal line.
  2. The first real performance dividend after MRV2 closed out: Model Runner V2 becomes the default execution path for all models (#53183); batch-sharded sampling (#50465) drops per-step logits VRAM to 1/TP; Mamba prefix cache improves TTFT by 9%–25% (#52789); --per-request-spec-decode-metrics (#48915) finally lets you quantify speculative-decode acceptance per request.
  3. A reverse signal in breaking changes: FlexOlmo / Olmo3 / Hunyuan V1 / VL move back to the Transformers backend (#53615), and the old api_server launch command is deprecated in favor of vllm serve. This runs opposite to the 0901 entry where vLLM “externalized bitsandbytes and pushed transformers out” — this round is “long-tail reuses transformers to save maintenance.”

Actionable conclusion: vLLM users should validate parity on staging before upgrading (594 commits is a wide span, and MRV2 default changes the execution path); last week’s “can it go to production” question gets an answer with the formal release — but the security advisories from the 0901 entry (video-decoder VRAM exhaustion / chat-audio decompression bomb) are best closed by upgrading to 0.29.0 rather than staying on the now-superseded 0.28.0.

Worth saying separately: this week’s contrast is the reverse of last week’s — the 0907 entry was “SGLang ships, vLLM stuck in candidate,” and this week vLLM finally delivers the formal release. But note a subtle signal: vLLM this round shows “moving back to the Transformers backend,” whereas 0901 was “pushing out.” The conclusion is clear — flagship models build their own PyTorch stack, long-tail models reuse transformers, tiered by popularity; there is no single direction.

2. vLLM & SGLang Community Tracking

Version status: vLLM v0.29.0 (formal, 594 commits / 277 contributors) replaces 8/24’s v0.28.0; SGLang v0.5.19 (09-05, 786 PRs / 214 contributors) remains the latest stable, no new tag in the last 72h.

vLLM v0.29.0

New features / architecture evolution:

Breaking changes:

ChangePRImpact
FlexOlmo / Olmo3 / Hunyuan V1 / VL move back to Transformers backend#53615Old api_server launch command deprecated → vllm serve; re-validate models depending on these
Old api_server launch command deprecated—Launch scripts must move to vllm serve
vLLM Ascend v0.26.0rc1 adds Step-3.5 / 3.7 Flash on Ascend 950—Huawei Ascend Step adaptation catches up

SGLang v0.5.19

New features:

Notable PRs (9/6–9/9): sgl-router bucket-aware policy domains + native cache indexing (cache-affinity routing, a productionization keystone); ROCm DSA indexer top-k precision fix; diffusion mixed INT8 + Comfy NVFP4 encoder; Qwen3.8 Flash adaptation.

Under the Hood: batch-sharded sampling (#50465)

The first typical performance dividend once vLLM v0.29.0 closed out MRV2.

Mechanism: each card computes only 1/TP of the vocab-shard logits, does local partial sampling, and finishes with one lightweight all-reduce to combine. VRAM drops linearly with TP — the larger the TP, the more saved.

Insight: users feel nothing (sampling results are identical), but large-TP deployments get clear throughput and peak-VRAM improvements. This is exactly the “eat the performance dividend” pattern after MRV2 closes — not a new operator, but removing redundancy from the execution path.

My read: after MRV2 becomes the default, vLLM’s optimization focus shifts from “adding features” to “closing paths” — batch-sharded sampling and the Mamba prefix cache both follow this logic. For the deployment side the takeaway is: upgrading to 0.29.0 is not just new features, it is running on a more VRAM-efficient execution path — the same cards fit a bigger batch.

Standing Topics

PD disaggregation: vLLM Kimi-K3 DCP with DSpark + DCP partial prefix cache; SGLang v0.5.19 DCP defaults to MLA plus 9/6 sgl-router cache-affinity routing (a productionization keystone).

Architecture evolution: state-space / hybrid-model prefix caching and KV migration become the new focus (echoing Agentic workloads); sparse MLA (SM100), Qwen3.8 MTP, W4A8 MoE.

Pure PyTorch vs transformers: a reverse signal this round — vLLM moves FlexOlmo / Olmo3 / Hunyuan back to the Transformers backend (reuse is good enough, saves maintenance); the conclusion flips: flagship models build their own PyTorch, long-tail models reuse transformers, tiered by popularity.

Step adaptation: vLLM is the first-class citizen (stepfun37 image + MTP + NVFP4 4-card) > SGLang (dev image + EAGLE); Ascend 950 already supports Step-3.5 / 3.7 Flash; native multimodal vision / video / speech architecture is taking shape.

3. AI Papers & Industry Hotspots

Highlight (one line)

PaLM-E uses a minimal scheme — unify image / text / robot proprioceptive state into the same token stream fed to a 562B model — to prove embodied intelligence = large model + multimodal input, pinning the humanoid race’s focus on the data efficiency of the brain.

Paper core (PaLM-E / Google Research, 2023, arXiv:2303.03378)

Operator explainer: NSA Native Sparse Attention (DeepSeek V4)

Performance optimization: Medusa Multi-Head Speculative Decoding (arXiv:2401.10774, ICLR 2024, Meta)

Industry hotspots (embodied-intelligence companies / chain speed)

⚠️ Industry dynamics do not constitute investment advice.

4. The One-Line Takeaway

vLLM finally delivers v0.29.0 (594 commits, MRV2 default + Mamba prefix cache −9~25% TTFT + per-request quantifiable speculative decode), closing last week’s “stuck in candidate, stable line carrying mines” suspense — but this round also shows the reverse signal of “moving back to the Transformers backend”: flagship in-house, long-tail reuse, tiered by popularity; SGLang holds at v0.5.19 with a maturing Step-3.7-Flash ecosystem; on the research side PaLM-E pins embodied intelligence on the data efficiency of the brain with a unified token, NSA native sparse attention and Medusa multi-head speculative decoding hand you copy-ready engineering homework, while Zhiyuan’s ball-walking, Unitree’s fully autonomous sparring, JD’s 3M-unit procurement, and XPeng’s first production line are each tagging the embodied thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。