★ Most Worth Your Attention Today
vLLM ships v0.29.0 (594 commits / 277 contributors — MRV2 default + Mamba prefix cache −9~25% TTFT + per-request speculative-decode acceptance metrics) — last week it idled in the candidate, this week it actually puts production-ready new stuff in your hands; SGLang holds at v0.5.19 with a maturing Step-3.7-Flash ecosystem.
Three layers of fact:
- vLLM v0.29.0 ships as a formal release (594 commits / 277 contributors), replacing 8/24’s v0.28.0. Last week (the 0907 entry) it was still in the rc1→rc4 candidate with no formal tag; this week it finally tags — the “candidate idling, stable line carrying advisories” worry from last week is replaced by a formal line.
- The first real performance dividend after MRV2 closed out: Model Runner V2 becomes the default execution path for all models (#53183); batch-sharded sampling (#50465) drops per-step logits VRAM to 1/TP; Mamba prefix cache improves TTFT by 9%–25% (#52789);
--per-request-spec-decode-metrics(#48915) finally lets you quantify speculative-decode acceptance per request. - A reverse signal in breaking changes: FlexOlmo / Olmo3 / Hunyuan V1 / VL move back to the Transformers backend (#53615), and the old
api_serverlaunch command is deprecated in favor ofvllm serve. This runs opposite to the 0901 entry where vLLM “externalized bitsandbytes and pushed transformers out” — this round is “long-tail reuses transformers to save maintenance.”
Actionable conclusion: vLLM users should validate parity on staging before upgrading (594 commits is a wide span, and MRV2 default changes the execution path); last week’s “can it go to production” question gets an answer with the formal release — but the security advisories from the 0901 entry (video-decoder VRAM exhaustion / chat-audio decompression bomb) are best closed by upgrading to 0.29.0 rather than staying on the now-superseded 0.28.0.
Worth saying separately: this week’s contrast is the reverse of last week’s — the 0907 entry was “SGLang ships, vLLM stuck in candidate,” and this week vLLM finally delivers the formal release. But note a subtle signal: vLLM this round shows “moving back to the Transformers backend,” whereas 0901 was “pushing out.” The conclusion is clear — flagship models build their own PyTorch stack, long-tail models reuse transformers, tiered by popularity; there is no single direction.
2. vLLM & SGLang Community Tracking
Version status: vLLM v0.29.0 (formal, 594 commits / 277 contributors) replaces 8/24’s v0.28.0; SGLang v0.5.19 (09-05, 786 PRs / 214 contributors) remains the latest stable, no new tag in the last 72h.
vLLM v0.29.0
New features / architecture evolution:
- Model Runner V2 becomes the default execution path for all models (#53183) — MRV2 closes out, execution path unified.
- batch-sharded sampling (#50465): each card computes only 1/TP of the vocab-shard logits, samples locally, and does one lightweight all-reduce to combine — VRAM drops linearly with TP; large-TP deployments gain throughput, users feel nothing (see the deep-dive).
- Mamba prefix cache improves TTFT by 9%–25% (#52789); new
prefix_cache_retention_intervalCLI. --per-request-spec-decode-metrics(#48915): per-request speculative-decode acceptance rate, finally quantifiable.- New models: Hy4-preview / Qwen3.8-Flash-Next / Kimi K3 NVFP4.
Breaking changes:
| Change | PR | Impact |
|---|---|---|
| FlexOlmo / Olmo3 / Hunyuan V1 / VL move back to Transformers backend | #53615 | Old api_server launch command deprecated → vllm serve; re-validate models depending on these |
Old api_server launch command deprecated | — | Launch scripts must move to vllm serve |
| vLLM Ascend v0.26.0rc1 adds Step-3.5 / 3.7 Flash on Ascend 950 | — | Huawei Ascend Step adaptation catches up |
SGLang v0.5.19
New features:
- beam search via
beam_width(not yet compatible with speculative decoding / PD disaggregation / DP attention / HiCache). - Unified Radix Tree forced default (already flagged in the 0907 entry).
- AMD Lean attention pushes GLM-5.2 disaggregated decode down to 8 ms/token;
--enable-layernorm-sp; W4A8 MoE on Hopper +12% (continuing the 0907 entry).
Notable PRs (9/6–9/9): sgl-router bucket-aware policy domains + native cache indexing (cache-affinity routing, a productionization keystone); ROCm DSA indexer top-k precision fix; diffusion mixed INT8 + Comfy NVFP4 encoder; Qwen3.8 Flash adaptation.
Under the Hood: batch-sharded sampling (#50465)
The first typical performance dividend once vLLM v0.29.0 closed out MRV2.
Mechanism: each card computes only 1/TP of the vocab-shard logits, does local partial sampling, and finishes with one lightweight all-reduce to combine. VRAM drops linearly with TP — the larger the TP, the more saved.
Insight: users feel nothing (sampling results are identical), but large-TP deployments get clear throughput and peak-VRAM improvements. This is exactly the “eat the performance dividend” pattern after MRV2 closes — not a new operator, but removing redundancy from the execution path.
My read: after MRV2 becomes the default, vLLM’s optimization focus shifts from “adding features” to “closing paths” — batch-sharded sampling and the Mamba prefix cache both follow this logic. For the deployment side the takeaway is: upgrading to 0.29.0 is not just new features, it is running on a more VRAM-efficient execution path — the same cards fit a bigger batch.
Standing Topics
PD disaggregation: vLLM Kimi-K3 DCP with DSpark + DCP partial prefix cache; SGLang v0.5.19 DCP defaults to MLA plus 9/6 sgl-router cache-affinity routing (a productionization keystone).
Architecture evolution: state-space / hybrid-model prefix caching and KV migration become the new focus (echoing Agentic workloads); sparse MLA (SM100), Qwen3.8 MTP, W4A8 MoE.
Pure PyTorch vs transformers: a reverse signal this round — vLLM moves FlexOlmo / Olmo3 / Hunyuan back to the Transformers backend (reuse is good enough, saves maintenance); the conclusion flips: flagship models build their own PyTorch, long-tail models reuse transformers, tiered by popularity.
Step adaptation: vLLM is the first-class citizen (stepfun37 image + MTP + NVFP4 4-card) > SGLang (dev image + EAGLE); Ascend 950 already supports Step-3.5 / 3.7 Flash; native multimodal vision / video / speech architecture is taking shape.
3. AI Papers & Industry Hotspots
Highlight (one line)
PaLM-E uses a minimal scheme — unify image / text / robot proprioceptive state into the same token stream fed to a 562B model — to prove embodied intelligence = large model + multimodal input, pinning the humanoid race’s focus on the data efficiency of the brain.
Paper core (PaLM-E / Google Research, 2023, arXiv:2303.03378)
- One-line positioning: project text / image (ViT) / proprioceptive state (MLP) into one sequence and run standard cross-entropy next-token prediction, with output containing both text and discrete action tokens — the foundational work of the VLA / world-model line.
- Core idea ①: multimodal embedding injection —
e_img = ViT(I),e_state = MLP(s)concatenated with text tokens; the training objective is identical to a normal LLM (cross-entropy only, no auxiliary loss, no separate policy network). - Core idea ②: action discretization — continuous joint angles binned into tokens, jointly predicted with text / special tokens in the same vocabulary.
- Core idea ③: multi-task mixed training (VQA + captioning + VLM + VLA); at 562B a positive transfer appears — joint training improves every task with no catastrophic forgetting.
- Impact: VLA success rate rises monotonically with model size (12× scaling, OK-Robot tabletop manipulation 91%); the methodological source of the 2026 Zhiyuan / Unitree brain race.
Operator explainer: NSA Native Sparse Attention (DeepSeek V4)
- Positioning: replace a single dense attention with three parallel trainable branches — compression + selection + sliding window — dropping long-sequence complexity from O(n²) to sparse, with 64k decode 11.6× faster and HBM access down ~11×.
- Mechanism: the compression branch scores block importance → top-N blocks get fine attention (soft, differentiable, gradient flows); a learnable gate fuses the three paths.
- Pitfall: sparsity must be natively trainable — a hard
argmaxselection breaks the gradient.
Performance optimization: Medusa Multi-Head Speculative Decoding (arXiv:2401.10774, ICLR 2024, Meta)
- Positioning: attach multiple MLP heads on the target model’s final hidden state, each predicting the future k-th token in parallel, building a draft tree verified in one pass to accept many tokens — no separate draft model needed.
- Gains: simpler than EAGLE, more VRAM-efficient than a small draft model, 2–3× over autoregressive.
- Embodied link: shorter VLA action-decode latency makes robots more responsive; KV quantization + Medusa let edge-side VLA run on onboard compute.
Industry hotspots (embodied-intelligence companies / chain speed)
- [Embodied] Zhiyuan AGILE2.0 (9/10, today): end-to-end sense-control integration; Lingxi X2 walks on a ball autonomously for the first time → bullish for Zhiyuan’s controlled platform [Shangwei New Material 688585], robot ETF; maps the cerebellum → brain progression.
- [Embodied] Unitree UnifoLM-X2-1.0 (9/7): a world-action model achieves fully autonomous humanoid sparring → Unitree (688836) technical anchor; partners with DeepSeek to fill the brain.
- [Embodied] JD Physical-AI Acceleration Plan (9/9): 3M robots / 1M unmanned vehicles / 100k drones procured over 5 years + 80 industry bases → bullish for actuators (Sanhua 002050 / Top Group 601689), reducers (Lead 688017), robot ETF (159039 / 159530).
- [Embodied] XPeng Robotics: the world’s first high-end general-purpose humanoid production line is live, first IRON unit off the line; first funding round exceeds $900M, a new China embodied single-round record → maps to the A-share supply chain.
- [Embodied] Defiance — the US’s first China humanoid-robot ETF (9/4): top-10 holds Lead / Sanhua / Top / Inovance / CATL → China’s supply chain becomes a globally bet-on object.
⚠️ Industry dynamics do not constitute investment advice.
4. The One-Line Takeaway
vLLM finally delivers v0.29.0 (594 commits, MRV2 default + Mamba prefix cache −9~25% TTFT + per-request quantifiable speculative decode), closing last week’s “stuck in candidate, stable line carrying mines” suspense — but this round also shows the reverse signal of “moving back to the Transformers backend”: flagship in-house, long-tail reuse, tiered by popularity; SGLang holds at v0.5.19 with a maturing Step-3.7-Flash ecosystem; on the research side PaLM-E pins embodied intelligence on the data efficiency of the brain with a unified token, NSA native sparse attention and Medusa multi-head speculative decoding hand you copy-ready engineering homework, while Zhiyuan’s ball-walking, Unitree’s fully autonomous sparring, JD’s 3M-unit procurement, and XPeng’s first production line are each tagging the embodied thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。