系列:每日AI热点

Daily AI Hotspot · 2026-09-26: vLLM Holds v0.30.0 and Cuts v0.30.1rc0 (AMD MI355 NVFP4); Step-5-Preview-BF16 Weights Quietly Appear on HF as the Most Notable Domestic Inference Progress; MAE Founds Self-Supervised Visual Front-End

★ Most Worth Your Attention Today

vLLM holds its annual major release v0.30.0 (9/21-22, 762 commits / 315 contributors) and cuts v0.30.1rc0 (9/23, adding AMD MI355 dense NVFP4 + MoRI kernel image #58281); SGLang still sits at v0.5.20 (main active, no new release); the most notable domestic inference-ecosystem progress in the window is StepFun’s Step-5-Preview-BF16 weights quietly appearing on HF (TypeSafeAI), with the community already giving vLLM/SGLang --trust-remote-code --reasoning-parser stepfun day-0 commands (⚠️ attribution unconfirmed, numbers self-reported, no independent benchmark). The one to put into capacity planning today is that vLLM also backfills ‘hardware quant parity (AMD NVFP4)’ into the 0.30 line, while the Step side warms up with community weights ahead of official framework adaptation.

This is actually the same sentence as today’s Fast Start (a resident weight-cache daemon turning cold-start from checkpoint reload into a single IPC map), seen from two angles: the framework is on two fronts ‘swapping the hard-to-carry thing for a layer you can reuse’ — weights resident in VRAM, restart only does IPC map; quantization pushes down to NVFP4 / MXFP8 KV / FP4, letting different hardware walk the same low-precision path.

Three layers of fact:

  1. What the release is: vLLM mainline v0.30.0 remains the recommended PD-disaggregation production upgrade target (carries the CVE-2026-93436 fix); v0.30.1rc0 only adds AMD MI355 dense NVFP4 + MoRI kernel image (#58281), keeping quant parity with NVIDIA; SGLang has no new tag in 72h, still v0.5.20.
  2. How cold-start is solved (continuing): Fast Start keeps a persistent per-GPU weight-cache daemon — on restart, --load-format ipc_cache walks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s and CUDA Graph capture 12s→2s, pushing down the ops cost of canary / self-heal / hot-weight (no new change this window, but still the primitive most worth landing on the v0.30 line).
  3. Domestic inference progress: Step-5-Preview-BF16 weights appear on HF (TypeSafeAI); community gives day-0 commands but attribution / numbers unconfirmed; Step-3.7-Flash remains the production mainstay (vLLM stepfun37 + MTP > SGLang dev + EAGLE, NVFP4 4 cards) — official framework adaptation PRs still await Step’s open-source countdown (10/15).

Actionable conclusion: AMD users can take v0.30.1rc0’s NVFP4; production PD keeps anchoring on v0.30.0; to try Step-5-Preview-BF16, validate small-traffic first via the community day-0 commands and don’t treat self-reported numbers as benchmarks; cold-start-sensitive services write Fast Start into the restart SOP.

Worth emphasizing: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, exactly the same source as my OpenInfer / Qwen3-4B DFlash — both push cost to a better operating point via ‘reuse existing capability, decouple lifecycle.’ Step-5-Preview-BF16 surfacing via community weights ahead of official adaptation pushes the 0921 ‘activated params decide the bill’ line one step further: once the 27B-activated 600B sparse MoE is merged on the framework side, the price-performance math has to be recomputed.

2. vLLM & SGLang Community Tracking

Version status: in the window, vLLM mainline v0.30.0 (9/21-22, 762 commits / 315 contributors / 104 new) is the recommended PD-disaggregation production upgrade target; SGLang has no new release in 72h, latest stable still v0.5.20 (9/18, 713 PRs / 237 contributors), main branch remains active (9/25 diffusion SRT prompt enhance, kv-hints envelope, sgl-router restructure, removed HiRadixCache) but no new tag.

vLLM (v0.30.0 · Fast Start + V4.1-Flash day-0 + HiSparse + dual-key watermark → v0.30.1rc0 AMD NVFP4)

Version / security landing:

SGLang (v0.5.20 · branch-point cache + sampling masks + DSpark-under-PD + diffusion SRT enhance)

Performance / hardware (continuing v0.5.20):

Under the Hood: Fast Start (weight-lifecycle decoupling)

My read: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, the same closing logic as my OpenInfer / DFlash ‘reuse already-computed capability’; this window vLLM also backfills AMD NVFP4 quant parity into the 0.30 line, flattening both cost-down lines — ‘lifecycle decoupling + low-precision path’.

Standing Topics

PD disaggregation: vLLM v0.30.0 lands the CVE-2026-93436 fix + HiSparse, v0.30.1rc0 adds AMD NVFP4; SGLang maintains DSpark-under-PD (#37709) + PD role hot-switch (#28403) + KV checksum (#39500), this period newly adds kv-hints envelope into the request transport layer (#38891) seeding a KV-affinity-routing mechanism; Step has no official PD recipe yet.

Architecture evolution: vLLM 0.30: Fast Start (restart opt) + HiSparse (sparse MLA host KV layer) + dual-key watermark + MRV2 full Graph; SGLang 0.5.20: branch-point cache + sampling mask overlap + DSpark-under-PD + diffusion SRT enhance; quantization pushes down to NVFP4 / MXFP8 KV / FP4.

PyTorch vs transformers: no ‘off transformers’ PR this window; vLLM --model-impl transformers + SGLang Transformers fallback page continue the re-layering ‘model layer converges to transformers, performance layer stays in engine kernel’; still costly, only mega-custom worth it.

Step adaptation: Step-5-Preview-BF16 weights appear on HF (TypeSafeAI); community gives vLLM/SGLang --trust-remote-code --reasoning-parser stepfun day-0 commands (⚠️ attribution unconfirmed, numbers self-reported, no independent benchmark); Step-3.7-Flash remains the production mainstay (vLLM stepfun37 + MTP > SGLang dev + EAGLE, NVFP4 4 cards).

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

MAE ports BERT-style masked modeling to vision (75% high mask ratio + asymmetric encoder-decoder, ViT learns general representations via pixel reconstruction), becoming one of the foundational paradigms for today’s VLA / world-model visual front-end; on the operator side GLA (Gated Linear Attention) drops self-attention to O(N·d²) linear, replacing KV cache with a fixed hidden-state matrix; on the performance side NIXL (NVIDIA Inference XL) gives PD-disaggregation zero-copy KV transport across GPU/CPU/NVMe/network; on the industry side Zhiyuan delivers its 20k-th unit to Chimelong, Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order, Digua Robot’s $400M Series C directly bullish on Horizon.

Paper Core (MAE · He et al., Meta FAIR, CVPR 2022)

Operator Deep-Dive: GLA (Gated Linear Attention)

Performance Optimization: NIXL KV Transport (NVIDIA Inference XL)

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

vLLM holds its annual major release v0.30.0 (762 commits / 315 contributors, carrying the CVE-2026-93436 fix, the recommended PD-disaggregation production upgrade target) and cuts v0.30.1rc0 (9/23, adding AMD MI355 dense NVFP4 + MoRI kernel image #58281, keeping quant parity with NVIDIA); SGLang has no new tag in 72h, still v0.5.20 (branch-point cache DSV4-Flash hit 43.8%→60.8%, sampling masks overlap Qwen3-8B +17%~+52%, DSpark-under-PD, Simulator TTFT prediction error ~6%), main active polishing diffusion SRT / kv-hints envelope; the most notable domestic inference progress is Step-5-Preview-BF16 weights quietly appearing on HF (community day-0 commands, attribution unconfirmed); on the papers side MAE (75% masking + asymmetric encoder-decoder, ViT-L 87.8% top-1) founds self-supervised visual pre-training and becomes the source of the VLA / world-model eye, operator GLA drops attention to O(N·d²) linear, NIXL gives PD-disaggregation zero-copy KV transport across nodes; on the industry side Zhiyuan delivers its 20k-th unit to Chimelong, Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order, Digua Robot’s $400M Series C directly bullish on Horizon — the embodied ‘smarter brain + real landing’ keeps getting a price tag.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。