★ Most Worth Your Attention Today
vLLM holds its annual major release v0.30.0 (9/21-22, 762 commits / 315 contributors) and cuts v0.30.1rc0 (9/23, adding AMD MI355 dense NVFP4 + MoRI kernel image #58281); SGLang still sits at v0.5.20 (main active, no new release); the most notable domestic inference-ecosystem progress in the window is StepFun’s Step-5-Preview-BF16 weights quietly appearing on HF (TypeSafeAI), with the community already giving vLLM/SGLang --trust-remote-code --reasoning-parser stepfun day-0 commands (⚠️ attribution unconfirmed, numbers self-reported, no independent benchmark). The one to put into capacity planning today is that vLLM also backfills ‘hardware quant parity (AMD NVFP4)’ into the 0.30 line, while the Step side warms up with community weights ahead of official framework adaptation.
This is actually the same sentence as today’s Fast Start (a resident weight-cache daemon turning cold-start from checkpoint reload into a single IPC map), seen from two angles: the framework is on two fronts ‘swapping the hard-to-carry thing for a layer you can reuse’ — weights resident in VRAM, restart only does IPC map; quantization pushes down to NVFP4 / MXFP8 KV / FP4, letting different hardware walk the same low-precision path.
Three layers of fact:
- What the release is: vLLM mainline v0.30.0 remains the recommended PD-disaggregation production upgrade target (carries the CVE-2026-93436 fix); v0.30.1rc0 only adds AMD MI355 dense NVFP4 + MoRI kernel image (
#58281), keeping quant parity with NVIDIA; SGLang has no new tag in 72h, still v0.5.20. - How cold-start is solved (continuing): Fast Start keeps a persistent per-GPU weight-cache daemon — on restart,
--load-format ipc_cachewalks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s and CUDA Graph capture 12s→2s, pushing down the ops cost of canary / self-heal / hot-weight (no new change this window, but still the primitive most worth landing on the v0.30 line). - Domestic inference progress: Step-5-Preview-BF16 weights appear on HF (TypeSafeAI); community gives day-0 commands but attribution / numbers unconfirmed; Step-3.7-Flash remains the production mainstay (vLLM
stepfun37+ MTP > SGLangdev+ EAGLE, NVFP4 4 cards) — official framework adaptation PRs still await Step’s open-source countdown (10/15).
Actionable conclusion: AMD users can take v0.30.1rc0’s NVFP4; production PD keeps anchoring on v0.30.0; to try Step-5-Preview-BF16, validate small-traffic first via the community day-0 commands and don’t treat self-reported numbers as benchmarks; cold-start-sensitive services write Fast Start into the restart SOP.
Worth emphasizing: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, exactly the same source as my OpenInfer / Qwen3-4B DFlash — both push cost to a better operating point via ‘reuse existing capability, decouple lifecycle.’ Step-5-Preview-BF16 surfacing via community weights ahead of official adaptation pushes the 0921 ‘activated params decide the bill’ line one step further: once the 27B-activated 600B sparse MoE is merged on the framework side, the price-performance math has to be recomputed.
2. vLLM & SGLang Community Tracking
Version status: in the window, vLLM mainline v0.30.0 (9/21-22, 762 commits / 315 contributors / 104 new) is the recommended PD-disaggregation production upgrade target; SGLang has no new release in 72h, latest stable still v0.5.20 (9/18, 713 PRs / 237 contributors), main branch remains active (9/25 diffusion SRT prompt enhance, kv-hints envelope, sgl-router restructure, removed HiRadixCache) but no new tag.
vLLM (v0.30.0 · Fast Start + V4.1-Flash day-0 + HiSparse + dual-key watermark → v0.30.1rc0 AMD NVFP4)
Version / security landing:
- Version:
v0.30.0(9/21-22) carries the CVE-2026-93436 fix (PD-disaggregated rejected-request KV metadata not released → memory exhaustion, CVSS 8.7, fixed#55677) — still the recommended upgrade target for PD-disaggregation production;v0.30.1rc0(9/23) only adds AMD MI355 dense NVFP4 + MoRI kernel image (#58281), keeping quant parity with NVIDIA. - New feature (cold-start): Fast Start persistent per-GPU weight-cache daemon — on restart,
--load-format ipc_cachewalks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s and CUDA Graph capture 12s→2s (no new change this window, still the primitive most worth landing on the v0.30 line). - Performance: Gemma 4 multimodal prefix 3.23× E2E (RTX PRO 6000); AMD ROCm packaged W4A16 median TPOT −26.4%; Kimi K3 hybrid batch +5.2~7.7% E2E; RL sampling mask GPU compression recovered to within 2~3% of no-mask.
- Breaking: scale-out now requires explicit
--enable-scale-out; removed GPTQ g_idx; Mamba cache deprecated; YaRN aligned to Transformers (max_model_lenor changes); default audio resampler switched to torchaudio — regression-test configs before upgrading.
SGLang (v0.5.20 · branch-point cache + sampling masks + DSpark-under-PD + diffusion SRT enhance)
Performance / hardware (continuing v0.5.20):
- Performance: unified radix-tree branch-point cache (DSV4-Flash hit rate 43.8%→60.8%, TTFT 1.57s→1.07s≈−32%, #34565);
sampling masks overlapscheduling (Qwen3-8B decode batch1 +17% / batch64 +52%, #36630/36631); DSpark-under-PD (8×B300 256K); SGLang Simulator (CPU TTFT prediction error ~6%, #33824); Responses API storage opt-in (#39122). - main active: 9/25 diffusion SRT prompt enhance, kv-hints envelope, sgl-router restructure, removed HiRadixCache — no upgrade needed yet, watch the v0.5.21 candidate.
Under the Hood: Fast Start (weight-lifecycle decoupling)
- Fast Start: weights live outside the engine in a resident daemon in ‘post-quantized, TP-sharded’ form on GPU VRAM; on restart, only CUDA IPC maps the VRAM pointers to the new engine process, skipping disk→HBM reload and heavy re-quantization — moving cold-start cost from checkpoint loading to a single IPC map, heavily cutting ops cost for canary / self-heal / hot-weight (#54921).
My read: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, the same closing logic as my OpenInfer / DFlash ‘reuse already-computed capability’; this window vLLM also backfills AMD NVFP4 quant parity into the 0.30 line, flattening both cost-down lines — ‘lifecycle decoupling + low-precision path’.
Standing Topics
PD disaggregation: vLLM v0.30.0 lands the CVE-2026-93436 fix + HiSparse, v0.30.1rc0 adds AMD NVFP4; SGLang maintains DSpark-under-PD (#37709) + PD role hot-switch (#28403) + KV checksum (#39500), this period newly adds kv-hints envelope into the request transport layer (#38891) seeding a KV-affinity-routing mechanism; Step has no official PD recipe yet.
Architecture evolution: vLLM 0.30: Fast Start (restart opt) + HiSparse (sparse MLA host KV layer) + dual-key watermark + MRV2 full Graph; SGLang 0.5.20: branch-point cache + sampling mask overlap + DSpark-under-PD + diffusion SRT enhance; quantization pushes down to NVFP4 / MXFP8 KV / FP4.
PyTorch vs transformers: no ‘off transformers’ PR this window; vLLM --model-impl transformers + SGLang Transformers fallback page continue the re-layering ‘model layer converges to transformers, performance layer stays in engine kernel’; still costly, only mega-custom worth it.
Step adaptation: Step-5-Preview-BF16 weights appear on HF (TypeSafeAI); community gives vLLM/SGLang --trust-remote-code --reasoning-parser stepfun day-0 commands (⚠️ attribution unconfirmed, numbers self-reported, no independent benchmark); Step-3.7-Flash remains the production mainstay (vLLM stepfun37 + MTP > SGLang dev + EAGLE, NVFP4 4 cards).
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
MAE ports BERT-style masked modeling to vision (75% high mask ratio + asymmetric encoder-decoder, ViT learns general representations via pixel reconstruction), becoming one of the foundational paradigms for today’s VLA / world-model visual front-end; on the operator side GLA (Gated Linear Attention) drops self-attention to O(N·d²) linear, replacing KV cache with a fixed hidden-state matrix; on the performance side NIXL (NVIDIA Inference XL) gives PD-disaggregation zero-copy KV transport across GPU/CPU/NVMe/network; on the industry side Zhiyuan delivers its 20k-th unit to Chimelong, Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order, Digua Robot’s $400M Series C directly bullish on Horizon.
Paper Core (MAE · He et al., Meta FAIR, CVPR 2022)
- One-line positioning: the foundational work of self-supervised visual pre-training — letting a vision Transformer learn general representations by ‘fill-in-the-blank’ like BERT; the 75% high mask ratio + asymmetric encoder-decoder is its most critical difference from earlier visual self-supervision.
- Core idea: a 224×224 image is split into 16×16 patches (196); randomly mask 75%, feed only the 25% visible patches into a heavy ViT encoder, then a light decoder reconstructs all patch pixels; loss is computed only on masked patches — forcing the encoder to infer global structure from few visible tokens.
- Impact: training FLOPs ~1/3 of a full ViT; ViT-L ImageNet fine-tune reaches 87.8% top-1; dense prediction significantly better than supervised pre-training; inspires DINOv2/v3, SigLIP 2, and becomes a visual-front-end candidate for OpenVLA / π0 / GR00T (one of the sources of the robot ‘eye’).
Operator Deep-Dive: GLA (Gated Linear Attention)
- Positioning: drops self-attention from O(N²·d) quadratic to O(N·d²) linear, replacing the KV cache with a fixed d×d hidden-state matrix — linearizing both VRAM and compute for long sequences.
- In use: Qwen3.5 / Qwen3-Next (Gated DeltaNet, 36 linear + 12 full-attention 3:1), MiniMax Lightning Attention, Mamba/S6 same family.
- Key point:
S_t = S_{t-1} ⊙ g_t + k_t v_t^T,o_t = (S_t q_t) ⊙ z_t; gateg_tcontrols history decay,z_tcontrols output; fits edge VLA long-history low-latency inference (no KV accumulation, fixed state).
Performance Optimization: NIXL KV Transport (NVIDIA Inference XL)
- Positioning: under PD-disaggregation, an efficient transport middleware for KV cache across GPU / CPU / NVMe / network — KV no longer has to live in a single card’s VRAM; the decode side pulls on demand.
- Representative work: NVIDIA Dynamo KV Router + NIXL; vLLM / SGLang community integrating.
- Gain: unified
nixl_agentAPI, zero-copy + RDMA / CUDA IPC, async prefetch overlap — KV transport overhead after TTFT becomes nearly invisible; supports multi-robot cloud-edge collaborative inference (KV reused near the edge).
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] Zhiyuan Robot: its 20k-th unit delivered to Chimelong, first 300+ units resident in the park → bullish for Shangwei New Material 688585.
- [Embodied] Tesla Optimus: Yangtze-delta audit finished, first ~5000-unit order → bullish for Tuopu 601689 / Sanhua 002050 / Leader 688017 / Joyson 600699 / Robot ETF 562500.
- [Embodied] Digua Robot: a Horizon spinoff, closed a $400M Series C → Horizon Robotics-W 09660.HK directly bullish.
- [General] Xiaopeng: second-gen VLA pushed to MONA Ultra SE, 4D spatiotemporal perception + 300% response boost → Xiaopeng Motors-W / Horizon 09660.HK.
- [General] Unitree: Wang Xingxing notes the core bottleneck is now ‘a few-millimeter matching error’; the embodied ChatGPT moment needs 80% of tasks done in 80% of unfamiliar scenarios.
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
vLLM holds its annual major release v0.30.0 (762 commits / 315 contributors, carrying the CVE-2026-93436 fix, the recommended PD-disaggregation production upgrade target) and cuts v0.30.1rc0 (9/23, adding AMD MI355 dense NVFP4 + MoRI kernel image #58281, keeping quant parity with NVIDIA); SGLang has no new tag in 72h, still v0.5.20 (branch-point cache DSV4-Flash hit 43.8%→60.8%, sampling masks overlap Qwen3-8B +17%~+52%, DSpark-under-PD, Simulator TTFT prediction error ~6%), main active polishing diffusion SRT / kv-hints envelope; the most notable domestic inference progress is Step-5-Preview-BF16 weights quietly appearing on HF (community day-0 commands, attribution unconfirmed); on the papers side MAE (75% masking + asymmetric encoder-decoder, ViT-L 87.8% top-1) founds self-supervised visual pre-training and becomes the source of the VLA / world-model eye, operator GLA drops attention to O(N·d²) linear, NIXL gives PD-disaggregation zero-copy KV transport across nodes; on the industry side Zhiyuan delivers its 20k-th unit to Chimelong, Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order, Digua Robot’s $400M Series C directly bullish on Horizon — the embodied ‘smarter brain + real landing’ keeps getting a price tag.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。