★ Most Worth Your Attention Today
vLLM shipped its annual major release v0.30.0 on 9/22 (762 commits / 315 contributors) — the only major version this period, landing Fast Start persistent weight-cache daemon, DeepSeek-V4.1-Flash day-0 support, HiSparse host-resident KV spillover and dual-key Gumbel-max watermark together, and officially fixing CVE-2026-93436, making it the recommended upgrade target for PD-disaggregation production (replacing the earlier rc0-only state). The one to put into your capacity planning today is that it closes two long-standing pain points — engine cold-start and PD security — in one release.
This is actually the same sentence as today’s HiSparse (sparse MLA decode spilling KV pages to pinned host memory), seen from two angles: the framework is decoupling ‘weight lifecycle’ and ‘KV storage’ from the engine process, giving cold-start and VRAM pressure each their own fallback layer — both swap ‘the hard-to-carry thing’ for ‘the layer you can reuse.’
Three layers of fact:
- What the release is: v0.30.0 (9/22, 762 commits / 315 contributors) carries the CVE-2026-93436 fix and is the recommended upgrade version for PD-disaggregation production; versus the earlier hanging state of ‘only rc0 carries the fix, production PD can’t sit on v0.29.0’, there is now a formal landing version.
- How cold-start is solved: Fast Start keeps a persistent per-GPU weight-cache daemon — on restart,
--load-format ipc_cachewalks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s, now covering FP4 and multi-node TP; the essence is ‘decoupling weight lifecycle from the engine process’, moving cold-start cost from checkpoint disk→HBM reload to a single IPC map. - Sparse + secure fallback: HiSparse spills sparse MLA decode KV pages to pinned host memory under GPU pressure (top-k miss → per-request GPU hot buffer, enabled via HiSparseConnector, exposing Prometheus counters); dual-key Gumbel-max watermark carries keyed PRF, per-request opt-out, and the dual-key variant is compatible with speculative decoding.
Actionable conclusion: move production PD deployments from ‘wait for 0.29.1+’ to ‘upgrade directly to v0.30.0’; cold-start-sensitive online services put Fast Start into the restart SOP; sparse long-context workloads use HiSparse to treat host memory as a KV fallback layer, first confirming CUDA Graph replay compatibility and Prometheus counter integration.
Worth emphasizing: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, exactly the same source as my OpenInfer / Qwen3-4B DFlash — both push cost to a better operating point via ‘reuse existing capability, decouple lifecycle.’ HiSparse spilling sparse KV to host memory pushes the 0913 ‘VRAM pooling’ idea from inside-VRAM to outside-VRAM; once v0.30.0 is measured, we can see whether Fast Start + HiSparse reproduces the same cold-start / VRAM cost-down on a single card.
2. vLLM & SGLang Community Tracking
Version status: in the window, vLLM v0.30.0 (9/22, 762 commits / 315 contributors) is the only formal major release this period; SGLang has no new release in 72h, latest stable still v0.5.20 (9/18, 713 PRs / 237 contributors), main branch remains active (9/24 AMD MI355X disagg nightly to ROCm 10, etc.) but no new tag.
vLLM (v0.30.0 · Fast Start + V4.1-Flash day-0 + HiSparse + dual-key watermark)
Version / security landing:
- Version:
v0.30.0(9/22) carries the CVE-2026-93436 fix (PD-disaggregated rejected-request KV metadata not released → memory exhaustion, CVSS 8.7, fixed#55677) — becomes the recommended upgrade target for PD-disaggregation production, replacing the earlier ‘only rc0 carries the fix’ hanging state. - New feature (cold-start): Fast Start persistent per-GPU weight-cache daemon — on restart,
--load-format ipc_cachewalks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s, now covering FP4 and multi-node TP; essence is ‘decoupling weight lifecycle from the engine process.’ - New model: DeepSeek-V4.1-Flash day-0 support — full KV stored via FlashMLA V4.1 in MXFP8 on SM100, with DeepGEMM Mega-mHC and async Engram prefetch (Engram DP sharding).
- New feature (sparse / secure): HiSparse host-resident layer — under GPU pressure, sparse MLA decode spills KV pages to pinned host memory, top-k miss walks a per-request GPU hot buffer, enabled via HiSparseConnector and exposing Prometheus counters; dual-key Gumbel-max watermark — keyed PRF for generation+detection, per-request opt-out, dual-key variant compatible with speculative decoding, with the Rust frontend forwarding per-request control.
SGLang (v0.5.20 · branch-point cache + sampling masks + DSpark-under-PD)
Performance / hardware (continuing v0.5.20):
- Performance: unified radix-tree branch-point cache (DSV4-Flash hit rate 43.8%→60.8%, TTFT 1.57s→1.07s),
sampling masks overlapscheduling (Qwen3-8B decode +17%~+52%), DSpark-under-PD, TRT-LLM attention kernel covering DSV4 Blackwell (B200 prefill ≈1.2× / decode ≈1.45×). - Hardware: v0.5.20 retired CUDA12 wheels, Intel XPU into the release image, ROCm loads GLM-5.2 505.7s→40.4s.
Under the Hood: Fast Start (weight-lifecycle decoupling) + HiSparse (sparse KV three-tier spillover)
- Fast Start: weights live outside the engine in a resident daemon in ‘post-quantized, TP-sharded’ form on GPU VRAM; on restart, only CUDA IPC maps the VRAM pointers to the new engine process, skipping disk→HBM reload — moving cold-start cost from checkpoint loading to a single IPC map.
- HiSparse: sparse attention’s top-k is the ‘working-set descriptor’, mapping each selected position to a three-tier resident GPU page / per-request GPU hot buffer / pinned host pool; resident hits return before the hot-cache lookup, all GPU-side and CUDA Graph replay compatible.
My read: these two decouple ‘weights’ and ‘KV’ out of the engine process respectively — Fast Start decouples weights, HiSparse decouples sparse KV — the same closing logic as my OpenInfer / DFlash ‘reuse already-computed capability.’ Next step: test whether ‘Fast Start cold-start + HiSparse host fallback’ reproduces the same cost-down on a single card.
Standing Topics
PD disaggregation: vLLM v0.30.0 is the landing version for the CVE-2026-93436 fix (PD-disaggregation production upgrade target); HiSparse spilling KV to host memory indirectly relieves decode-side VRAM pressure under disagg; SGLang maintains v0.5.20’s DSpark-under-PD (#37709) + PD role hot-switch (#28403) + KV checksum (#39500), main keeps polishing.
Architecture evolution: vLLM 0.30: Fast Start (restart opt), HiSparse (sparse MLA host KV layer), dual-key watermark + speculative decoding, MRV2 graph-capture GC freeze (capture 12s→2s), online acceptance estimator extending adaptive verify to all draft-model spec decoders; SGLang 0.5.20 multi-pronged with sampling masks + branch-point cache + DSV4 Blackwell kernel.
PyTorch vs transformers: no ‘off transformers’ PR this period; boundary signal continues re-layering — vLLM has --model-impl transformers direct HF, SGLang docs add a ‘Transformers fallback’ page, model-definition layer converges to transformers, performance layer stays in engine kernel. Horizontal conclusion: stop treating ‘in-house model implementation’ as a performance prerequisite; only hot operators deserve in-house.
Step adaptation: Step-3.7-Flash deploy path stable (vLLM stepfun37 image + MTP num_speculative_tokens=3; SGLang dev-step-3.7-flash + multi-layer EAGLE; NVFP4 4 cards, FP8 KV); Step 5 Preview (600B/27B, 1M context, native text+vision, AA composite 44 top-3 open) weights to open-source 10/15, adaptation PRs will emerge then, no merge on the framework side now.
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
Meta DINOv3 uses Gram anchoring to preserve dense features, letting a 7B frozen visual backbone beat weak-supervision SOTA for the first time — robots get a label-free default ‘eye’; on the operator side Muon (Newton-Schulz orthogonalization) pushes every direction evenly and MegaBlocks no-drop MoE (block-sparse GEMM) drops capacity_factor, pushing training and inference cost down together; on the industry side Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order.
Paper Core (DINOv3 · Meta FAIR, arXiv:2508.10104)
- One-line positioning: a frozen visual backbone matches or beats SigLIP 2 weak-supervision SOTA across ~60 benchmarks / 15 tasks — a 7B-param, 1.7B-unlabeled-image self-supervised model proving ‘a frozen backbone can be a general visual front-end.’
- Core idea: ① teacher-student self-distillation (EMA teacher) + DINO/iBOT/Koleo unlabeled recipe; ② Gram anchoring — anchor the student’s patch-feature Gram matrix (patch-pairwise similarity structure) to an early checkpoint, constraining only the similarity structure not the absolute values, curing ‘longer training → blurrier local features’; ③ constant lr 1M steps + Gram stage + high-res post-training (extrapolate 4096²+) + multi-student distillation (ConvNeXt small cup, edge-friendly).
- Impact: ADE20k 53.0→63.0 mIoU, ImageNet linear probe 88.2%; usable as a VLA visual front-end (OpenVLA’s DINOv2 can swap to DINOv3), NASA JPL Mars robots already use DINO.
Operator Deep-Dive: Muon Optimizer (Newton-Schulz orthogonalization)
- Positioning: AdamW ignores matrix singular-direction structure; Muon does an orthogonalizing update on momentum so every direction is pushed evenly; Moonlight matches AdamW at about 52% FLOPs.
- In use: Kimi K2 (MuonClip, 1.04T/15.5T tokens zero loss spike), GLM-4.5 (all but embeddings, NS 5 steps/μ=0.95/RMS 0.2), Moonlight.
- Key point: only governs 2D hidden weights; large batch needs QK-Clip / QK-Norm to suppress logit blow-up.
Performance Optimization: MegaBlocks No-Drop MoE (block-sparse GEMM)
- Positioning: solves the MoE ‘capacity drop or padding waste’ dilemma by rewriting the MoE FFN into a block-sparse GEMM that drops
capacity_factor. - Representative work: MegaBlocks (MLSys 2023) + modern grouped GEMM / DeepGEMM Mega MoE (FP8×FP4 single mega-kernel).
- Gain: 40% faster than Tutel’s best config, 2.4× faster than dense Megatron-LM, zero token drop.
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] Tesla Optimus: Tuopu / Sanhua / Joyson upgraded to certified mass-production partners, first ~5000-unit order placed, ~1000 units/week by end of Sept, ~50k units in 2026 → bullish for Tuopu 601689 / Sanhua 002050 / Joyson 600699 / Leader 688017 / Robot ETF 562500.
- [Embodied] financing shift: after Unitree’s IPO institutions paused new whole-robot projects, regulators raised the IPO bar, Mech-Mind 3800× oversubscribed yet still broke issue → neutral-to-cautious on Shangwei 688585 / Unitree 688836, capital rotating to upstream joints / dexterous hands / world models.
- [Embodied] volume & price: 2026 H1 global humanoid shipments 19.1k (+272%), China holds 97%+; Unitree R1 dual-arm from ¥26,900, own average price three years 593k→167k.
- [General]: DeepSeek ARR breaks $1B + rumored Shanghai listing raising ¥50B / ¥500B valuation (API already up 2.3-4.5×); OpenAI Agent first confirmed intrusion into a government system (Australia’s Medicare); Gemini 4 into post-training, Alibaba plans 5-10T-param Qwen.
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
vLLM shipped annual major release v0.30.0 on 9/22 (762 commits / 315 contributors) — Fast Start persistent per-GPU weight-cache daemon (CUDA IPC mapping, H200 engine init 28.9s→8.2s), DeepSeek-V4.1-Flash day-0 support (MXFP8 KV on SM100), HiSparse host-resident KV spillover and dual-key Gumbel-max watermark landed together, and the CVE-2026-93436 fix officially lands, making v0.30.0 the recommended upgrade target for PD-disaggregation production (replacing the earlier rc0-only state); SGLang no new tag in 72h, still v0.5.20 (branch-point cache DSV4-Flash hit 43.8%→60.8%, sampling masks overlap +17%~+52%, DSpark-under-PD, B200 Blackwell 1.2×/1.45×) keeps leading; on the papers side Meta DINOv3 (7B self-supervised, Gram anchoring) is the first frozen visual backbone to beat weak-supervision SOTA, a label-free default eye for robots, operator Muon (Newton-Schulz orthogonalization) and MegaBlocks no-drop MoE (block-sparse GEMM) cut training and inference cost, and on the industry side Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order — the embodied ‘smarter brain + real landing’ keeps getting a price tag.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。