系列:每日AI热点

Daily AI Hotspot · 2026-09-25: vLLM v0.30.0 Annual Major Release (762 commits, Fast Start Persistent Weight Cache + HiSparse Host KV Spillover, CVE-2026-93436 Fix Landed); DINOv3 7B Frozen Visual Backbone Beats Weak-Supervision SOTA for the First Time

★ Most Worth Your Attention Today

vLLM shipped its annual major release v0.30.0 on 9/22 (762 commits / 315 contributors) — the only major version this period, landing Fast Start persistent weight-cache daemon, DeepSeek-V4.1-Flash day-0 support, HiSparse host-resident KV spillover and dual-key Gumbel-max watermark together, and officially fixing CVE-2026-93436, making it the recommended upgrade target for PD-disaggregation production (replacing the earlier rc0-only state). The one to put into your capacity planning today is that it closes two long-standing pain points — engine cold-start and PD security — in one release.

This is actually the same sentence as today’s HiSparse (sparse MLA decode spilling KV pages to pinned host memory), seen from two angles: the framework is decoupling ‘weight lifecycle’ and ‘KV storage’ from the engine process, giving cold-start and VRAM pressure each their own fallback layer — both swap ‘the hard-to-carry thing’ for ‘the layer you can reuse.’

Three layers of fact:

  1. What the release is: v0.30.0 (9/22, 762 commits / 315 contributors) carries the CVE-2026-93436 fix and is the recommended upgrade version for PD-disaggregation production; versus the earlier hanging state of ‘only rc0 carries the fix, production PD can’t sit on v0.29.0’, there is now a formal landing version.
  2. How cold-start is solved: Fast Start keeps a persistent per-GPU weight-cache daemon — on restart, --load-format ipc_cache walks CUDA IPC to map TP-sharded weights, cutting H200 engine init 28.9s→8.2s, now covering FP4 and multi-node TP; the essence is ‘decoupling weight lifecycle from the engine process’, moving cold-start cost from checkpoint disk→HBM reload to a single IPC map.
  3. Sparse + secure fallback: HiSparse spills sparse MLA decode KV pages to pinned host memory under GPU pressure (top-k miss → per-request GPU hot buffer, enabled via HiSparseConnector, exposing Prometheus counters); dual-key Gumbel-max watermark carries keyed PRF, per-request opt-out, and the dual-key variant is compatible with speculative decoding.

Actionable conclusion: move production PD deployments from ‘wait for 0.29.1+’ to ‘upgrade directly to v0.30.0’; cold-start-sensitive online services put Fast Start into the restart SOP; sparse long-context workloads use HiSparse to treat host memory as a KV fallback layer, first confirming CUDA Graph replay compatibility and Prometheus counter integration.

Worth emphasizing: Fast Start makes ‘weights resident in VRAM, restart only does IPC map’ a primitive, exactly the same source as my OpenInfer / Qwen3-4B DFlash — both push cost to a better operating point via ‘reuse existing capability, decouple lifecycle.’ HiSparse spilling sparse KV to host memory pushes the 0913 ‘VRAM pooling’ idea from inside-VRAM to outside-VRAM; once v0.30.0 is measured, we can see whether Fast Start + HiSparse reproduces the same cold-start / VRAM cost-down on a single card.

2. vLLM & SGLang Community Tracking

Version status: in the window, vLLM v0.30.0 (9/22, 762 commits / 315 contributors) is the only formal major release this period; SGLang has no new release in 72h, latest stable still v0.5.20 (9/18, 713 PRs / 237 contributors), main branch remains active (9/24 AMD MI355X disagg nightly to ROCm 10, etc.) but no new tag.

vLLM (v0.30.0 · Fast Start + V4.1-Flash day-0 + HiSparse + dual-key watermark)

Version / security landing:

SGLang (v0.5.20 · branch-point cache + sampling masks + DSpark-under-PD)

Performance / hardware (continuing v0.5.20):

Under the Hood: Fast Start (weight-lifecycle decoupling) + HiSparse (sparse KV three-tier spillover)

My read: these two decouple ‘weights’ and ‘KV’ out of the engine process respectively — Fast Start decouples weights, HiSparse decouples sparse KV — the same closing logic as my OpenInfer / DFlash ‘reuse already-computed capability.’ Next step: test whether ‘Fast Start cold-start + HiSparse host fallback’ reproduces the same cost-down on a single card.

Standing Topics

PD disaggregation: vLLM v0.30.0 is the landing version for the CVE-2026-93436 fix (PD-disaggregation production upgrade target); HiSparse spilling KV to host memory indirectly relieves decode-side VRAM pressure under disagg; SGLang maintains v0.5.20’s DSpark-under-PD (#37709) + PD role hot-switch (#28403) + KV checksum (#39500), main keeps polishing.

Architecture evolution: vLLM 0.30: Fast Start (restart opt), HiSparse (sparse MLA host KV layer), dual-key watermark + speculative decoding, MRV2 graph-capture GC freeze (capture 12s→2s), online acceptance estimator extending adaptive verify to all draft-model spec decoders; SGLang 0.5.20 multi-pronged with sampling masks + branch-point cache + DSV4 Blackwell kernel.

PyTorch vs transformers: no ‘off transformers’ PR this period; boundary signal continues re-layering — vLLM has --model-impl transformers direct HF, SGLang docs add a ‘Transformers fallback’ page, model-definition layer converges to transformers, performance layer stays in engine kernel. Horizontal conclusion: stop treating ‘in-house model implementation’ as a performance prerequisite; only hot operators deserve in-house.

Step adaptation: Step-3.7-Flash deploy path stable (vLLM stepfun37 image + MTP num_speculative_tokens=3; SGLang dev-step-3.7-flash + multi-layer EAGLE; NVFP4 4 cards, FP8 KV); Step 5 Preview (600B/27B, 1M context, native text+vision, AA composite 44 top-3 open) weights to open-source 10/15, adaptation PRs will emerge then, no merge on the framework side now.

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

Meta DINOv3 uses Gram anchoring to preserve dense features, letting a 7B frozen visual backbone beat weak-supervision SOTA for the first time — robots get a label-free default ‘eye’; on the operator side Muon (Newton-Schulz orthogonalization) pushes every direction evenly and MegaBlocks no-drop MoE (block-sparse GEMM) drops capacity_factor, pushing training and inference cost down together; on the industry side Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order.

Paper Core (DINOv3 · Meta FAIR, arXiv:2508.10104)

Operator Deep-Dive: Muon Optimizer (Newton-Schulz orthogonalization)

Performance Optimization: MegaBlocks No-Drop MoE (block-sparse GEMM)

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

vLLM shipped annual major release v0.30.0 on 9/22 (762 commits / 315 contributors) — Fast Start persistent per-GPU weight-cache daemon (CUDA IPC mapping, H200 engine init 28.9s→8.2s), DeepSeek-V4.1-Flash day-0 support (MXFP8 KV on SM100), HiSparse host-resident KV spillover and dual-key Gumbel-max watermark landed together, and the CVE-2026-93436 fix officially lands, making v0.30.0 the recommended upgrade target for PD-disaggregation production (replacing the earlier rc0-only state); SGLang no new tag in 72h, still v0.5.20 (branch-point cache DSV4-Flash hit 43.8%→60.8%, sampling masks overlap +17%~+52%, DSpark-under-PD, B200 Blackwell 1.2×/1.45×) keeps leading; on the papers side Meta DINOv3 (7B self-supervised, Gram anchoring) is the first frozen visual backbone to beat weak-supervision SOTA, a label-free default eye for robots, operator Muon (Newton-Schulz orthogonalization) and MegaBlocks no-drop MoE (block-sparse GEMM) cut training and inference cost, and on the industry side Tesla Optimus finishes its Yangtze-delta audit with the first ~5000-unit order — the embodied ‘smarter brain + real landing’ keeps getting a price tag.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。