★ Most Worth Your Attention Today
On a non-trading Sunday the frameworks shipped no new formal releases in the last 24h: vLLM still holds v0.30.0 + v0.30.1rc0, SGLang still holds v0.5.20, but SGLang’s main branch is refactoring densely (logic-kernel-group restore, unified MemCache eviction, sgl-router forwards built-body only, HiRadixCache→Unified Radix) — a sign v0.5.21 will close the “cache + routing” base. The single PR to watch is #38891, which turns kv-hints into a request-transport envelope so KV affinity routing no longer leans on out-of-band signals. StepFun’s Step 5 Preview is about 18 days from its 10-15 open-source drop.
This is really the same sentence written twice: both engines are converging “cache affinity + PD disaggregation + elastic scaling” into the next-phase spine, differing only in release cadence and degree of defaulting — vLLM already put the CVE fix and AMD NVFP4 quantization onto the 0.30 stable line, while SGLang is keeping the base refactor buried in main, to close it all at once in v0.5.21.
Three layers of fact:
- What the versions are: vLLM’s stable line is still v0.30.0 (09-22, 762 commits / 315 contributors), with the release line at v0.30.1rc0 (09-23, #58281 AMD MI355 dense NVFP4); SGLang’s stable line is still v0.5.20 (09-18, 713 PRs / 237 contributors), no v0.5.21, with main still committing densely on 09-26 (cumulative 18,939 commits).
- What SGLang is sitting on: #41243 logic-kernel-group restore, #41276 unified MemCache eviction cursor + lock receipt, sgl-router forwarding input_ids only for the built body, #40787 HiRadixCache→Unified Radix, #38891 kv-hints envelope into the request transport — the cache and routing base is being closed and hardened, expected to land together in v0.5.21.
- Where Step stands: Step 5 Preview (09-20, 600B/27B ultra-sparse MoE, 1M context, top-3 among open-weight on composite intelligence, single-task cost ≈ 1/8 of Claude Opus 5) will open its full weights 10-15 (about 18 days out); the community already has Step-5-Preview-BF16 weights (TypeSafeAI), runnable with
--trust-remote-code --reasoning-parser stepfun, but attribution and license are unconfirmed; on the IPO, no HKEX prospectus is public yet and the media line is “entered the listing-review stage” with no company comment — do not write “has filed.”
Actionable conclusion: SGLang users don’t chase main, wait for v0.5.21 (upgrade once the cache/routing base closes); production PD should keep anchoring on vLLM v0.30.0 (the line carrying the CVE-2026-93436 fix); for Step 5 Preview wait for the official 10-15 open-source drop for day-0, and validate the community BF16 weights at small traffic first — don’t treat self-reported numbers as a benchmark.
Worth emphasizing: SGLang stuffing kv-hints into the request envelope, together with #40787’s “HiRadixCache→Unified Radix,” are two faces of the same closing move — the former solves “de-out-of-band-ing routing decisions,” the latter solves “a single source of truth for prefix cache.” Advancing both in parallel pushes the most painful PD-disaggregation coupling, “cache-hit-rate ↔ affinity-routing,” out of the out-of-band-signal era into the request-inlined era. For my own OpenInfer / Qwen3-4B DFlash this is the same closing logic: make the reusable state (KV, draft) a unit that can be transparently carried and aligned, rather than coordinated through external side channels.
2. vLLM & SGLang Community Tracking
Version status: in-window vLLM still holds v0.30.0 (09-22, 762 commits / 315 contributors) + v0.30.1rc0 (09-23) on two lines; SGLang shipped no new release, latest stable still v0.5.20 (09-18, 713 PRs / 237 contributors), with main still committing densely on 09-26 (cumulative 18,939 commits) but no v0.5.21 tag.
vLLM (v0.30.0 + v0.30.1rc0 · no new formal release, security and landing features carry over)
Version / security status (no new formal release):
- Version: stable still v0.30.0 (09-22, 762 commits / 315 contributors); release line at v0.30.1rc0 (09-23, #58281 AMD MI355 dense NVFP4) — no new tag since 09-23, AMD users may try NVFP4 but it is still RC.
- Security carries over: CVE-2026-93436 (CVSS 8.7, PD-disaggregation rejected request’s decode-side KV metadata not freed → memory-exhaustion DoS) needs ≥0.29.1 (still only rc0 available) — production PD must move to the 0.30 line.
- Landing-feature recap: v0.30.0’s Fast Start (H200 engine init 28.9s→8.2s), HiSparse host-resident KV spillover, dual-key Gumbel-max watermark, MRV2 graph-capture freeze (12s→2s) → zero-change rolling-restart cost-down + sparse-MLA decode VRAM relief; no new changes this window but still the most worthwhile primitives on that line.
SGLang (v0.5.20 · main refactoring densely, v0.5.21 to close cache/routing base)
Community dev / refactor (recent-window additions):
- Cache/routing base closing: #41243 logic-kernel-group restore (09-26), #41276 unified MemCache eviction cursor + lock receipt (09-26), sgl-router forwards input_ids only for the built body (09-26), #40787 HiRadixCache→Unified Radix (09-22).
- kv-hints envelope: #38891 turns “KV hit hint” into a request-transport envelope (09-23, see the deep-dive below).
- Model / structured: #39026 XGrammar V4.1 DSML parameter constraints (09-21) → Agent tool-calling / structured output more robust.
Under the Hood: SGLang PR #38891 (kv-hints Envelope Into the Request Transport)
The most worthwhile technical read of the day.
Mechanism: PR #38891 (09-23) turns “KV hit hint” into a request-transport envelope — the request carries prefix-cache location / hit hints and transparently passes them along the transport. Routing decisions therefore no longer depend on out-of-band signals, enabling stable affinity routing under PD disaggregation, cross-node, and elastic scaling.
Direct impact:
- Hit-rate↑ → TTFT↓ + prefill compute cost↓, with the largest gains on multi-turn / Agent / RAG workloads.
- No strong-consistency guarantee: must pair with the radix index and KV event publishing; expected to land in v0.5.21.
My read: turning “cache affinity” from an out-of-band signal into a “request-inlined envelope” is a critical step in PD-disaggregation engineering. The problem with side-channel signals is that their causal link to the request is loosely coupled and drifts the moment scaling gets frequent; an inlined envelope makes “which prefiller should this request go to” a self-carried, verifiable field. SGLang advancing it alongside #40787’s Unified Radix aligns “where to route” and “the cache source of truth” in the same generation.
Standing Topics
PD disaggregation: vLLM 0.30 lands CVE-2026-93436 fix + HiSparse; SGLang kv-hints envelope #38891 + sgl-router cache_aware + Dynamo KV Router (x-prefiller-host-port); llm-d combines PD with prefix / load-aware scheduling.
Architecture evolution: vLLM Fast Start + HiSparse + dual-key watermark + MRV2 graph-capture freeze; SGLang Unified Radix (branch-point cache #34565) + sampling masks + DSpark-under-PD + HiCache closing (#40787).
PyTorch vs transformers: no “off-transformers in-house stack” PR, the boundary re-layering continues (vLLM --model-impl transformers + SGLang Transformers fallback); the model-definition layer converges on transformers, the performance layer stays in engine kernels.
Step adaptation: Step-3.7-Flash dual-framework deployment is mature (vLLM stepfun37 + MTP > SGLang dev + EAGLE, NVFP4 4 cards); Step 5 Preview open-sources 10-15, community BF16 weights already appear, official day-0 pending 10-15.
3. AI Papers & Industry Hotspots
Today’s Highlight (1 sentence)
DriveVLM closes the autonomous-driving dimension loop: a VLM slow-thinking cognitive anchor plus a fast-planner high-frequency fallback as a dual-system paradigm, inherited within a year by Helix/GR00T/Xpeng/Li Auto; on the operator side Rectified Flow straightens generation to 1 step, and the Tensor Parallel deep-dive gives the sharding baseline for running VLA on dual in-vehicle chips.
Paper Core (DriveVLM · Tsinghua × CASIA “Tiangong”, arXiv:2402.12289, 2024-02)
- One-line positioning: the founding work that put VLM slow-thinking into the autonomous-driving loop and proposed the System1/System2 dual-system — decoupling “cognition” from “execution” into two systems at different frequencies.
- Core idea: a five-stage CoT (scene description → key-object 3D boxes → driving intention → discrete meta-action dictionary → trajectory waypoints), where meta-actions collapse free-form outputs into a finite decision set — interpretable and constrainable; a vision-language alignment module uses 3D perceptual features to patch VLM’s spatial weakness.
- Impact confirmed: EMMA (fully textualized single model, the research ceiling) vs DriveVLM-Dual (CoT anchor + fast-system fallback, the production architecture) vs Helix (latent vector, whole-robot) — three paradigms contrasted; production autonomous-driving VLAs are all engineering enlargements of it — today’s Xpeng / Li Auto VLA production architectures are, at the bone, enlargements of this dual system.
Operator Deep-Dive: Rectified Flow (Liu et al., ICLR 2023)
- Positioning: straight-line interpolation x_t=(1−t)x₀+t·x₁, velocity u=x₁−x₀, MSE regression; reflow iterates straightening → 1-step generation.
- In use: SD3 / FLUX (logit-normal t) for image generation; π0 / π0.5 / GR00T N2 attach the action head on Rectified Flow for robot policies.
- Difference from Flow Matching: Flow Matching is the “transport framework” (any probability path gets a velocity field); Rectified Flow is the “transport-geometry viewpoint” (take the straight path, use reflow to force it straight) — on the robot side 10 steps → 1-2 steps = control rate 5Hz → 30Hz+, directly lifting closed-loop responsiveness.
Performance Optimization: Tensor Parallel Deep-Dive (Megatron-LM)
- Positioning: column-parallel zero-communication + row-parallel AllReduce, f/g conjugate operators do 2 AllReduces per layer forward (≈
2(p−1)/p·b·s·hbytes). - Constraints: TP≤8 and ≤ number of heads, stay within a single NVLink machine; TP8 cuts VRAM ÷8 near-linearly.
- Embodied link: dual in-vehicle chips run VLA at TP2 (Xpeng dual-Turing / Thor) — moving the attention-parallel sharding baseline straight onto the board-level interconnect of dual Orin/Thor.
Industry Hotspots (embodied companies / supply-chain dispatches · pinned)
- [Embodied] Tesla Optimus: Gen3 finalized + Yangtze-delta audit wrapped (weekly output 500+ → 1000 by end-Sept → 2000-2500/week by year-end, ~50k in 2026) → bullish on Tuopu 601689 / Sanhua 002050 / Joyson 600699 / Lepu 688017 / Robot ETF 562500.
- [Embodied] AgiBot (Zhiyuan): 20,000th unit delivered to Chimelong + 300 units in steady park operation; Galbot (Yinhe Tongyong) ran 7×24 at CATL for 3+ months steady; Zhijian Power’s 100-unit delivery landed in Lepu harmonic reducers and Dongshan Precision lines; Morgan Stanley raised its 2026 China humanoid forecast to 50k twice in a row.
- [General] LLM price war: V4.1 Flash / MiMo-V2.6 / Opus 5.5 / GPT-6 Luna all cutting prices (Zhipu market cap fell 300B HKD); Xpeng’s 2nd-gen VLA pushed to MONA (response +300%, 720B distilled to 508TOPS); DrivingBench measured general LLM direct-driving at ~¥40/km → dedicated VLA architecture is the right answer; Waymo World Model (Genie 3) simulates extreme scenarios.
⚠️ Industry dynamics do not constitute investment advice.
4. The One-Line Takeaway
On a non-trading Sunday the frameworks shipped no new formal releases in the last 24h: vLLM still holds v0.30.0 + v0.30.1rc0 (CVE-2026-93436 fix is on the 0.30 line, production PD must upgrade), SGLang still holds v0.5.20 but its main branch refactors densely (#41243/#41276/#40787/#38891) signaling v0.5.21 will close the “cache + routing” base; the PR to watch, #38891, turns kv-hints into a request-transport envelope so affinity routing under PD disaggregation no longer depends on out-of-band signals. StepFun’s Step 5 Preview is about 18 days from its 10-15 open-source drop, with community BF16 weights already present but attribution unconfirmed. On the papers side, DriveVLM puts a VLM slow-thinking + fast-planner dual-system into the autonomous-driving loop (five-stage CoT + meta-actions), inherited within a year by Helix/GR00T/Xpeng/Li Auto; operators Rectified Flow straightens generation to 1 step and the Tensor Parallel deep-dive gives the sharding baseline for dual in-vehicle-chip VLA; on the industry side Tesla Optimus Gen3 finalized + Yangtze-delta audit wrapped, AgiBot delivered its 20,000th unit to Chimelong, and Morgan Stanley twice raised its 2026 China humanoid forecast to 50k — embodied “brain-building + landing” keeps getting itemized and priced.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。