系列:每日AI热点

Daily AI Hotspot · 2026-09-27: No New Release on Sunday — SGLang main Closes Cache/Routing Base, kv-hints Envelope Enters Transport; DriveVLM Puts VLM Slow-Thinking Into the Autonomous-Driving Loop

★ Most Worth Your Attention Today

On a non-trading Sunday the frameworks shipped no new formal releases in the last 24h: vLLM still holds v0.30.0 + v0.30.1rc0, SGLang still holds v0.5.20, but SGLang’s main branch is refactoring densely (logic-kernel-group restore, unified MemCache eviction, sgl-router forwards built-body only, HiRadixCache→Unified Radix) — a sign v0.5.21 will close the “cache + routing” base. The single PR to watch is #38891, which turns kv-hints into a request-transport envelope so KV affinity routing no longer leans on out-of-band signals. StepFun’s Step 5 Preview is about 18 days from its 10-15 open-source drop.

This is really the same sentence written twice: both engines are converging “cache affinity + PD disaggregation + elastic scaling” into the next-phase spine, differing only in release cadence and degree of defaulting — vLLM already put the CVE fix and AMD NVFP4 quantization onto the 0.30 stable line, while SGLang is keeping the base refactor buried in main, to close it all at once in v0.5.21.

Three layers of fact:

  1. What the versions are: vLLM’s stable line is still v0.30.0 (09-22, 762 commits / 315 contributors), with the release line at v0.30.1rc0 (09-23, #58281 AMD MI355 dense NVFP4); SGLang’s stable line is still v0.5.20 (09-18, 713 PRs / 237 contributors), no v0.5.21, with main still committing densely on 09-26 (cumulative 18,939 commits).
  2. What SGLang is sitting on: #41243 logic-kernel-group restore, #41276 unified MemCache eviction cursor + lock receipt, sgl-router forwarding input_ids only for the built body, #40787 HiRadixCache→Unified Radix, #38891 kv-hints envelope into the request transport — the cache and routing base is being closed and hardened, expected to land together in v0.5.21.
  3. Where Step stands: Step 5 Preview (09-20, 600B/27B ultra-sparse MoE, 1M context, top-3 among open-weight on composite intelligence, single-task cost ≈ 1/8 of Claude Opus 5) will open its full weights 10-15 (about 18 days out); the community already has Step-5-Preview-BF16 weights (TypeSafeAI), runnable with --trust-remote-code --reasoning-parser stepfun, but attribution and license are unconfirmed; on the IPO, no HKEX prospectus is public yet and the media line is “entered the listing-review stage” with no company comment — do not write “has filed.”

Actionable conclusion: SGLang users don’t chase main, wait for v0.5.21 (upgrade once the cache/routing base closes); production PD should keep anchoring on vLLM v0.30.0 (the line carrying the CVE-2026-93436 fix); for Step 5 Preview wait for the official 10-15 open-source drop for day-0, and validate the community BF16 weights at small traffic first — don’t treat self-reported numbers as a benchmark.

Worth emphasizing: SGLang stuffing kv-hints into the request envelope, together with #40787’s “HiRadixCache→Unified Radix,” are two faces of the same closing move — the former solves “de-out-of-band-ing routing decisions,” the latter solves “a single source of truth for prefix cache.” Advancing both in parallel pushes the most painful PD-disaggregation coupling, “cache-hit-rate ↔ affinity-routing,” out of the out-of-band-signal era into the request-inlined era. For my own OpenInfer / Qwen3-4B DFlash this is the same closing logic: make the reusable state (KV, draft) a unit that can be transparently carried and aligned, rather than coordinated through external side channels.

2. vLLM & SGLang Community Tracking

Version status: in-window vLLM still holds v0.30.0 (09-22, 762 commits / 315 contributors) + v0.30.1rc0 (09-23) on two lines; SGLang shipped no new release, latest stable still v0.5.20 (09-18, 713 PRs / 237 contributors), with main still committing densely on 09-26 (cumulative 18,939 commits) but no v0.5.21 tag.

vLLM (v0.30.0 + v0.30.1rc0 · no new formal release, security and landing features carry over)

Version / security status (no new formal release):

SGLang (v0.5.20 · main refactoring densely, v0.5.21 to close cache/routing base)

Community dev / refactor (recent-window additions):

Under the Hood: SGLang PR #38891 (kv-hints Envelope Into the Request Transport)

The most worthwhile technical read of the day.

Mechanism: PR #38891 (09-23) turns “KV hit hint” into a request-transport envelope — the request carries prefix-cache location / hit hints and transparently passes them along the transport. Routing decisions therefore no longer depend on out-of-band signals, enabling stable affinity routing under PD disaggregation, cross-node, and elastic scaling.

Direct impact:

My read: turning “cache affinity” from an out-of-band signal into a “request-inlined envelope” is a critical step in PD-disaggregation engineering. The problem with side-channel signals is that their causal link to the request is loosely coupled and drifts the moment scaling gets frequent; an inlined envelope makes “which prefiller should this request go to” a self-carried, verifiable field. SGLang advancing it alongside #40787’s Unified Radix aligns “where to route” and “the cache source of truth” in the same generation.

Standing Topics

PD disaggregation: vLLM 0.30 lands CVE-2026-93436 fix + HiSparse; SGLang kv-hints envelope #38891 + sgl-router cache_aware + Dynamo KV Router (x-prefiller-host-port); llm-d combines PD with prefix / load-aware scheduling.

Architecture evolution: vLLM Fast Start + HiSparse + dual-key watermark + MRV2 graph-capture freeze; SGLang Unified Radix (branch-point cache #34565) + sampling masks + DSpark-under-PD + HiCache closing (#40787).

PyTorch vs transformers: no “off-transformers in-house stack” PR, the boundary re-layering continues (vLLM --model-impl transformers + SGLang Transformers fallback); the model-definition layer converges on transformers, the performance layer stays in engine kernels.

Step adaptation: Step-3.7-Flash dual-framework deployment is mature (vLLM stepfun37 + MTP > SGLang dev + EAGLE, NVFP4 4 cards); Step 5 Preview open-sources 10-15, community BF16 weights already appear, official day-0 pending 10-15.

3. AI Papers & Industry Hotspots

Today’s Highlight (1 sentence)

DriveVLM closes the autonomous-driving dimension loop: a VLM slow-thinking cognitive anchor plus a fast-planner high-frequency fallback as a dual-system paradigm, inherited within a year by Helix/GR00T/Xpeng/Li Auto; on the operator side Rectified Flow straightens generation to 1 step, and the Tensor Parallel deep-dive gives the sharding baseline for running VLA on dual in-vehicle chips.

Paper Core (DriveVLM · Tsinghua × CASIA “Tiangong”, arXiv:2402.12289, 2024-02)

Operator Deep-Dive: Rectified Flow (Liu et al., ICLR 2023)

Performance Optimization: Tensor Parallel Deep-Dive (Megatron-LM)

Industry Hotspots (embodied companies / supply-chain dispatches · pinned)

⚠️ Industry dynamics do not constitute investment advice.

4. The One-Line Takeaway

On a non-trading Sunday the frameworks shipped no new formal releases in the last 24h: vLLM still holds v0.30.0 + v0.30.1rc0 (CVE-2026-93436 fix is on the 0.30 line, production PD must upgrade), SGLang still holds v0.5.20 but its main branch refactors densely (#41243/#41276/#40787/#38891) signaling v0.5.21 will close the “cache + routing” base; the PR to watch, #38891, turns kv-hints into a request-transport envelope so affinity routing under PD disaggregation no longer depends on out-of-band signals. StepFun’s Step 5 Preview is about 18 days from its 10-15 open-source drop, with community BF16 weights already present but attribution unconfirmed. On the papers side, DriveVLM puts a VLM slow-thinking + fast-planner dual-system into the autonomous-driving loop (five-stage CoT + meta-actions), inherited within a year by Helix/GR00T/Xpeng/Li Auto; operators Rectified Flow straightens generation to 1 step and the Tensor Parallel deep-dive gives the sharding baseline for dual in-vehicle-chip VLA; on the industry side Tesla Optimus Gen3 finalized + Yangtze-delta audit wrapped, AgiBot delivered its 20,000th unit to Chimelong, and Morgan Stanley twice raised its 2026 China humanoid forecast to 50k — embodied “brain-building + landing” keeps getting itemized and priced.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。