系列:每日AI热点

Daily AI Hotspot · 2026-09-19: SGLang v0.5.20 Ships (713 PRs), #37709 DSpark-under-PD Makes Draft KV Reusable Across Nodes; Decision Transformer Anchors VLA's 'Action as Token' Lineage

★ Most Worth Your Attention Today

SGLang v0.5.20 (shipped 9/18 22:41, 190 commits / 713 PRs / 237 contributors) — the only formal major release in the window, far ahead of vLLM’s same-period rc2 polishing pace. Among it, #37709 DSpark under PD directly connects to my own OpenInfer / DFlash / DSpark main line, and #34565 SWA branch-point cache measured TTFT −32%. vLLM, meanwhile, is on the “stable version sitting still, release line patching bugfixes” rhythm and v0.29.0 still carries two CVEs — exactly the same conclusion as recent weeks: “don’t sit on v0.29.0, follow main or wait for 0.29.1.”

It is actually the same thread as today’s papers headline: the framework makes draft KV reusable across nodes, the modeling side treats action as token — both are closing moves that “reuse existing capability, push it into usable form.” Decision Transformer (2021) is precisely the ideological ancestor of today’s every-VLA “action as token” (control = generation). The framework runs draft KV across nodes, the modeling side treats action as token — two facets of the same underlying move.

Three layers of fact:

  1. SGLang v0.5.20 is the window’s only major release (713 PRs / 237 contributors ≫ vLLM’s rc2 polish): day-0 new models dropped in a burst — GLM-5.3-Flash (#36507/#38621), Hy4-Preview (#36805), Qwen3.8-Flash-Next (#37500), K2 Horizon (#37654/#38033), Nanbeige4.2 (#32151), SenseNova-U1.5-8B-MoT diffusion (#36606), FastH3 (#37480), VDN-H3 (#37903). Not “a few PRs” — a release cadence visibly faster than vLLM’s.
  2. #37709 DSpark under PD (same source as OpenInfer/DFlash): one DCP1 prefill passes its DSpark draft KV straight to DCP-N decode, letting hybrid models like Kimi-Linear run DSpark in a “disaggregated + context-parallel” topology; verified on 8×B300 NIXL / Mooncake to 256K input. It turns cross-node draft-KV transfer into a reusable engineering paradigm, the same source as my Qwen3-4B DFlash.
  3. #34565 unified prefix-tree SWA branch-point cache measured TTFT −32%: under a shared system prompt, token hit rate 43.8%→60.8%, mean TTFT 1.57s→1.07s (~−32%); an optional external linker (#37381) addresses a shared global memory pool via Mooncake / UMBP. Plus sampling-masks overlap scheduling (#36630/#36631): Qwen3-8B decode throughput batch1 +17%, batch64 +52%.

Actionable conclusion: if you run disaggregated PD, SGLang v0.5.20’s #37709 is the signal to put DSpark draft-KV transfer on the default path — “reuse draft KV across nodes” directly rewrites the speculative-decode topology assumption; for my own OpenInfer / Qwen3-4B DFlash the direction is consistent: cross-node draft-KV reuse is the next engineering form of the same tradeoff. On the vLLM side: upgrade to 0.29.1+ — the PD-metadata CVE-2026-93436 (CVSS 8.7) sits on the KV-transfer path, the current main PD attack surface.

Worth saying separately: vLLM v0.29.0 still carries CVE-2025-30165 and the more serious CVE-2026-93436 (PD-disaggregation rejected-request metadata not released, CVSS 8.7, fixed #55677), the latter needing 0.29.1+. The KV-transfer path has become the main PD attack surface; every cycle now fixed-vuln scans it. The conclusion is blunt: don’t sit on v0.29.0, follow main or wait for 0.29.1+.

2. vLLM & SGLang Community Tracking

Version status: SGLang v0.5.20 (9/18 22:41, 190 commits / 713 PRs / 237 contributors) — the window’s only formal major release, with a visibly faster cadence than vLLM; vLLM stable still v0.29.0 (09-08/09-09), release line advanced to v0.30.0rc2 (~9/19, rc2 adds NIXL bugfix #57285 over 9/18’s rc1), also v0.30.0rc1 (9/18), v0.29.1rc0 (9/13), proto-v0.2.0 (9/17). This period’s increments come mainly from the main branch + SGLang’s formal release; vLLM has no new stable features.

vLLM (main · 0.30 line BF16 autotuning + NIXL robustness)

Version / feature continuation:

SGLang (v0.5.20 · day-0 models + DSpark-under-PD + SWA cache)

New features / major adaptations:

Under the Hood: #37709 DSpark under PD (Cross-Node Draft-KV Reuse, Same Source as OpenInfer/DFlash)

Standing Topics

PD disaggregation: SGLang leads on two lines — #37709 DSpark-under-PD (DCP1→DCP-N draft-KV transfer) + the prior #28403 role hot-switch (KV pool role-agnostic). vLLM’s 0.30 line continues NIXL bugfixes.

Architecture evolution: SGLang unified prefix tree + SWA branch-point cache (#34565, +17pp hit / TTFT −32%) + sampling-masks overlap scheduling (#36630/#36631, decode +17~52%); vLLM 0.30 focuses on autotuning / bugfix.

PyTorch vs transformers: the boundary re-layers — vLLM --model-impl transformers runs HF directly; SGLang new models still follow the “cookbook + HF weights” dual track. Conclusion: the model-definition layer converges to transformers, the performance layer stays in engine kernels.

Step adaptation: no framework-side merge; Step-3.7-Flash (198B/11B/256K/~400TPS) deploy still vLLM stepfun37+MTP first-class > SGLang dev+EAGLE, NVFP4 4 cards; IPO still no formal filing.

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

Chen Lili et al.’s Decision Transformer (UC Berkeley/FAIR, arXiv:2106.01345, 2021) reframes RL as “conditional sequence modeling” — control=generation, the ideological ancestor of today’s every VLA (RT-2/π0/OpenVLA) “action as token,” and Unitree/UBTECH/Zhiyuan/XPeng “brains” are its enlarged versions; meanwhile FP8 and SmoothQuant complete the quantization chain, and the industry’s “Tesla Optimus China audit rushing 1000 units/week, but Counterpoint says 2026H1 global shipments hit 22k with only 13% truly entering factories” is the other face of the same “bodies get smarter, deploy lags” bottleneck — and Unitree’s UnifoLM-WLA-1.0, UBTECH’s Liuzhou factory, Digua Robot’s $400M Series C (bullish 09660.HK) are each putting a price tag on the embodied stack.

Paper Core (Decision Transformer · Chen et al., arXiv:2106.01345)

Operator Deep-Dive: FP8 Data Format & Quantization

Performance Optimization: SmoothQuant (arXiv:2211.10438)

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

On the framework side, SGLang v0.5.20 (9/18, 713 PRs / 237 contributors, the window’s only major release) — #37709 DSpark-under-PD lets one DCP1 prefill pass its draft KV straight to DCP-N decode (same source as OpenInfer/DFlash Qwen3-4B), #34565 SWA branch-point cache measured TTFT −32%, sampling-masks overlap scheduling decode +17~52%; vLLM stable stays v0.29.0 but carries two CVEs (incl. PD-metadata CVE-2026-93436 CVSS 8.7, fixed #55677, needs 0.29.1+) and the release line advances to v0.30.0rc2 — conclusion unchanged: don’t sit on v0.29.0, follow main or wait 0.29.1+. On the papers side, Decision Transformer (arXiv:2106.01345) reframes RL as conditional sequence modeling (control=generation), anchoring today’s every-VLA “action as token” lineage, with FP8 and SmoothQuant completing the quantization chain; on the industry side Tesla Optimus’s China audit rushes 1000 units/week but Counterpoint says 2026H1 only 13% truly entered factories, Unitree shipped UnifoLM-WLA-1.0, UBTECH’s Liuzhou factory went live, and Digua Robot’s $400M Series C is bullish for 09660.HK — each putting a price tag on the embodied “smarter brain + real landing” stack.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。