★ Most Worth Your Attention Today
SGLang v0.5.20 (shipped 9/18 22:41, 190 commits / 713 PRs / 237 contributors) — the only formal major release in the window, far ahead of vLLM’s same-period rc2 polishing pace. Among it, #37709 DSpark under PD directly connects to my own OpenInfer / DFlash / DSpark main line, and #34565 SWA branch-point cache measured TTFT −32%. vLLM, meanwhile, is on the “stable version sitting still, release line patching bugfixes” rhythm and v0.29.0 still carries two CVEs — exactly the same conclusion as recent weeks: “don’t sit on v0.29.0, follow main or wait for 0.29.1.”
It is actually the same thread as today’s papers headline: the framework makes draft KV reusable across nodes, the modeling side treats action as token — both are closing moves that “reuse existing capability, push it into usable form.” Decision Transformer (2021) is precisely the ideological ancestor of today’s every-VLA “action as token” (control = generation). The framework runs draft KV across nodes, the modeling side treats action as token — two facets of the same underlying move.
Three layers of fact:
- SGLang v0.5.20 is the window’s only major release (713 PRs / 237 contributors ≫ vLLM’s rc2 polish): day-0 new models dropped in a burst — GLM-5.3-Flash (
#36507/#38621), Hy4-Preview (#36805), Qwen3.8-Flash-Next (#37500), K2 Horizon (#37654/#38033), Nanbeige4.2 (#32151), SenseNova-U1.5-8B-MoT diffusion (#36606), FastH3 (#37480), VDN-H3 (#37903). Not “a few PRs” — a release cadence visibly faster than vLLM’s. - #37709 DSpark under PD (same source as OpenInfer/DFlash): one DCP1 prefill passes its DSpark draft KV straight to DCP-N decode, letting hybrid models like Kimi-Linear run DSpark in a “disaggregated + context-parallel” topology; verified on 8×B300 NIXL / Mooncake to 256K input. It turns cross-node draft-KV transfer into a reusable engineering paradigm, the same source as my Qwen3-4B DFlash.
- #34565 unified prefix-tree SWA branch-point cache measured TTFT −32%: under a shared system prompt, token hit rate 43.8%→60.8%, mean TTFT 1.57s→1.07s (~−32%); an optional external linker (
#37381) addresses a shared global memory pool via Mooncake / UMBP. Plus sampling-masks overlap scheduling (#36630/#36631): Qwen3-8B decode throughput batch1 +17%, batch64 +52%.
Actionable conclusion: if you run disaggregated PD, SGLang v0.5.20’s #37709 is the signal to put DSpark draft-KV transfer on the default path — “reuse draft KV across nodes” directly rewrites the speculative-decode topology assumption; for my own OpenInfer / Qwen3-4B DFlash the direction is consistent: cross-node draft-KV reuse is the next engineering form of the same tradeoff. On the vLLM side: upgrade to 0.29.1+ — the PD-metadata CVE-2026-93436 (CVSS 8.7) sits on the KV-transfer path, the current main PD attack surface.
Worth saying separately: vLLM v0.29.0 still carries CVE-2025-30165 and the more serious CVE-2026-93436 (PD-disaggregation rejected-request metadata not released, CVSS 8.7, fixed
#55677), the latter needing 0.29.1+. The KV-transfer path has become the main PD attack surface; every cycle now fixed-vuln scans it. The conclusion is blunt: don’t sit on v0.29.0, follow main or wait for 0.29.1+.
2. vLLM & SGLang Community Tracking
Version status: SGLang v0.5.20 (9/18 22:41, 190 commits / 713 PRs / 237 contributors) — the window’s only formal major release, with a visibly faster cadence than vLLM; vLLM stable still v0.29.0 (09-08/09-09), release line advanced to v0.30.0rc2 (~9/19, rc2 adds NIXL bugfix #57285 over 9/18’s rc1), also v0.30.0rc1 (9/18), v0.29.1rc0 (9/13), proto-v0.2.0 (9/17). This period’s increments come mainly from the main branch + SGLang’s formal release; vLLM has no new stable features.
vLLM (main · 0.30 line BF16 autotuning + NIXL robustness)
Version / feature continuation:
- Version: stable still v0.29.0; release line advanced to v0.30.0rc2 (rc1→rc2 only adds NIXL bugfix
#57285); the 0.30 line focuses on FlashInfer BF16 autotuning isolation and NIXL robustness, no new stable features;--model-impl transformersrunning HF stays the main-line signal. - Security (continuing 9/18): v0.29.0 is hit by
CVE-2025-30165(1 item);CVE-2026-93436(CVSS 8.7, PD-disaggregation rejected-request metadata not released, fixed#55677) needs 0.29.1+. The KV-transfer path has become the main PD attack surface; every cycle now fixed-vuln scans it.
SGLang (v0.5.20 · day-0 models + DSpark-under-PD + SWA cache)
New features / major adaptations:
- day-0 new models: GLM-5.3-Flash (
#36507/#38621), Hy4-Preview (#36805), Qwen3.8-Flash-Next (#37500), K2 Horizon (#37654/#38033), Nanbeige4.2 (#32151), SenseNova-U1.5-8B-MoT diffusion (#36606), FastH3 (#37480), VDN-H3 (#37903); impact = SGLang still leads day-0 new-model support. - Sampling masks (
#36630/#36631):return_sampling_maskreturns per-step support set + logp, so RL trainers can replay rollout without rebuilding top-k/top-p; under overlap scheduling Qwen3-8B decode throughput batch1 +17%, batch64 +52%. - Unified prefix-tree SWA branch-point cache (
#34565): shared system prompt token hit rate 43.8%→60.8%, mean TTFT 1.57s→1.07s (~−32%); optional external linker (#37381) addresses a shared global memory pool via Mooncake / UMBP. - DSpark under PD (
#37709): DCP1 prefill passes DSpark draft KV to DCP-N decode; Kimi-Linear runs DSpark in disaggregated + context-parallel topology; verified on 8×B300 NIXL / Mooncake to 256K input. - Other: Responses API storage becomes opt-in (
#39122, must NOT be enabled for PD deployments); SGLang Simulator (#33824, pure-CPU TTFT prediction error ~6%); removed strategy-based prefill CP v1.
Under the Hood: #37709 DSpark under PD (Cross-Node Draft-KV Reuse, Same Source as OpenInfer/DFlash)
- Traditional approach: DSpark under PD runs speculative only on the decode side; prefill and decode each manage their own KV.
- What #37709 changes: one DCP1 prefill transfers its DSpark draft KV directly to DCP-N decode, letting hybrid models like Kimi-Linear run DSpark in a “disaggregated + context-parallel” topology — turning cross-node draft-KV transfer into a reusable engineering paradigm, the same source as OpenInfer / DFlash (Qwen3-4B).
- My read: the value isn’t “resend the KV” but “make draft-KV transfer a default-path primitive” — the same DNA as 0913’s “VRAM tiered pooling,” 0916’s “DFlash2 pushes E[L] up,” 0917’s “fused kernel + NVFP4 compressed KV”: all hunt the optimal operating point on a tradeoff surface. For my own OpenInfer / Qwen3-4B DFlash, the next step is to test whether a “DCP1→DCP-N draft-KV transfer” reproduces the same E[L] lift on a single card.
Standing Topics
PD disaggregation: SGLang leads on two lines — #37709 DSpark-under-PD (DCP1→DCP-N draft-KV transfer) + the prior #28403 role hot-switch (KV pool role-agnostic). vLLM’s 0.30 line continues NIXL bugfixes.
Architecture evolution: SGLang unified prefix tree + SWA branch-point cache (#34565, +17pp hit / TTFT −32%) + sampling-masks overlap scheduling (#36630/#36631, decode +17~52%); vLLM 0.30 focuses on autotuning / bugfix.
PyTorch vs transformers: the boundary re-layers — vLLM --model-impl transformers runs HF directly; SGLang new models still follow the “cookbook + HF weights” dual track. Conclusion: the model-definition layer converges to transformers, the performance layer stays in engine kernels.
Step adaptation: no framework-side merge; Step-3.7-Flash (198B/11B/256K/~400TPS) deploy still vLLM stepfun37+MTP first-class > SGLang dev+EAGLE, NVFP4 4 cards; IPO still no formal filing.
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
Chen Lili et al.’s Decision Transformer (UC Berkeley/FAIR, arXiv:2106.01345, 2021) reframes RL as “conditional sequence modeling” — control=generation, the ideological ancestor of today’s every VLA (RT-2/π0/OpenVLA) “action as token,” and Unitree/UBTECH/Zhiyuan/XPeng “brains” are its enlarged versions; meanwhile FP8 and SmoothQuant complete the quantization chain, and the industry’s “Tesla Optimus China audit rushing 1000 units/week, but Counterpoint says 2026H1 global shipments hit 22k with only 13% truly entering factories” is the other face of the same “bodies get smarter, deploy lags” bottleneck — and Unitree’s UnifoLM-WLA-1.0, UBTECH’s Liuzhou factory, Digua Robot’s $400M Series C (bullish 09660.HK) are each putting a price tag on the embodied stack.
Paper Core (Decision Transformer · Chen et al., arXiv:2106.01345)
- One-line positioning: reframes RL as “conditional sequence modeling” — doesn’t fit a value function, doesn’t compute a policy gradient; instead uses a GPT-style causal Transformer with Return-to-Go (RTG, future cumulative reward) as prompt to autoregressively generate actions; mathematically control = generation.
- Offline-RL fit: no bootstrapping, no value overestimation, sparse rewards don’t collapse; matches or beats CQL/BCQ on Atari/Gym/Key-to-Door.
- VLA lineage: the ideological ancestor of today’s every-VLA “treat action as token” — RT-2, π0, OpenVLA are all direct descendants; Unitree/UBTECH/Zhiyuan/XPeng “brains” are their engineering enlargements. The framework side reuses draft KV across nodes, the modeling side treats action as token — the same closing logic underneath.
Operator Deep-Dive: FP8 Data Format & Quantization
- Positioning: 8-bit float; on Hopper/Blackwell matmul throughput is 2× FP16, bandwidth/VRAM halved. Two formats E4M3 (forward/weights) and E5M2 (gradients), landed via per-tensor/per-block scaling + stochastic rounding; MXFP8 uses one scale shared across a group.
- Usage: DeepSeek V4/V4.1, Qwen3, MiniMax use it at scale.
- Distinction: unlike GPTQ/AWQ (weight-only INT quantization) covered in the performance column, FP8 is the matmul itself at low precision — exactly the “lower precision” main line that 0917’s
#56935fused kernel + NVFP4 compressed KV and SGLang’s FP4 packing land on.
Performance Optimization: SmoothQuant (arXiv:2211.10438)
- Problem: a few LLM-activation channels have huge magnitudes (outliers), making activations un-quantizable.
- Method: a smoothing factor
s_j = max(|X_j|)^α / max(|W_j|)^(1-α)“transfers” outliers from activation to weights, so both fall into INT8 — achieving W8A8 with <1% precision loss and ~1.5–2× speedup. - Position: it is the first step of INT8; FP8 is the more aggressive next stop. Same source as 0917’s “move scheduling cost into the kernel, use lower-precision KV” — both thin the precision ledger.
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] Tesla Optimus: China supply-chain audit (Ningbo/Shanghai, rushing toward 1000 units/week by end of Sept, ~50k units in 2026); Toyota plans 400k into its factories.
- [Embodied] Counterpoint: 2026H1 global humanoid shipments 22k (+300%), top 5 all China, but 60%+ “performing,” only 13% truly entered factories → same thread as “bid ≠ commercial revenue.”
- [Embodied] Unitree 688836: shipped UnifoLM-WLA-1.0 base model (6B / 64 tasks / 7 open-source bests).
- [Embodied] UBTECH 09880.HK: Liuzhou 10k-unit factory in production, U1 starts delivery.
- [Embodied] Digua Robot (Horizon spinoff): $400M Series C → bullish for 09660.HK.
- [Embodied]: Tashi/Songyan/Zhiyuan “collectively supplement brains”; robot summit in Hangzhou; regulator moves against次新 speculation.
- [General]: Zhipu $5B financing, Huawei Ascend 960, OpenAI alignment-failure record, 10Y US Treasury breaks 5%.
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
On the framework side, SGLang v0.5.20 (9/18, 713 PRs / 237 contributors, the window’s only major release) — #37709 DSpark-under-PD lets one DCP1 prefill pass its draft KV straight to DCP-N decode (same source as OpenInfer/DFlash Qwen3-4B), #34565 SWA branch-point cache measured TTFT −32%, sampling-masks overlap scheduling decode +17~52%; vLLM stable stays v0.29.0 but carries two CVEs (incl. PD-metadata CVE-2026-93436 CVSS 8.7, fixed #55677, needs 0.29.1+) and the release line advances to v0.30.0rc2 — conclusion unchanged: don’t sit on v0.29.0, follow main or wait 0.29.1+. On the papers side, Decision Transformer (arXiv:2106.01345) reframes RL as conditional sequence modeling (control=generation), anchoring today’s every-VLA “action as token” lineage, with FP8 and SmoothQuant completing the quantization chain; on the industry side Tesla Optimus’s China audit rushes 1000 units/week but Counterpoint says 2026H1 only 13% truly entered factories, Unitree shipped UnifoLM-WLA-1.0, UBTECH’s Liuzhou factory went live, and Digua Robot’s $400M Series C is bullish for 09660.HK — each putting a price tag on the embodied “smarter brain + real landing” stack.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。