★ Most Worth Your Attention Today
NVIDIA GR00T N2 / DreamZero World Action Model (WAM): it upgrades the robot brain from a ‘see-image → output-action’ VLA to one that ‘first simulates in its head what the world will look like after it acts, then decides how to move’ — ~2× SOTA VLA success on new tasks + new environments. This is the single most worth-remembering item today, because it rewrites the parts list of a ‘robot brain’: the world model goes from optional module to standard equipment.
It is actually the same sentence as today’s framework / performance lines seen from different facets: the framework is making draft KV and prefix caches ‘reusable across nodes / across replicas,’ while WAM turns ‘future video-frame prediction’ into the criterion for action feasibility — both swap ‘the hard-to-learn thing’ for ‘the thing you can learn from massive data’: the framework makes PD draft-KV reuse a default primitive; WAM swaps ‘is the action right’ for ‘is the video prediction right’ (the latter learns from massive unlabeled video). Underneath is the same closing logic of ‘reuse existing capability, push it into usable form.’
Three layers of fact:
- What WAM is: VLA is a no-time-arrow ‘see → act’ mapping that does not know consequences; WAM uses a shared diffusion Transformer (~14B@7Hz) to jointly predict future video frames + actions, rolling latent ẑ_{t+1}=f_θ(z_t,a_t), with the action head on a flow-matching ODE. The criterion is neat: an imagined future frame being correct → the action is probably feasible; imagined drift → the action probably fails.
- Result: new task + new environment success ≈ 2× SOTA VLA, double-first on RoboArena + MolmoSpaces; a new body needs only 30 minutes of play data to transfer; the failure mode shifts from ‘wrong action’ to ‘world-model hallucination’ (model-plant mismatch) — the edge-compute bar rises, bullish for the robot-chip segment.
- Matching engineering cost-down: SGLang v0.5.20’s
#34565branch-point cache (free cost-down for Agent / multi-turn / RAG) + NVIDIA Dynamo KV-aware routing (4 replicas TTFT −50%, 89% hit) eat the ‘duplicate prefill’ — the same source as my OpenInfer / DFlash: all push ‘already-computed capability’ into maximum reuse so inference cost lands on a better operating point.
Worth saying separately: WAM pushes the ‘world model’ from an optional paper module to a robot-brain standard, meaning the edge has to run this 14B latent roll + video diffusion — which closes the loop exactly with today’s operator column KDA (‘fixed state, no KV growth → end-side VLA long-horizon tasks run’) and the 09660.HK (Horizon) robot-chip logic. World model + attention-saving + KV-saving are three lines of the same embodied-compute ledger.
2. vLLM & SGLang Community Tracking
Version status: no new release in the last 24h; this period fills in SGLang v0.5.20 details + corrects vLLM v0.30.0rc2 fix attribution on top of 9/18–9/19. SGLang v0.5.20 (9/18, 190 commits / 713 PRs / 237 contributors) remains the window’s only formal major release; vLLM stable still v0.29.0 (09-08), release line v0.30.0rc2, 0.30 line only bugfix polish.
vLLM (main · 0.30 line NIXL robustness + transformers direct-run)
Version / security continuation:
- Version correction:
v0.30.0rc2is actually #57570 (NIXL does not generate a receive report for notification-only requests) — a PD-disaggregation counting / stability fix; the 9/19 mis-attributed #57285 was actually rc1’s FlashInfer BF16 autotuning isolation. Stable still v0.29.0;v0.29.1rc0(#56122 speculative-decode dual-key watermark) not formally released; 0.30 line only patches bugfixes. - Security (continuing):
CVE-2026-93436(≤0.29.0, PD rejected-request metadata not released → memory exhaustion, CVSS 8.7, fixed#55677) needs 0.29.1+; the KV-transfer path has become the main PD attack surface; every cycle now fixed-vuln scans it. - Architecture signal:
--model-impl transformersrunning HF directly is now a stable capability — the model-definition layer converges to transformers, the performance layer stays in engine kernels.
SGLang (v0.5.20 · HRRN scheduling + TRT-LLM Blackwell + Mamba SSM + branch-point cache)
New features / major adaptations:
- Scheduling: v0.5.20 adds HRRN policy (GLM-5.2 trace ~69% lower TTFT than FCFS), with a CPU Simulator for pre-launch planning.
- New hardware: Blackwell TRT-LLM kernel vs FlashMLA on B200 prefill 1.2× / decode 1.45×; AMD GLM-5.2 4×MI355X load 505.7s → 40.4s; ROCm10 default.
- Models: first Mamba 1/2 SSM support + GLM-5.3-Flash / Hy4-Preview / Qwen3.8-Flash-Next etc.; retired CUDA12 wheels (last in v0.5.19).
Under the Hood: #34565 Unified Prefix-Tree SWA Branch-Point Cache (Free Cost-Down for Agent / Multi-turn / RAG)
- Traditional approach: after a request forks from a shared prefix, downstream branches must recompute the SWA window state.
- What #34565 changes: it additionally caches the SWA window state at the fork point, so downstream branches reuse instead of recompute — extending ‘prefix tree’ from ‘token hit’ to ‘window-state hit’.
- Measured: on DSV4-Flash token hit rate 43.8% → 60.8%, mean TTFT 1.57s → 1.07s (~−32%). Workloads like Agent / multi-turn / RAG with ‘shared long system prompt + many branches’ get a free cost-down.
My read: this is the same DNA as 0919’s
#37709 DSpark-under-PD(cross-node draft-KV reuse) and today’s performance column KV-aware routing (cross-replica prefix-block reuse) — all build reuse primitives on ‘already-computed capability.’ For my own OpenInfer / Qwen3-4B DFlash, the next step is to test whether ‘branch-point cache + draft-KV reuse’ reproduces the same E[L] lift on a single card.
Standing Topics
PD disaggregation: vLLM #57570 purifies NIXL counting; SGLang #37709 DSpark-under-PD + #28403 role hot-switch lead on two lines; /v1/responses not persisted by default (PD must not be enabled).
Architecture evolution: SGLang v0.5.20 multi-pronged (prefix tree + branch point, sampling masks overlap, TRT-LLM Blackwell, Mamba SSM, HRRN, Simulator); vLLM 0.30 polish.
PyTorch vs transformers: no ‘off transformers’ PR; boundary signal continues (model layer transformers, hot operators in-house).
Step adaptation: vLLM stepfun37 (MTP k=3) > SGLang dev (EAGLE draft=4), both NVFP4 4 cards + FP8 KV; StepAudio 3 edge-cloud cohere; Step IPO enters review stage (media wording, no public prospectus seen).
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
NVIDIA GR00T N2 / DreamZero World Action Model (WAM) upgrades the robot brain from VLA to ‘simulate the future in its head before acting’ — 2× success on new environments, world model from optional module to standard, edge-compute bar raised bullish for robot chips; the KDA operator makes linear attention beat MLA on every task with KV −75% and 6× decode at 1M context, and NVIDIA Dynamo KV-aware routing measured TTFT −50% with 89% hit on 4 replicas forms a trio with PD disaggregation / KV Offload; on the industry side Tesla Optimus starts a new China supply-chain audit, UBTECH’s Liuzhou factory is in production, and Digua Robot’s $400M Series C is bullish for 09660.HK — the embodied ‘smarter brain + real landing’ keeps getting a price tag.
Paper Core (NVIDIA GR00T N2 / DreamZero World Action Model)
- One-line positioning: VLA is a no-time-arrow ‘see → act’ mapping that ignores consequences; WAM uses a shared diffusion Transformer (~14B@7Hz) to jointly predict future video frames + actions, rolling latent ẑ_{t+1}=f_θ(z_t,a_t), with the action head on a flow-matching ODE.
- Core criterion: an imagined future frame being correct → the action is probably feasible; imagined drift → the action probably fails — swapping ‘action correctness’ (hard to learn) for ‘video-prediction correctness’ (learnable from massive unlabeled video).
- Result / impact: new task + new environment success ≈ 2× SOTA VLA, double-first on RoboArena + MolmoSpaces; a new body needs only 30 minutes of play data to transfer; the failure mode shifts from ‘wrong action’ to ‘world-model hallucination’ (model-plant mismatch) — world model from optional module to robot-brain standard, edge-compute bar raised bullish for the robot-chip segment.
Operator Deep-Dive: KDA (Kimi Delta Attention, Moonshot arXiv:2510.26692)
- Positioning: linear attention’s ‘per-channel forgetting gate’ — GDN gives one scalar decay per head (all channels forget at the same speed), KDA gives each channel an independent forgetting rate Diag(α_t); the transition matrix takes the diagonal + rank-1 (DPLR) special case a=b=k, making chunkwise ~2× faster.
- Effect: under the same training recipe, beats full MLA on every task, KV cache −75%, 6× decode throughput at 1M context.
- In use / parallel: Kimi Linear 48B-A3B (KDA:MLA=3:1), Kimi K3 (2.8T / 1M context); parallel with Qwen3-Next GDN, DeepSeek V4 MLA+NSA, MiniMax Lightning — four ‘attention-saving’ routes.
- Embodied relevance: fixed state with no KV growth → end-side VLA long-horizon tasks run — the same source as WAM’s edge-compute bar.
Performance Optimization: KV Cache-Aware Routing (prefix-aware routing, NVIDIA Dynamo)
- Problem: ‘duplicate prefill’ in multi-replica clusters — the same prefix recomputed by several replicas.
- Method: a global radix tree indexes each replica’s cached prefix blocks and selects a replica by ‘cache overlap + decode load’ cost.
- Measured (Baseten, Qwen3 Coder 480B, 4 replicas): TTFT −50%, TPOT −34%, 89% hit; production traffic P95 −48%.
- Position: forms a trio with PD disaggregation and KV Offload.
- Embodied relevance: a robot fleet shares a cloud-brain cache, and TTFT decides first-response latency — the same closing logic as the framework side’s ‘draft KV / prefix reuse’.
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] Tesla Optimus: a new round of China supply-chain mass-production audit (9/16-17 Ningbo → Shanghai/Hangzhou/Xiamen, orders placed, ~50k units target in 2026) → bullish for Tuopu 601689 / Sanhua 002050 / Leader 688017; note Musk’s ‘high volume by 2027’ wording.
- [Embodied] UBTECH 09880.HK: Liuzhou 10k-unit factory in production (10 min/unit, 1.3万+ orders); Digua Robot (Horizon spinoff) $400M Series C (Sunrise chip 8M+ shipped) → bullish for Horizon 09660.HK.
- [Embodied]: 2025 global general humanoid shipments 18k (+508%), China holds 97% of manufacturing + export; MIIT expects China 2026 production >100k, parts localization >75% → Unitree 688836 / Robot ETF 562500 / Shangwei 688585.
- [General]: Qwen3.8-LiveTranslate real-time interpreting (LAAL 2.8s→2.3s); Gemini autonomously hacked 3 real companies in a safety exercise then self-terminated (Agent safety goes practical); NYT v. OpenAI / Microsoft enters summary judgment (authorized-data + synthetic-data routes benefit).
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
On the papers side, NVIDIA GR00T N2 / DreamZero World Action Model (WAM) upgrades the robot brain from VLA to ‘simulate the future in its head before acting’ — ~2× SOTA VLA success on new environments, world model to standard, edge-compute bar raised bullish for robot chips; the KDA operator makes linear attention beat MLA on every task with KV −75% and 6× decode at 1M context, and NVIDIA Dynamo KV-aware routing measured TTFT −50% with 89% hit on 4 replicas forms a trio with PD disaggregation / KV Offload; on the framework side SGLang v0.5.20 fills in details — HRRN scheduling (GLM-5.2 trace ~69% lower TTFT than FCFS), TRT-LLM Blackwell 1.2× prefill / 1.45× decode over FlashMLA, first Mamba SSM support, and #34565 branch-point cache measured TTFT −32%; vLLM stable stays v0.29.0 carrying CVE-2026-93436 (fixed #55677, needs 0.29.1+) with a corrected rc2 attribution to #57570 — conclusion unchanged: don’t sit on v0.29.0, follow main or wait for 0.29.1+; on the industry side Tesla Optimus starts a new China supply-chain audit, UBTECH’s Liuzhou factory is in production, and Digua Robot’s $400M Series C is bullish for 09660.HK — the embodied ‘smarter brain + real landing’ keeps getting a price tag.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。