系列:每日AI热点

Daily AI Hotspot · 2026-09-20: NVIDIA GR00T N2 Upgrades the Robot Brain From VLA to a World Action Model (2× Success on New Environments); SGLang v0.5.20 HRRN Scheduling Cuts TTFT 69%, NVIDIA Dynamo KV-Aware Routing Cuts TTFT −50%

★ Most Worth Your Attention Today

NVIDIA GR00T N2 / DreamZero World Action Model (WAM): it upgrades the robot brain from a ‘see-image → output-action’ VLA to one that ‘first simulates in its head what the world will look like after it acts, then decides how to move’ — ~2× SOTA VLA success on new tasks + new environments. This is the single most worth-remembering item today, because it rewrites the parts list of a ‘robot brain’: the world model goes from optional module to standard equipment.

It is actually the same sentence as today’s framework / performance lines seen from different facets: the framework is making draft KV and prefix caches ‘reusable across nodes / across replicas,’ while WAM turns ‘future video-frame prediction’ into the criterion for action feasibility — both swap ‘the hard-to-learn thing’ for ‘the thing you can learn from massive data’: the framework makes PD draft-KV reuse a default primitive; WAM swaps ‘is the action right’ for ‘is the video prediction right’ (the latter learns from massive unlabeled video). Underneath is the same closing logic of ‘reuse existing capability, push it into usable form.’

Three layers of fact:

  1. What WAM is: VLA is a no-time-arrow ‘see → act’ mapping that does not know consequences; WAM uses a shared diffusion Transformer (~14B@7Hz) to jointly predict future video frames + actions, rolling latent ẑ_{t+1}=f_θ(z_t,a_t), with the action head on a flow-matching ODE. The criterion is neat: an imagined future frame being correct → the action is probably feasible; imagined drift → the action probably fails.
  2. Result: new task + new environment success ≈ 2× SOTA VLA, double-first on RoboArena + MolmoSpaces; a new body needs only 30 minutes of play data to transfer; the failure mode shifts from ‘wrong action’ to ‘world-model hallucination’ (model-plant mismatch) — the edge-compute bar rises, bullish for the robot-chip segment.
  3. Matching engineering cost-down: SGLang v0.5.20’s #34565 branch-point cache (free cost-down for Agent / multi-turn / RAG) + NVIDIA Dynamo KV-aware routing (4 replicas TTFT −50%, 89% hit) eat the ‘duplicate prefill’ — the same source as my OpenInfer / DFlash: all push ‘already-computed capability’ into maximum reuse so inference cost lands on a better operating point.

Worth saying separately: WAM pushes the ‘world model’ from an optional paper module to a robot-brain standard, meaning the edge has to run this 14B latent roll + video diffusion — which closes the loop exactly with today’s operator column KDA (‘fixed state, no KV growth → end-side VLA long-horizon tasks run’) and the 09660.HK (Horizon) robot-chip logic. World model + attention-saving + KV-saving are three lines of the same embodied-compute ledger.

2. vLLM & SGLang Community Tracking

Version status: no new release in the last 24h; this period fills in SGLang v0.5.20 details + corrects vLLM v0.30.0rc2 fix attribution on top of 9/18–9/19. SGLang v0.5.20 (9/18, 190 commits / 713 PRs / 237 contributors) remains the window’s only formal major release; vLLM stable still v0.29.0 (09-08), release line v0.30.0rc2, 0.30 line only bugfix polish.

vLLM (main · 0.30 line NIXL robustness + transformers direct-run)

Version / security continuation:

SGLang (v0.5.20 · HRRN scheduling + TRT-LLM Blackwell + Mamba SSM + branch-point cache)

New features / major adaptations:

Under the Hood: #34565 Unified Prefix-Tree SWA Branch-Point Cache (Free Cost-Down for Agent / Multi-turn / RAG)

My read: this is the same DNA as 0919’s #37709 DSpark-under-PD (cross-node draft-KV reuse) and today’s performance column KV-aware routing (cross-replica prefix-block reuse) — all build reuse primitives on ‘already-computed capability.’ For my own OpenInfer / Qwen3-4B DFlash, the next step is to test whether ‘branch-point cache + draft-KV reuse’ reproduces the same E[L] lift on a single card.

Standing Topics

PD disaggregation: vLLM #57570 purifies NIXL counting; SGLang #37709 DSpark-under-PD + #28403 role hot-switch lead on two lines; /v1/responses not persisted by default (PD must not be enabled).

Architecture evolution: SGLang v0.5.20 multi-pronged (prefix tree + branch point, sampling masks overlap, TRT-LLM Blackwell, Mamba SSM, HRRN, Simulator); vLLM 0.30 polish.

PyTorch vs transformers: no ‘off transformers’ PR; boundary signal continues (model layer transformers, hot operators in-house).

Step adaptation: vLLM stepfun37 (MTP k=3) > SGLang dev (EAGLE draft=4), both NVFP4 4 cards + FP8 KV; StepAudio 3 edge-cloud cohere; Step IPO enters review stage (media wording, no public prospectus seen).

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

NVIDIA GR00T N2 / DreamZero World Action Model (WAM) upgrades the robot brain from VLA to ‘simulate the future in its head before acting’ — 2× success on new environments, world model from optional module to standard, edge-compute bar raised bullish for robot chips; the KDA operator makes linear attention beat MLA on every task with KV −75% and 6× decode at 1M context, and NVIDIA Dynamo KV-aware routing measured TTFT −50% with 89% hit on 4 replicas forms a trio with PD disaggregation / KV Offload; on the industry side Tesla Optimus starts a new China supply-chain audit, UBTECH’s Liuzhou factory is in production, and Digua Robot’s $400M Series C is bullish for 09660.HK — the embodied ‘smarter brain + real landing’ keeps getting a price tag.

Paper Core (NVIDIA GR00T N2 / DreamZero World Action Model)

Operator Deep-Dive: KDA (Kimi Delta Attention, Moonshot arXiv:2510.26692)

Performance Optimization: KV Cache-Aware Routing (prefix-aware routing, NVIDIA Dynamo)

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

On the papers side, NVIDIA GR00T N2 / DreamZero World Action Model (WAM) upgrades the robot brain from VLA to ‘simulate the future in its head before acting’ — ~2× SOTA VLA success on new environments, world model to standard, edge-compute bar raised bullish for robot chips; the KDA operator makes linear attention beat MLA on every task with KV −75% and 6× decode at 1M context, and NVIDIA Dynamo KV-aware routing measured TTFT −50% with 89% hit on 4 replicas forms a trio with PD disaggregation / KV Offload; on the framework side SGLang v0.5.20 fills in details — HRRN scheduling (GLM-5.2 trace ~69% lower TTFT than FCFS), TRT-LLM Blackwell 1.2× prefill / 1.45× decode over FlashMLA, first Mamba SSM support, and #34565 branch-point cache measured TTFT −32%; vLLM stable stays v0.29.0 carrying CVE-2026-93436 (fixed #55677, needs 0.29.1+) with a corrected rc2 attribution to #57570 — conclusion unchanged: don’t sit on v0.29.0, follow main or wait for 0.29.1+; on the industry side Tesla Optimus starts a new China supply-chain audit, UBTECH’s Liuzhou factory is in production, and Digua Robot’s $400M Series C is bullish for 09660.HK — the embodied ‘smarter brain + real landing’ keeps getting a price tag.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。