★ Most Worth Your Attention Today
vLLM’s HiSparse trio (#53781 + #56061 + #56629) pushes the “not enough VRAM” fix from “save it” to “tier and pool” — sparse MLA decode turns pinned host memory into a GPU VRAM extension layer for the first time, with full GPU-side resolution, CUDA Graph replay compatibility, and one host pool shared across TP ranks, directly lifting the long-context concurrency ceiling; SGLang’s graph-pool four-in-a-row (#39176~#39180) turns CUDA Graph VRAM from “per-graph exclusive” into a “borrowable pool.” All three frameworks have no new tag (vLLM v0.29.0 · 09-09 / SGLang v0.5.19 · 09-05 / Step frozen); all increments come from main.
Three layers of fact:
- vLLM HiSparse (#53781, merged 09-12, +14373/−1017, 124 files): sparse MLA decode adds a host cache, prefers device, and only overflows to pinned host memory when GPU capacity must be reclaimed (supersedes #46326 “always host-resident”), relaxing the decode VRAM ceiling from GPU capacity to host RAM.
- Cross-rank sharing + metrics (#56629 + #56061): an mmap-backed pool shares one host cache across local TP ranks, dropping host cache footprint to ~1/TP; KV connector stats expose cache metrics for observability.
- SGLang graph-pool four-in-a-row (09-12~09-13): #39176 reuses live graph executables, #39177 folds borrow state into the runtime + pre-cuts reusable segments + lifecycle checks (less fragmentation), #39178 borrow-capacity checks, #39180 borrows must be on their own allocation stream (cross-stream strands segments → OOM) — CUDA Graph VRAM goes from “per-graph exclusive” to “borrowable pool.”
Actionable conclusion: the three-tier resolve + LRU “pull only the selected rows on demand” is isomorphic to my own OpenInfer / Qwen3-4B DFlash speculative decoding — the lessons of cache boundaries are directly borrowable: draft/verify KV can be tiered the same way. On the deployment side, watch the VRAM ceiling freed by HiSparse + graph-pool for long-context high concurrency; validate parity on staging before upgrading (both are main-branch increments, not tags, so stability is unproven).
Worth saying separately: the lead swings back from “race to release” (0907/0910/0912 consecutive tags) to “VRAM engineering” this week — no new tags, but both KV and CUDA Graph VRAM are now made into “tiered / borrowable” resource pools. The moat of an inference framework increasingly lands on “VRAM boundary management.”
2. vLLM & SGLang Community Tracking
Version status: vLLM v0.29.0 (09-09) remains the latest stable with no new tag this window; SGLang v0.5.19 (09-05) remains the latest stable with no new tag; Step frozen. All increments this round come from main, with no formal release.
vLLM (main · VRAM-tiering main line)
New features / architecture evolution:
- HiSparse three-tier host cache (#53781): sparse MLA decode adds a pinned host overflow layer, superseding #46326 “always host-resident” — overflow only when GPU capacity must be reclaimed, relaxing sparse MLA decode VRAM ceiling from GPU capacity to host RAM.
- Cross-TP-rank host cache sharing (#56629): mmap-backed pool, rank0 writes, peers wait via IPC event, host cache footprint drops to ~1/TP; #56061 exposes cache metrics via KV connector stats.
- Sparse MLA PCP + DCP (#56157): GB200×4 / GLM-5.3 NVFP4 / TP1·PCP4·DCP4, new MoE all-reduce fast path, dropping the per-PCP decode-row copy.
- Rust frontend completion: #54821/#56378/#56386/#55047; scaling/warmup path fixes #56610 (ROCm elastic EP deadlock), #56526, #54416, #56452; DSv4 warmup series #50178/#56323/#53566 (cold-start compression).
Breaking changes:
| Change | Impact |
|---|---|
| HiSparse supersedes always-host-resident (#53781) | Sparse MLA decode behavior changes; old config must re-verify VRAM and hit rate |
| MoE all-reduce fast path for PCP+DCP (#56157) | Parallel path change; re-verify parity for GLM-5.3 NVFP4 etc. |
| DSv4 warmup series (#50178/#56323/#53566) | Cold-start path change; watch first-request latency |
SGLang (main · graph-VRAM-pooling main line)
New features / major adaptation:
- graph-pool four-in-a-row (09-12~09-13): #39176 reuses live graph executables, #39177 folds borrow state into runtime + pre-cuts reusable segments + lifecycle checks (less fragmentation), #39178 borrow-capacity checks, #39180 borrows must be on their own allocation stream (cross-stream strands segments → OOM) — CUDA Graph VRAM goes from “per-graph exclusive” to “borrowable pool.”
- Cache Rust-ification: #37584 ports SWA branch-point cache to Rust TreeCore; #38482 keeps aux LRU recency on node split.
- Default change: #39165
--dcp-comm-backendresolves per-model to fi_a2a/a2a by default. - PD / speculative / new models: #36651 Qwen3.8-Next PD state transfer; #39190 Kimi-K3 DCP under HiCache L1+L2 with DSPARK; #39154 DSpark per-step metadata on device; #34556 supports Mamba 2 & 1; GLM-5.3-Flash cookbook (#39213, MTP 5/1/6, EP1+flashinfer_trtllm on Blackwell); Responses API extensions (custom tools / encrypted inference replay).
Production practice: after graph-pool makes CUDA Graph VRAM a borrowable pool, multi-graph concurrency’s VRAM fragmentation and exclusive waste drop, and long-context steady-state throughput is smoother (main-branch increment, not yet in the v0.5.19 tag).
Under the Hood: HiSparse Three-Tier Resolution (#53781 source comments)
HiSparse uses a fused CUDA resolver to map each top-k selected position to: ① a regular resident GPU page; ② rows already in the request’s own GPU hot buffer; ③ the pinned host pool. The order matters — a resident hit returns before the hot-cache lookup (no needless host pull), a hot hit updates the GPU-side LRU, and a miss only gathers the selected host rows to replace an LRU entry. Resolution / replacement / LRU update all stay on the GPU side and are CUDA Graph replay-compatible; the active indexer cache stays device-resident and is not tiered.
Insight: in sparse attention, top-k itself is a “working-set descriptor” — pull only the selected rows on demand, not the whole page. This is isomorphic to my own OpenInfer (Qwen3-4B speculative decode, sparse validation set, VRAM-sensitive): the three-tier resolve + LRU are directly borrowable for draft/verify KV. It shares roots with V4.1’s “bounded recompute for less cache”: both are Pareto tradeoffs on the cache boundary.
My read: this week’s main line is “VRAM pooling / tiering.” HiSparse turns host into a GPU extension layer, graph-pool turns CUDA Graph VRAM into a borrowable pool — the decisive factor in inference frameworks increasingly lands on “VRAM boundary management” rather than simply stacking operators.
Standing Topics
PD disaggregation: SGLang #36651 (Qwen3.8-Next PD state transfer) + #39190 (Kimi-K3 DCP under HiCache L1+L2); vLLM HiSparse is the VRAM extension of PD’s “D” segment, complementing (not replacing) external KV pooling (Mooncake) → the D segment is being split into device / host tiers.
Architecture evolution: this window’s main line = VRAM pooling / tiering (HiSparse three-tier + graph-pool); second = systematic compile warmup (vLLM DSv4 warmup) + Rust frontend / TreeCore completion.
PyTorch vs transformers: no new “leave transformers” PR this window; still tiered — model definitions rely on transformers, data plane & runtime rely on in-house.
Step adaptation: no new adaptation near-window; the only new signal = the Steptron training framework open-sourced (SteptronOss, 09-04, 587★). Inference side still recommends vLLM stepfun37 + MTP first; if MTP is a dependency, first test whether --speculative-config '{"method":"mtp"}' can launch.
3. AI Papers & Industry Hotspots
Highlight (one line)
Zhiyuan’s GO-1 uses ViLLA latent action to free action knowledge from “scarce real-world labeled data” to “massive unlabeled video,” giving VLA its first cross-embodiment transfer capability; while AdaLN-Zero and CUDA Graphs complete the “deterministic-latency execution” chain, exactly matching today’s embodied-industry “act” main line.
Paper core (GO-1 / ViLLA · Zhiyuan × Shanghai AI Lab, released 2025-03-10, open-sourced 2025-09-19)
- One-line positioning: real-world action-labeled data is scarce (a million trajectories is a ceiling) while internet unlabeled video is massive; GO-1 upgrades VLA to ViLLA (Vision-Language-Latent-Action) — inserting a Latent Action layer between perception and execution so action knowledge can migrate across embodiments from unlabeled video.
- Core idea ① (three modules): VLM (InternVL-2B) + MoE dual experts (Latent Planner implicit planner / Action Expert diffusion action expert), sharing a Transformer backbone with independent FFN and Q/K/V/O.
- Core idea ② (LAM latent action model): encoder Spatio-temporal Transformer + temporal causal mask, decoder Spatial Transformer (initial frame + discrete token → reconstruct target frame), token quantized via VQ-VAE.
- Core idea ③ (data base AgiBot World): 4000㎡, five domains, 100+ scenes, >1M trajectories, 30Hz, ~1% failure-recovery trajectories.
- Effect: 5 task types average success rate 46%→78% (+32%); ablation adding Latent Planner +12% (66%→78%).
Operator explainer: AdaLN-Zero
- Positioning:
AdaLN(h|c)=γ(c)⊙LayerNorm(h)+β(c), where γ/β are dynamically generated by an MLP from the condition vector c and modulate per-channel (vs traditional LayerNorm’s globally-learned, condition-independent scale/shift). - Zero key: the MLP also outputs a gate α, and the γ/β/α three linear layers are zero-initialized → when α=0 the residual branch outputs 0 and the network starts as an identity mapping, so the early-training signal is clean and deep stacking does not diverge.
- Advantage: O(N) element-wise vs Cross-Attn O(N×M), cheap, stable, easy to deepen; used in DiT / SD3-MMDiT / Diffusion Policy / RDT-1B / VLA action heads.
Performance optimization: CUDA Graphs (graph capture · eliminate kernel launch overhead)
- Problem: each GPU step needs the CPU to launch hundreds/thousands of kernels one by one (single ~5–10μs); autoregressive decode produces only 1 token per step → CPU becomes the bottleneck, GPU idles (launch gap), latency & jitter rise.
- Solution: stream-capture the repeating compute sequence (whole forward / one decode step) as a CUDA Graph, replay once to rerun all kernels.
- Hard constraints: static shape / static VRAM address (fixed batch & seq len → vLLM/SGLang align batch to padding buckets); no dynamic control flow; PagedAttention’s dynamic page table needs a fixed-address variant; common “capture big, release small” (only capture small batches).
- Benefit: decode latency down 10–30%, jitter drops sharply; vLLM V1 / SGLang / TensorRT-LLM on by default. Embodied link: a control loop demands deterministic latency, and de-jittering is the key to a real-time closed loop — exactly the pit SGLang’s graph-pool “cross-stream strand → OOM” is designed to avoid.
Industry hotspots (embodied-intelligence companies / chain speed · pinned)
- [Embodied] UBTECH (09880.HK): on 9/13 won the Leshan commercial-service humanoid project ~¥150.8M; on 9/12 the Liuzhou “10k-unit-level industrial humanoid super smart factory” went into production (world’s first adapted to 10k-unit capacity, ~1 unit off the line every 10 min, annual planned capacity >10k units). H1 revenue ¥1.27B (+104.2%), humanoid total sales 16,123 units (+268.3%), full-size embodied revenue ¥590M (+1445%).
- [Embodied] Zhiyuan / Shangwei New Material (688585·A-share): AGILE 2.0 sense-control integration (Lingxi X2 ball-walking / fire-ring drilling / long-rope jumping; ball-walking is the industry’s first bipedal autonomous); with Chimelong building the world’s first large embodied-intelligence theme park; with Minth/Serbia building a factory with 5000+ units/yr capacity.
- [Embodied] Unitree (688836·A-share): closed at ¥477 on 9/11, mcap
¥192.978B, vs the 8/19 first-day high (¥1100) retraced nearly 55%; 2025 revenue >70% from research & education, industry only ~9%. Regulators tightening the bar for listing humanoids. - [Chain] XPeng IRON: on 9/8 the world’s first high-end general humanoid automated production line was enabled, first fully-automatic final assembly off the line, core-process automation rate >80%.
- [Primary]: Chaowei Power (Kinetix AI) >¥500M angel+ round (KAI Hand 37 DOF, fingertip force >30N already on global sale); Xieyue Intelligence tens-of-millions angel+ round (home-scenario general embodied base model); Zilian Robot confidential HKEX filing (valuation >¥20B). H1 domestic embodied financing ~¥93.5B / 322 deals (YoY ~5×).
- [Chain]: Lingyi Zhizao × Zhiyuan JV “Lingzhi Innovation” first embodied-intelligent line (thousands of whole-machine assemblies); Baolong Tech won robot rotary-joint encoder designation (2026-12 mass production, ±10 arcsec); dexterous-hand annual shipments break 10k units, tactile penetration 20%→60%. Maps to: Sanhua 002050 / Top Group 601689 / Leader 688017 / robot ETF 562500 / Shangwei 688585.
- [General]: Anthropic 9/12 “We Must Pace the Frontier” advocates slowing frontier models + employee-level third-party embedded evaluation; State Council 9/11 deploys a compute network (green-power direct connect / source-grid-load-storage); Google >$1.5B acquires AI-coding startup Mechanize; global humanoid H1 shipments break 22k units (YoY near +300%), MIIT expects full-year whole-machine output to break 100k units.
⚠️ Industry dynamics do not constitute investment advice.
4. The One-Line Takeaway
vLLM’s HiSparse trio (#53781+#56061+#56629) turns pinned host memory into a GPU VRAM extension layer for sparse MLA decode for the first time — full GPU-side resolution, CUDA Graph replay-compatible, one host pool shared across TP ranks, directly lifting the long-context concurrency ceiling; SGLang’s graph-pool four-in-a-row turns CUDA Graph VRAM from “per-graph exclusive” into a “borrowable pool”; both are main-branch increments with no new tag, and the main line is “VRAM pooling/tiering” rather than releases — the inference-framework moat increasingly lands on VRAM boundary management. On the research side GO-1 uses ViLLA latent action to migrate action knowledge across embodiments (5 task types +32%), AdaLN-Zero and CUDA Graphs complete the “deterministic-latency execution” chain, while UBTECH’s ¥150.8M win, Unitree’s −55% retrace, and embodied financing 5× YoY each tag the “see + act” embodied thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。