系列:每日AI热点

Daily AI Hotspot · 2026-09-13: vLLM HiSparse Trio Turns Host Memory into a GPU VRAM Extension Layer, SGLang graph-pool Makes CUDA Graph VRAM a Borrowable Pool

★ Most Worth Your Attention Today

vLLM’s HiSparse trio (#53781 + #56061 + #56629) pushes the “not enough VRAM” fix from “save it” to “tier and pool” — sparse MLA decode turns pinned host memory into a GPU VRAM extension layer for the first time, with full GPU-side resolution, CUDA Graph replay compatibility, and one host pool shared across TP ranks, directly lifting the long-context concurrency ceiling; SGLang’s graph-pool four-in-a-row (#39176~#39180) turns CUDA Graph VRAM from “per-graph exclusive” into a “borrowable pool.” All three frameworks have no new tag (vLLM v0.29.0 · 09-09 / SGLang v0.5.19 · 09-05 / Step frozen); all increments come from main.

Three layers of fact:

  1. vLLM HiSparse (#53781, merged 09-12, +14373/−1017, 124 files): sparse MLA decode adds a host cache, prefers device, and only overflows to pinned host memory when GPU capacity must be reclaimed (supersedes #46326 “always host-resident”), relaxing the decode VRAM ceiling from GPU capacity to host RAM.
  2. Cross-rank sharing + metrics (#56629 + #56061): an mmap-backed pool shares one host cache across local TP ranks, dropping host cache footprint to ~1/TP; KV connector stats expose cache metrics for observability.
  3. SGLang graph-pool four-in-a-row (09-12~09-13): #39176 reuses live graph executables, #39177 folds borrow state into the runtime + pre-cuts reusable segments + lifecycle checks (less fragmentation), #39178 borrow-capacity checks, #39180 borrows must be on their own allocation stream (cross-stream strands segments → OOM) — CUDA Graph VRAM goes from “per-graph exclusive” to “borrowable pool.”

Actionable conclusion: the three-tier resolve + LRU “pull only the selected rows on demand” is isomorphic to my own OpenInfer / Qwen3-4B DFlash speculative decoding — the lessons of cache boundaries are directly borrowable: draft/verify KV can be tiered the same way. On the deployment side, watch the VRAM ceiling freed by HiSparse + graph-pool for long-context high concurrency; validate parity on staging before upgrading (both are main-branch increments, not tags, so stability is unproven).

Worth saying separately: the lead swings back from “race to release” (0907/0910/0912 consecutive tags) to “VRAM engineering” this week — no new tags, but both KV and CUDA Graph VRAM are now made into “tiered / borrowable” resource pools. The moat of an inference framework increasingly lands on “VRAM boundary management.”

2. vLLM & SGLang Community Tracking

Version status: vLLM v0.29.0 (09-09) remains the latest stable with no new tag this window; SGLang v0.5.19 (09-05) remains the latest stable with no new tag; Step frozen. All increments this round come from main, with no formal release.

vLLM (main · VRAM-tiering main line)

New features / architecture evolution:

Breaking changes:

ChangeImpact
HiSparse supersedes always-host-resident (#53781)Sparse MLA decode behavior changes; old config must re-verify VRAM and hit rate
MoE all-reduce fast path for PCP+DCP (#56157)Parallel path change; re-verify parity for GLM-5.3 NVFP4 etc.
DSv4 warmup series (#50178/#56323/#53566)Cold-start path change; watch first-request latency

SGLang (main · graph-VRAM-pooling main line)

New features / major adaptation:

Production practice: after graph-pool makes CUDA Graph VRAM a borrowable pool, multi-graph concurrency’s VRAM fragmentation and exclusive waste drop, and long-context steady-state throughput is smoother (main-branch increment, not yet in the v0.5.19 tag).

Under the Hood: HiSparse Three-Tier Resolution (#53781 source comments)

HiSparse uses a fused CUDA resolver to map each top-k selected position to: ① a regular resident GPU page; ② rows already in the request’s own GPU hot buffer; ③ the pinned host pool. The order matters — a resident hit returns before the hot-cache lookup (no needless host pull), a hot hit updates the GPU-side LRU, and a miss only gathers the selected host rows to replace an LRU entry. Resolution / replacement / LRU update all stay on the GPU side and are CUDA Graph replay-compatible; the active indexer cache stays device-resident and is not tiered.

Insight: in sparse attention, top-k itself is a “working-set descriptor” — pull only the selected rows on demand, not the whole page. This is isomorphic to my own OpenInfer (Qwen3-4B speculative decode, sparse validation set, VRAM-sensitive): the three-tier resolve + LRU are directly borrowable for draft/verify KV. It shares roots with V4.1’s “bounded recompute for less cache”: both are Pareto tradeoffs on the cache boundary.

My read: this week’s main line is “VRAM pooling / tiering.” HiSparse turns host into a GPU extension layer, graph-pool turns CUDA Graph VRAM into a borrowable pool — the decisive factor in inference frameworks increasingly lands on “VRAM boundary management” rather than simply stacking operators.

Standing Topics

PD disaggregation: SGLang #36651 (Qwen3.8-Next PD state transfer) + #39190 (Kimi-K3 DCP under HiCache L1+L2); vLLM HiSparse is the VRAM extension of PD’s “D” segment, complementing (not replacing) external KV pooling (Mooncake) → the D segment is being split into device / host tiers.

Architecture evolution: this window’s main line = VRAM pooling / tiering (HiSparse three-tier + graph-pool); second = systematic compile warmup (vLLM DSv4 warmup) + Rust frontend / TreeCore completion.

PyTorch vs transformers: no new “leave transformers” PR this window; still tiered — model definitions rely on transformers, data plane & runtime rely on in-house.

Step adaptation: no new adaptation near-window; the only new signal = the Steptron training framework open-sourced (SteptronOss, 09-04, 587★). Inference side still recommends vLLM stepfun37 + MTP first; if MTP is a dependency, first test whether --speculative-config '{"method":"mtp"}' can launch.

3. AI Papers & Industry Hotspots

Highlight (one line)

Zhiyuan’s GO-1 uses ViLLA latent action to free action knowledge from “scarce real-world labeled data” to “massive unlabeled video,” giving VLA its first cross-embodiment transfer capability; while AdaLN-Zero and CUDA Graphs complete the “deterministic-latency execution” chain, exactly matching today’s embodied-industry “act” main line.

Paper core (GO-1 / ViLLA · Zhiyuan × Shanghai AI Lab, released 2025-03-10, open-sourced 2025-09-19)

Operator explainer: AdaLN-Zero

Performance optimization: CUDA Graphs (graph capture · eliminate kernel launch overhead)

Industry hotspots (embodied-intelligence companies / chain speed · pinned)

⚠️ Industry dynamics do not constitute investment advice.

4. The One-Line Takeaway

vLLM’s HiSparse trio (#53781+#56061+#56629) turns pinned host memory into a GPU VRAM extension layer for sparse MLA decode for the first time — full GPU-side resolution, CUDA Graph replay-compatible, one host pool shared across TP ranks, directly lifting the long-context concurrency ceiling; SGLang’s graph-pool four-in-a-row turns CUDA Graph VRAM from “per-graph exclusive” into a “borrowable pool”; both are main-branch increments with no new tag, and the main line is “VRAM pooling/tiering” rather than releases — the inference-framework moat increasingly lands on VRAM boundary management. On the research side GO-1 uses ViLLA latent action to migrate action knowledge across embodiments (5 task types +32%), AdaLN-Zero and CUDA Graphs complete the “deterministic-latency execution” chain, while UBTECH’s ¥150.8M win, Unitree’s −55% retrace, and embodied financing 5× YoY each tag the “see + act” embodied thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。