系列:每日AI热点

Daily AI Hotspot · 2026-09-17: vLLM #56935 Makes DSv4.1 Fused Mega Attention + NVFP4 Compressed KV SM100 Default; π*0.6/RECAP Shifts Robot Deployment from 'Model Strength' to 'Real Deploy Data'

★ Most Worth Your Attention Today

vLLM #56935 (merged 09-16) — wires DeepSeek-V4.1’s FlashMLA mega attention + NVFP4 compressed KV into one path and sets it SM100 default: the first change to make both a “fused kernel” and “lower-precision KV” a default path at once (compressed record 45% smaller than fp8_ds_mla), and it ships rare end-to-end precision evidence (under DSpark k=5 true rejection sampling, gsm8k 0.9318 vs 0.9265, gpqa 0.9053 vs 0.9066), with three boundaries clearly drawn — noticeably faster at TP1, prefill untested, no effect off-SM100. This is the cleanest signal on the framework side today: it hands the KV-capacity accounting back to the hardware default path.

It is actually the same thread as today’s papers headline: the robot-deployment bottleneck is shifting from “is the model strong enough” to “can we keep collecting real deploy data” — π*0.6/RECAP produces deployable policies from “pretrained general VLA + task-level real-robot RL fine-tune” (hardest-task throughput ×2+, failure rate ~halved, 13 hours of continuous coffee making), proving the model side is already sufficient and the constraint is now data; while the industry’s “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders” is the other face of that same bottleneck. The framework thins the KV ledger, the robot thickens the deploy-data ledger — both are closing moves that push existing capability into usable form.

Three layers of fact:

  1. Fused kernel + lower-precision KV default at once (#56935): mega-attention compresses “Q RoPE → sparse attention → inverse RoPE → FP8 quantize” into a single kernel launch, writing straight into the buffer wo_a consumes; the NVFP4 record is 288 B/token (256 B e2m1 + 32 e4m3 scales, group 16), 45% smaller than fp8_ds_mla. Not “a format swap” — “move scheduling cost from Python side into the kernel.”
  2. Precision cost negligible (first end-to-end evidence): GB300 / CUDA Graph / decode, mega vs sparse microseconds, TP1 s_q=512 is 96/140 (1.45×), flat from TP2 up — the deciding variable is live heads / padded heads, not batch size; draft acceptance is nearly identical, so the fused kernel + smaller compressed record carry no measurable precision loss.
  3. PD disaggregation verifiable for the first time (SGLang #39500): adds an optional KV-transfer Adler-32 checksum (--disaggregation-enable-kv-checksum, off by default) that aborts and counts a request on mismatch — PD goes from “runs” to “verifiable”; turn it on during troubleshooting, off for steady-state stress tests.

Actionable conclusion: if you track the DeepSeek-V4.1 deploy stack, #56935 is the signal to recompute your KV-capacity math on B200/GB300 — “45% less KV” directly rewrites the concurrency ceiling assumption; meanwhile SGLang’s six-PR chain pushes multimodal V4.1 ahead of vLLM, and the PD checksum is a must-have production troubleshooting switch. For my own OpenInfer / Qwen3-4B DFlash the direction is consistent: make compression / KV-lifecycle a default path rather than widening the draft tree.

Worth saying separately: all four vLLM bugfixes this issue (#57152/#57132/#56930/#57104) sit on the “V4.1 + speculative decode + PD” cross-path — causal image SWA, ROCm V4 precision collapse, EAGLE/dense draft failing under EP, KV Connector + MTP deadlock under KV pressure. The conclusion is blunt: don’t sit on v0.29.0, follow main or wait for v0.29.1.

2. vLLM & SGLang Community Tracking

Version status: vLLM stable v0.29.0 (09-09), next candidate v0.29.1rc0 (09-13); SGLang v0.5.19 (09-05) remains the latest stable, no new tag in-window. This period’s increments come mainly from the main branch, with no formal release (same cadence as 0916).

vLLM (main · DSv4.1 deep water + cross-path bugfix)

New features / architecture evolution:

Bug fixes (cross-path, follow main): #57152 (V4.1 causal image SWA), #57132 (ROCm V4 precision collapse), #56930 (EAGLE/dense draft failing under EP), #57104 (KV Connector + MTP deadlock under KV pressure) — all four on the “V4.1 + speculative decode + PD” intersection.

SGLang (v0.5.19 · multimodal V4.1 six-PR chain + verifiable PD)

New features / major adaptations:

Other: #38526 supports Ling-3.0-flash-VL, #37810 enables breakable CUDA graph prefill for DSV4 on ROCm, #39875 fixes AMD dsv4 server startup — AMD-side V4 availability fills in fast.

Under the Hood: #56935’s Mega-Attention Fused Kernel + NVFP4 Compressed KV (Thin the KV Ledger, Not Just a Format Swap)

My read: the point of #56935 is not the “45% less KV” number but that it makes “fused kernel + low-precision KV” a default path — moving Python-side scheduling cost into the kernel, fundamentally the same DNA as 0913’s “VRAM tiered pooling” and 0916’s “DFlash2 pushes E[L] up”: all hunt the optimal operating point on a tradeoff surface. For my own OpenInfer / Qwen3-4B DFlash, the next step is to empirically test whether a “mega-attention-style fused kernel + NVFP4 compressed KV” can reproduce a 45% KV cut on Qwen3-4B with no precision loss.

Standing Topics

PD disaggregation: SGLang #39500 adds an optional checksum to KV transfer; vLLM #57077 fixes HiSparse cross-logical-block pull alignment, #54222 flows hit rate into token line-items. The 09-15 community review flags an unresolved contradiction: chunked prefill is on by default at both vendors, critics call it a local optimum, but the “context length × concurrency” cross-curve has still never been published.

Architecture evolution: KV numeric format keeps dropping (vLLM NVFP4 288 B/token, SGLang FP4 packing + MXFP8 scale backup); fused kernels replace multi-kernel orchestration (moving Python-side scheduling cost into the kernel); speculative decode is deeply coupled with scheduling / KV lifecycle (all four bugfixes this round sit at that intersection).

PyTorch vs transformers: no new “off transformers” PR/discussion this cycle (151 vLLM and 132 SGLang merges, none). The boundary holds — the model layer still uses transformers, only hot operators like MoE/RoPE/FP4 packing get in-house kernels.

Step adaptation: no framework-side Step merge in the last 72h, reuse the standing conclusion (vLLM stepfun37 + MTP first-class > SGLang dev image + EAGLE). On the industry side, 09-16 CEO Jiang Daxing laid out a finance-vertical “FDE + SaS” paradigm and edge lead Yu Gang talked AI 2.0 generalization and quality; StepAudio 3’s five models are live, only Realtime over WebSocket, the rest over HTTP.

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

Physical Intelligence’s π*0.6 / RECAP (arXiv:2511.14759) proves that “pretrained general VLA + task-level real-robot RL fine-tune” yields deployable policies — the robot-deployment bottleneck is shifting from “is the model strong enough” to “can we keep collecting real deploy data,” the mirror image of today’s industry “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders”; Selective Scan / S6 (Mamba) uses O(N) associative scan to push inference KV cache down to a few MB and elastic inference lifts GPU utilization from ~20% to 60%+, rounding out the “edge brain + real-robot deploy” cost model from the algorithm layer down to the system layer — and UBTECH’s 10k-unit factory, Zhiyuan’s A3 Ultra thousand-unit mass production, Unitree’s 57% market-cap drawdown, Qianxun’s three rounds >¥4.5B, and the regulator’s “bid ≠ commercial revenue” call are each putting a price tag on the embodied “smarter brain + real landing” stack.

Paper Core (π*0.6 / RECAP · Physical Intelligence, arXiv:2511.14759)

Operator Deep-Dive: Selective Scan / S6 (Mamba)

Performance Optimization: Elastic Inference / Serverless LLM Serving (Scale-to-Zero)

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

On the framework side, vLLM #56935 (merged 09-16) wires DeepSeek-V4.1’s FlashMLA mega attention + NVFP4 compressed KV into one path and sets it SM100 default — the first change to make both a “fused kernel” and “lower-precision KV” a default path at once (45% less KV, 1.45× faster at TP1, no precision loss), with three boundaries clearly drawn; SGLang lands a six-PR multimodal V4.1 chain and PD disaggregation gets an optional Adler-32 checksum (#39500, from “runs” to “verifiable”); four cross-path bugfixes say don’t sit on v0.29.0, follow main or wait for v0.29.1. On the papers side, π*0.6/RECAP (Physical Intelligence, arXiv:2511.14759) proves “pretrained general VLA + task-level real-robot RL fine-tune” is deployable (throughput ×2+, failure rate halved, 13 hours of coffee), shifting the robot bottleneck from model strength to real deploy data, the same thread as the industry’s “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders”; Selective Scan/S6 (Mamba) uses O(N) associative scan to push inference KV cache to a few MB and elastic inference lifts GPU utilization from ~20% to 60%+, rounding out the “edge brain + real-robot deploy” cost model from algorithm to system — and UBTECH’s 10k-unit factory, Zhiyuan’s A3 Ultra thousand-unit production, Unitree’s 57% drawdown, Qianxun’s three rounds >¥4.5B, and the regulator’s “bid ≠ commercial revenue” are each putting a price tag on the embodied “smarter brain + real landing” stack.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。