★ Most Worth Your Attention Today
vLLM #56935 (merged 09-16) — wires DeepSeek-V4.1’s FlashMLA mega attention + NVFP4 compressed KV into one path and sets it SM100 default: the first change to make both a “fused kernel” and “lower-precision KV” a default path at once (compressed record 45% smaller than fp8_ds_mla), and it ships rare end-to-end precision evidence (under DSpark k=5 true rejection sampling, gsm8k 0.9318 vs 0.9265, gpqa 0.9053 vs 0.9066), with three boundaries clearly drawn — noticeably faster at TP1, prefill untested, no effect off-SM100. This is the cleanest signal on the framework side today: it hands the KV-capacity accounting back to the hardware default path.
It is actually the same thread as today’s papers headline: the robot-deployment bottleneck is shifting from “is the model strong enough” to “can we keep collecting real deploy data” — π*0.6/RECAP produces deployable policies from “pretrained general VLA + task-level real-robot RL fine-tune” (hardest-task throughput ×2+, failure rate ~halved, 13 hours of continuous coffee making), proving the model side is already sufficient and the constraint is now data; while the industry’s “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders” is the other face of that same bottleneck. The framework thins the KV ledger, the robot thickens the deploy-data ledger — both are closing moves that push existing capability into usable form.
Three layers of fact:
- Fused kernel + lower-precision KV default at once (#56935): mega-attention compresses “Q RoPE → sparse attention → inverse RoPE → FP8 quantize” into a single kernel launch, writing straight into the buffer
wo_aconsumes; the NVFP4 record is 288 B/token (256 B e2m1 + 32 e4m3 scales, group 16), 45% smaller than fp8_ds_mla. Not “a format swap” — “move scheduling cost from Python side into the kernel.” - Precision cost negligible (first end-to-end evidence): GB300 / CUDA Graph / decode, mega vs sparse microseconds, TP1 s_q=512 is 96/140 (1.45×), flat from TP2 up — the deciding variable is live heads / padded heads, not batch size; draft acceptance is nearly identical, so the fused kernel + smaller compressed record carry no measurable precision loss.
- PD disaggregation verifiable for the first time (SGLang #39500): adds an optional KV-transfer Adler-32 checksum (
--disaggregation-enable-kv-checksum, off by default) that aborts and counts a request on mismatch — PD goes from “runs” to “verifiable”; turn it on during troubleshooting, off for steady-state stress tests.
Actionable conclusion: if you track the DeepSeek-V4.1 deploy stack, #56935 is the signal to recompute your KV-capacity math on B200/GB300 — “45% less KV” directly rewrites the concurrency ceiling assumption; meanwhile SGLang’s six-PR chain pushes multimodal V4.1 ahead of vLLM, and the PD checksum is a must-have production troubleshooting switch. For my own OpenInfer / Qwen3-4B DFlash the direction is consistent: make compression / KV-lifecycle a default path rather than widening the draft tree.
Worth saying separately: all four vLLM bugfixes this issue (#57152/#57132/#56930/#57104) sit on the “V4.1 + speculative decode + PD” cross-path — causal image SWA, ROCm V4 precision collapse, EAGLE/dense draft failing under EP, KV Connector + MTP deadlock under KV pressure. The conclusion is blunt: don’t sit on v0.29.0, follow main or wait for v0.29.1.
2. vLLM & SGLang Community Tracking
Version status: vLLM stable v0.29.0 (09-09), next candidate v0.29.1rc0 (09-13); SGLang v0.5.19 (09-05) remains the latest stable, no new tag in-window. This period’s increments come mainly from the main branch, with no formal release (same cadence as 0916).
vLLM (main · DSv4.1 deep water + cross-path bugfix)
New features / architecture evolution:
- #56935 (merged 09-16): DSv4.1’s FlashMLA mega attention + NVFP4 compressed KV, new backend
FLASHMLA_MEGA_ATTN_DSV41becomes SM100 default, compressed record 45% smaller than fp8_ds_mla; impact = recompute the KV-capacity ledger on B200/GB300. - Perf #57204 / #57140: drop MegaMoE padding and the shared-padding workaround, GDN scatters mixed speculative output back into the caller’s buffer; impact = MoE stops doing useless pad, long sequences save VRAM and compute.
- PD disaggregation #57077 / #54222: HiSparse aligns region-mapped pull to logical-block size under PD (previously silently pulled wrong KV); prefill worker hit rate flows into
prompt_tokens_details; impact = PD cost/benefit can finally be reconciled against token line-items. - HW/ecosystem #55934 / #57210 / #55557: ROCm triton 3.8 mxfp4 MoE (gpt-oss + DSv4), DeepEPv2 onto the SP support list, Qwen4Exp uses fp8_e4m3 main KV on the QSA path; impact = AMD-side MXFP4 MoE catches up to the NVIDIA route.
Bug fixes (cross-path, follow main): #57152 (V4.1 causal image SWA), #57132 (ROCm V4 precision collapse), #56930 (EAGLE/dense draft failing under EP), #57104 (KV Connector + MTP deadlock under KV pressure) — all four on the “V4.1 + speculative decode + PD” intersection.
SGLang (v0.5.19 · multimodal V4.1 six-PR chain + verifiable PD)
New features / major adaptations:
- DSv4.1 six-PR chain: #39652 compression/KV I-O/metadata kernels, #39656 RoPE and FP4 packing, #39657 Hopper FP8 matmul and tuning, #39664 mHC computation and compensation projection, #39668 vision tower and image preprocessing, #39671 candidate indexer library; impact = multimodal V4.1 currently moves ahead of vLLM in SGLang.
- PD disaggregation #39500 (high priority): adds an optional KV-transfer Adler-32 checksum (
--disaggregation-enable-kv-checksum, off by default, aborts and counts a request on mismatch); impact = PD goes from “runs” to “verifiable.” - KV architecture #37615 [kv-shard 2/4] Sharded pools (Blackwell / memory-pool / unified-radix-cache), #39089 HiCache backs up MXFP8 KV scales in the host pool; impact = paves the way for Blackwell big pools + unified Radix tree, KV offload/reload no longer drops precision metadata.
- Router productionization #39002 / #39016 / #38774: sampling contract 3/3 (splice injection without re-serialization), revoke readiness before SIGTERM so k8s drains the Pod, fix NIXL backend device-context init; impact = rolling releases drop no requests.
Other: #38526 supports Ling-3.0-flash-VL, #37810 enables breakable CUDA graph prefill for DSV4 on ROCm, #39875 fixes AMD dsv4 server startup — AMD-side V4 availability fills in fast.
Under the Hood: #56935’s Mega-Attention Fused Kernel + NVFP4 Compressed KV (Thin the KV Ledger, Not Just a Format Swap)
- Four steps in one: mega-attention compresses “Q RoPE → sparse attention → inverse RoPE → FP8 quantize” into a single kernel launch, writing straight into the buffer
wo_aconsumes; the new capability bitaccepts_unnormed_unroped_querylets the kernel do Q norm/RoPE itself, and at TP1 the 64 live heads exactly equal the kernel head count, so the Q-padding step is skipped outright. - One step, one buffer: output goes through QuantizedActivation (fp8 e4m3 + packed ue8m0 scales); prefill chunk and decode segment fill disjoint token ranges, one fp8_einsum covers all N tokens; the NVFP4 record is 288 B/token (256 B e2m1 + 32 e4m3 scales, group 16), 45% smaller than fp8_ds_mla.
- Perf (GB300 / CUDA graph / decode, mega vs sparse microseconds): TP1 s_q=512 is 96/140 (1.45×), flat from TP2 up — the deciding variable is live heads / padded heads, not batch size. Precision: under DSpark k=5 true rejection sampling gsm8k 0.9318 vs 0.9265, gpqa 0.9053 vs 0.9066, draft acceptance nearly identical, so the fused kernel + smaller compressed record carry no measurable precision cost. Untested: prefill unmeasured, SM100 only.
My read: the point of #56935 is not the “45% less KV” number but that it makes “fused kernel + low-precision KV” a default path — moving Python-side scheduling cost into the kernel, fundamentally the same DNA as 0913’s “VRAM tiered pooling” and 0916’s “DFlash2 pushes E[L] up”: all hunt the optimal operating point on a tradeoff surface. For my own OpenInfer / Qwen3-4B DFlash, the next step is to empirically test whether a “mega-attention-style fused kernel + NVFP4 compressed KV” can reproduce a 45% KV cut on Qwen3-4B with no precision loss.
Standing Topics
PD disaggregation: SGLang #39500 adds an optional checksum to KV transfer; vLLM #57077 fixes HiSparse cross-logical-block pull alignment, #54222 flows hit rate into token line-items. The 09-15 community review flags an unresolved contradiction: chunked prefill is on by default at both vendors, critics call it a local optimum, but the “context length × concurrency” cross-curve has still never been published.
Architecture evolution: KV numeric format keeps dropping (vLLM NVFP4 288 B/token, SGLang FP4 packing + MXFP8 scale backup); fused kernels replace multi-kernel orchestration (moving Python-side scheduling cost into the kernel); speculative decode is deeply coupled with scheduling / KV lifecycle (all four bugfixes this round sit at that intersection).
PyTorch vs transformers: no new “off transformers” PR/discussion this cycle (151 vLLM and 132 SGLang merges, none). The boundary holds — the model layer still uses transformers, only hot operators like MoE/RoPE/FP4 packing get in-house kernels.
Step adaptation: no framework-side Step merge in the last 72h, reuse the standing conclusion (vLLM stepfun37 + MTP first-class > SGLang dev image + EAGLE). On the industry side, 09-16 CEO Jiang Daxing laid out a finance-vertical “FDE + SaS” paradigm and edge lead Yu Gang talked AI 2.0 generalization and quality; StepAudio 3’s five models are live, only Realtime over WebSocket, the rest over HTTP.
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
Physical Intelligence’s π*0.6 / RECAP (arXiv:2511.14759) proves that “pretrained general VLA + task-level real-robot RL fine-tune” yields deployable policies — the robot-deployment bottleneck is shifting from “is the model strong enough” to “can we keep collecting real deploy data,” the mirror image of today’s industry “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders”; Selective Scan / S6 (Mamba) uses O(N) associative scan to push inference KV cache down to a few MB and elastic inference lifts GPU utilization from ~20% to 60%+, rounding out the “edge brain + real-robot deploy” cost model from the algorithm layer down to the system layer — and UBTECH’s 10k-unit factory, Zhiyuan’s A3 Ultra thousand-unit mass production, Unitree’s 57% market-cap drawdown, Qianxun’s three rounds >¥4.5B, and the regulator’s “bid ≠ commercial revenue” call are each putting a price tag on the embodied “smarter brain + real landing” stack.
Paper Core (π*0.6 / RECAP · Physical Intelligence, arXiv:2511.14759)
- One-line positioning: RECAP (RL with Experience and Corrections via Advantage-conditioned Policies) unifies demonstration + autonomous rollout + human correction — three heterogeneous data types — into one offline-RL loop; π*0.6 = π0.6 plus advantage conditioning.
- Three-generation evolution: π0.5 → π0.6 (more robot platforms / backbone swapped to Gemma 3 4B / action expert grown to 860M) → π*0.6 (advantage conditioning). Policy outputs 50 Hz action chunk + discrete subtask text ℓ̂.
- Distributed value function V: B = 201 discrete value bins, cross-entropy trained; “progress-based” reward (successful episode reward rises linearly to 0 with steps, failure gets a large negative constant) → V is essentially “how many steps from success,” normalized to (−1, 0); V backbone is only 670M.
- Advantage conditioning (core): I_t is an extra text token inserted after the subtask, before the action → affects only the action log-likelihood; training randomly drops I_t, inference sets I_t=True (β=1) or CFG; human-intervention actions force I_t=True. Loss Eq.3 pulls policy extraction from PPO online interaction back to offline weighted regression, sidestepping large-model PPO instability.
- Results: hardest-task throughput ×2+, failure rate ~halved; packing success 90%+; 13 hours of continuous coffee making, >2 hours folding unfamiliar clothes in a new home. Failures are mostly “policy times out” (slow actions) not wrong actions → what’s bought is mainly speed.
- Limitations: human-intervention quality not guaranteed, exploration greedy, limited help on tasks “demos can’t do”, not open-sourced (openpi ships only π0 / π0-FAST / π0.5).
Operator Deep-Dive: Selective Scan / S6 (Mamba)
- Positioning: makes the SSM’s Δ_t, B_t, C_t functions of the input x_t; S4 (LTI, foldable into one global convolution) → S6 (time-varying, gains content-aware gating).
- Engineering core: associative scan (associative parallel prefix scan) + hardware-aware CUDA kernel (discretize/scan/gate fused, hidden state stays in SRAM, never written back to HBM) — the role equivalent to FlashAttention in attention.
- Complexity: compute O(N) vs O(N²); each inference layer has a fixed-size hidden state, no KV cache (1M tokens: several GB → several MB); paper reports up to ~5× inference throughput.
- Weakness: associative-recall (MQAR) weak → long-doc QA / exact citation still leans Transformer; landing form is the hybrid architecture: Jamba, Nemotron-H, Granite 4.0, Zamba, Codestral/Falcon Mamba; domestic kin Qwen3-Next Gated DeltaNet, MiniMax Lightning Attention.
- Pit: judging only by perplexity misleads; the SSM:attention layer ratio is a hyperparameter (directly sets KV cache size); the training kernel (parallel scan) and inference kernel (single-step recurrence) are not the same and must be profiled separately.
Performance Optimization: Elastic Inference / Serverless LLM Serving (Scale-to-Zero)
- Standard checklist: quantization (GPTQ/AWQ/FP8), pruning, distillation, PagedAttention, continuous batching, LoRA, FlashAttention, CUDA Graphs, KV quantization.
- Frontier checklist: sparse MoE + EP, MLA/NSA, linear attention and SSM, speculative decode (MTP/EAGLE/Medusa/DSpark), PD disaggregation, BitNet 1-bit, KV quantization and offload, compile optimization (TorchCompile/FA-4), sequence & context parallel, Expert Offload, MoD, semantic cache.
- Deep points: GPU billed by the second + traffic tides → fixed always-on utilization is only 15–30%. Four things: ① SLO-aware scaling (by queue wait / P99 TTFT, not utilization); ② Scale-to-Zero; ③ cold-start elimination (loading 70B bf16 ≈140GB over 1GbE is ~100s → GPUDirect Storage / sequential read / weight snapshot to seconds; runtime CUDA context + graph capture + JIT → keep-warm pool / CRIU snapshot); ④ multi-model share one card (cache-aware routing, S-LoRA unified paging, off-peak co-residency).
- Most practical combo: stack with PD disaggregation (elastic preemptible prefill + guaranteed-resident decode), then add Spot + checkpoint resume. Payoff: cold start tens of seconds → sub-second to seconds; GPU utilization ~20% → 60%+; per-token cost down 2–5×.
- Embodied link: cloud “brain” elastic on demand + edge fixed-latency budget = the underlying variable in the robot-as-a-service cost model; today’s evidence = the services-fair compute-loan/compute-token-loan, DeepSeek’s Ulanqab ~1 GW (≥160k Ascend 950DT), Wuqin Core’s APXInf edge engine (as low as ~26 ms).
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] UBTECH 09880.HK: 10k-unit industrial-humanoid super-smart factory online, one unit rolls off every 10 minutes, annual capacity over 10k; 2026H1 revenue ¥1.269B (+104.2%), full-size humanoid revenue ¥590M became the #1 source (>70% gross), new orders >¥1B (Airbus / top automaker / State Grid); but U1 delivery guidance cut to 1500–2000 units, targeting Q4 single-quarter adjusted EBITDA positive → neutral-slightly-positive but兑现 discounted (capacity ≠ orders ≠ revenue).
- [Embodied] Zhiyuan → Shangwei 688585: 9/16 unveiled the A3 Ultra, the world’s first scaled-mass-production full-size humanoid commercial landing scene (4S store / hotel / convenience store), 700 TOPS + 20 DoF omnidirectional tactile dexterous hand + 360° perception + UWB/RTK, completed thousand-unit automotive-grade mass-production validation, 7×24 duty; AGILE 2.0 + GE-Act 2.0 together → bullish.
- [Embodied] Unitree 688836: one month post-IPO mcap ¥444.9B → ~¥190B (−57%, >¥220B evaporated); 2026H1 revenue ¥1.152B (+48.54%) but attributable-net −19.34%, research/education 73.6%, industrial logistics only ~9%; open-sourced UnifoLM-WLA-1.0 → valuation regression.
- [Embodied] Qianxun (Spirit AI): three rounds in three months totaling >¥4.5B (Shunwei / Yunfeng rare joint appearance, CATL / JD / TCL venture capital in); Spirit v1.6 topped North America’s RoboArena global #1 (first Chinese embodied model); Moz1 industrial humanoid + Moz2 commercial service, first line already running at CATL’s Zhongzhou base → strong bullish.
- [Regulation] Industry shakeout ⚠️: Mech-Mind founder called out Galaxy General’s “data-collection center + related-party transactions fabricating fake revenue,” Qianxun / Star Air map reported; 2026H1 domestic humanoid bid total ¥1.723B but real commercial orders only 21.1%; rumored informal window guidance on humanoid IPOs → “bid ≠ commercial revenue.”
- [Chain] Primary market & training grounds: ENCOS (Inx) >¥300M Series B (joint module → platformization), RobotPlusPlus ¥100M+ Series C (hull derusting 200万+ hours, “work is data collection”), CAICT says >70 embodied training grounds are built and running nationwide → maps to Leaderdrive 688017 / Sanhua 002050 / Tuopu 601689 / robot ETF 562500 / Horizon 09660.HK.
- [Edge VLA] 5 models from 3 vendors in 8 days: Wuqin Core + Tsinghua/SJTU open-sourced APXInf edge engine (4090/Jetson Orin/Thor, as low as ~26 ms), Juna Cog-WM Cog-WM 1.0 brain-like world model, Force Lingji DM0.5 tops six leaderboards — bodies collectively “get smarter.”
- [General]: OpenAI launch week (GPT-6 Sol possibly Thursday), DeepSeek open-sourced Harness agent framework + plans Ulanqab ~1 GW data center (≥160k Ascend 950DT) + preparing Sci-Tech board IPO, Doubao Seed-2.1-Pro-0915 (image/video inference Token down 30%+), nine-ministry “15th Five-Year” smart-NEV plan sets “AI + auto” (Horizon 09660.HK benefits), Ant open-sourced safety guardrail SingProbe Infra (fits 29 models, plugs into SGLang/vLLM, overhead <0.5%).
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
On the framework side, vLLM #56935 (merged 09-16) wires DeepSeek-V4.1’s FlashMLA mega attention + NVFP4 compressed KV into one path and sets it SM100 default — the first change to make both a “fused kernel” and “lower-precision KV” a default path at once (45% less KV, 1.45× faster at TP1, no precision loss), with three boundaries clearly drawn; SGLang lands a six-PR multimodal V4.1 chain and PD disaggregation gets an optional Adler-32 checksum (#39500, from “runs” to “verifiable”); four cross-path bugfixes say don’t sit on v0.29.0, follow main or wait for v0.29.1. On the papers side, π*0.6/RECAP (Physical Intelligence, arXiv:2511.14759) proves “pretrained general VLA + task-level real-robot RL fine-tune” is deployable (throughput ×2+, failure rate halved, 13 hours of coffee), shifting the robot bottleneck from model strength to real deploy data, the same thread as the industry’s “bodies get smarter but only 21.1% of 2026H1 humanoid bids were real commercial orders”; Selective Scan/S6 (Mamba) uses O(N) associative scan to push inference KV cache to a few MB and elastic inference lifts GPU utilization from ~20% to 60%+, rounding out the “edge brain + real-robot deploy” cost model from algorithm to system — and UBTECH’s 10k-unit factory, Zhiyuan’s A3 Ultra thousand-unit production, Unitree’s 57% drawdown, Qianxun’s three rounds >¥4.5B, and the regulator’s “bid ≠ commercial revenue” are each putting a price tag on the embodied “smarter brain + real landing” stack.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。