系列:每日AI热点

Daily AI Hotspot · 2026-08-25: vLLM Ships a ~3× Decode Kernel for 4090D/H20, SGLang Spreads DSpark Verification Across Four Quantization Tiers

★ Most Worth Your Attention Today

vLLM #53247: architecture-tuned batch-invariant persistent matmul — roughly 3× on RTX 4090D / H20 decode kernels.

It wins today because it is the only change here that pays out an order of magnitude with zero changes to your workload.

The mechanism is not complicated. The core contract of a batch-invariant kernel is “compile the CUDA graph once per shape, then reuse it across requests” — which is a natural fit for a persistent matmul, a hotspot repeatedly invoked with the same handful of shapes. The step that makes this PR valuable is that it pre-tunes tile/block configurations for 4090D / H20, the consumer- and workstation-class parts. Why that matters: the decode phase is exactly the region that is bound by both bandwidth and kernel launch. There is no large compute density to amortize launch overhead, and you are capped by HBM round-trips anyway. Get tile/block right and per-step latency drops sharply — measured at roughly 3×.

How to turn it on: set VLLM_BATCH_INVARIANT. No model change, no operator change, no service-code change.

The other side worth noting: this belongs to the same family as today’s #52193 (speculative decoding workspace taking the max hidden dim under TP>1) — polishing an existing mechanism until it is actually usable in production. There is no big release today, and all of the value lives in details like these. If your upgrade policy is “wait for a tag,” you will miss a fortnight of exactly this kind of gain.

1. AI Industry & Paper Highlights

Paper: ACT (Action Chunking with Transformers, ICRA 2023, Stanford)

The foundational work for action control in embodied AI. The core move is to predict a block of k future actions in one Transformer pass, decoupling low-frequency observation from high-frequency control: observations arrive a few times per second at best, but control has to run at 50 Hz, and action chunks are precisely what bridges that frequency gap.

Two design decisions carry the paper:

Today, the action-generation modules of modern VLAs (RT-1 / π0 / OpenVLA) are essentially variants of this.

Operator: Chunked Prefill

The core inference mechanism in vLLM / SGLang: split a long prompt into chunks and accumulate them into the KV cache block by block, letting long prefills interleave with short decodes and removing the head-of-line blocking that long sequences impose on short requests.

The math relies on an online softmax to guarantee that chunked results are strictly equivalent to a full prefill — which is the precondition for having it on by default with confidence. It is not an approximation.

In embodied settings this is concrete: multi-frame visual instructions in a VLA easily reach 2K+ tokens. Without chunking, the robot stalls wholesale during the “make sense of the scene” phase and every decode behind it is blocked.

Performance: Prefix Caching

A Radix Tree detects shared prefixes automatically, so the KV cache gets computed once and used many times. Measured on SGLang: 2.2× throughput on multi-turn conversation, 3.5× on code-generation workloads.

Combined with PagedAttention it supports copy-on-write data isolation — separate requests share one physical prefix, but each can write without disturbing the others. It fits multi-task VLA scenarios especially well: the visual prefix shared across a batch of tasks is reused outright.

Industry: Embodied AI’s Triple Catalyst — Mass Production, IPOs, Record Runs

CompanyMoveNumbers
Unitree Robotics (688836)Post-IPO allocationCumulative output ≈ 18,000 units; ¥2.022B into embodied foundation models — for the first time more than the hardware body itself
AgiBot / Zhiyuan RoboticsHong Kong IPO filingTarget valuation HK$40–50B; A-share proxy 688585 Shangwei New Materials
Tiangong UltraHumanoid Robot GamesWon the 1500 m in 2 min 21.63 s, beating the human world record
Galaxy Botan / Dobot (WRC)Competing roadmaps“One brain, many capabilities” vs. “one brain, many bodies,” breaking the one-model-per-robot lock-in
UBTech (09880.HK)Automotive robot procurement win¥90.51M — the largest single humanoid robot order worldwide

The item most worth chewing on is Unitree’s: for the first time, capital is going from the body toward the embodied foundation model. That is a hardware maker stating publicly where the value sits — the chassis is no longer the moat, the brain is.

2. vLLM & SGLang Community Tracking

Version baseline: vLLM v0.27.1 (08-11) / v0.28.0rc2 (08-21), SGLang v0.5.18 (08-22). No new tags from any of the three projects in the last 72 hours — the movement is all on main.

vLLM

PRCategoryWhat it doesImpact
#53247Perf · kernelBatch-invariant persistent matmul with pre-tuned tile/block for 4090D/H20~3× decode kernel ★
#52193Bugfix · speculative decodingUnder TP>1, workspace takes the max of target/draft hidden dimFixes multi-GPU MTP/DFlash shape mismatch
#52157Feature · speculative decodingAdaptive verification for varlen trtllm-gen decodeBetter acceptance length and quality at high temperature
#52676Perf · kernelFused QK-norm + partial MRoPE + gate for Qwen3.6Fewer HBM round-trips
#53165Bugfix · multimodalFixes mixed CLIP/SigLIP pooling text encodingMultimodal robustness

#52193 deserves a paragraph of its own. It is the key correctness fix for running MTP on Step-3.7-Flash at TP=8: when the draft and target models disagree on hidden dim, the old logic sized the workspace off a single dimension, and multi-GPU MTP / DFlash crashed at startup. Taking the max of the two makes it come up cleanly.

What makes bugs like this obnoxious is that they do not produce wrong answers — they stop you from starting at all, and the error message usually points at a completely unrelated shape assertion.

Also worth flagging: #52157 extends adaptive verification to variable-length trtllm-gen paths, echoing the main DFlash2 adaptive-verification line of work. The “fixed draft length” assumption is being retired path by path.

SGLang

SGLang’s theme for the day is unmistakable: pushing from “does it run” to “does it run correctly at every quantization tier.” The component-level work on the diffusion side is especially notable — different components in image/video generation vary enormously in quantization sensitivity, and whole-model one-size-fits-all quantization has always been a compromise.

3. The One-Line Takeaway

No new releases today, but #53247 trades one environment variable for a ~3× decode kernel on 4090D/H20, and #52193 stops multi-GPU speculative decoding from dying at startup — on days with no tag, what actually pays is exactly this class of commit: no headline, but it changes what you can run.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。