★ Most Worth Your Attention Today
vLLM #53247: architecture-tuned batch-invariant persistent matmul — roughly 3× on RTX 4090D / H20 decode kernels.
It wins today because it is the only change here that pays out an order of magnitude with zero changes to your workload.
The mechanism is not complicated. The core contract of a batch-invariant kernel is “compile the CUDA graph once per shape, then reuse it across requests” — which is a natural fit for a persistent matmul, a hotspot repeatedly invoked with the same handful of shapes. The step that makes this PR valuable is that it pre-tunes tile/block configurations for 4090D / H20, the consumer- and workstation-class parts. Why that matters: the decode phase is exactly the region that is bound by both bandwidth and kernel launch. There is no large compute density to amortize launch overhead, and you are capped by HBM round-trips anyway. Get tile/block right and per-step latency drops sharply — measured at roughly 3×.
How to turn it on: set VLLM_BATCH_INVARIANT. No model change, no operator change, no service-code change.
The other side worth noting: this belongs to the same family as today’s #52193 (speculative decoding workspace taking the max hidden dim under TP>1) — polishing an existing mechanism until it is actually usable in production. There is no big release today, and all of the value lives in details like these. If your upgrade policy is “wait for a tag,” you will miss a fortnight of exactly this kind of gain.
1. AI Industry & Paper Highlights
Paper: ACT (Action Chunking with Transformers, ICRA 2023, Stanford)
The foundational work for action control in embodied AI. The core move is to predict a block of k future actions in one Transformer pass, decoupling low-frequency observation from high-frequency control: observations arrive a few times per second at best, but control has to run at 50 Hz, and action chunks are precisely what bridges that frequency gap.
Two design decisions carry the paper:
- A cVAE models the multimodal distribution over actions, avoiding regression to the mean. This is close to decisive for manipulation tasks: under one visual observation, “go around the left” and “go around the right” are both correct, and averaging them gives you “drive straight into it.”
- Temporal Ensemble smooths chunk boundaries, killing the jitter where adjacent action blocks meet.
Today, the action-generation modules of modern VLAs (RT-1 / π0 / OpenVLA) are essentially variants of this.
Operator: Chunked Prefill
The core inference mechanism in vLLM / SGLang: split a long prompt into chunks and accumulate them into the KV cache block by block, letting long prefills interleave with short decodes and removing the head-of-line blocking that long sequences impose on short requests.
The math relies on an online softmax to guarantee that chunked results are strictly equivalent to a full prefill — which is the precondition for having it on by default with confidence. It is not an approximation.
In embodied settings this is concrete: multi-frame visual instructions in a VLA easily reach 2K+ tokens. Without chunking, the robot stalls wholesale during the “make sense of the scene” phase and every decode behind it is blocked.
Performance: Prefix Caching
A Radix Tree detects shared prefixes automatically, so the KV cache gets computed once and used many times. Measured on SGLang: 2.2× throughput on multi-turn conversation, 3.5× on code-generation workloads.
Combined with PagedAttention it supports copy-on-write data isolation — separate requests share one physical prefix, but each can write without disturbing the others. It fits multi-task VLA scenarios especially well: the visual prefix shared across a batch of tasks is reused outright.
Industry: Embodied AI’s Triple Catalyst — Mass Production, IPOs, Record Runs
| Company | Move | Numbers |
|---|---|---|
| Unitree Robotics (688836) | Post-IPO allocation | Cumulative output ≈ 18,000 units; ¥2.022B into embodied foundation models — for the first time more than the hardware body itself |
| AgiBot / Zhiyuan Robotics | Hong Kong IPO filing | Target valuation HK$40–50B; A-share proxy 688585 Shangwei New Materials |
| Tiangong Ultra | Humanoid Robot Games | Won the 1500 m in 2 min 21.63 s, beating the human world record |
| Galaxy Botan / Dobot (WRC) | Competing roadmaps | “One brain, many capabilities” vs. “one brain, many bodies,” breaking the one-model-per-robot lock-in |
| UBTech (09880.HK) | Automotive robot procurement win | ¥90.51M — the largest single humanoid robot order worldwide |
The item most worth chewing on is Unitree’s: for the first time, capital is going from the body toward the embodied foundation model. That is a hardware maker stating publicly where the value sits — the chassis is no longer the moat, the brain is.
2. vLLM & SGLang Community Tracking
Version baseline: vLLM v0.27.1 (08-11) / v0.28.0rc2 (08-21), SGLang v0.5.18 (08-22). No new tags from any of the three projects in the last 72 hours — the movement is all on main.
vLLM
| PR | Category | What it does | Impact |
|---|---|---|---|
| #53247 | Perf · kernel | Batch-invariant persistent matmul with pre-tuned tile/block for 4090D/H20 | ~3× decode kernel ★ |
| #52193 | Bugfix · speculative decoding | Under TP>1, workspace takes the max of target/draft hidden dim | Fixes multi-GPU MTP/DFlash shape mismatch |
| #52157 | Feature · speculative decoding | Adaptive verification for varlen trtllm-gen decode | Better acceptance length and quality at high temperature |
| #52676 | Perf · kernel | Fused QK-norm + partial MRoPE + gate for Qwen3.6 | Fewer HBM round-trips |
| #53165 | Bugfix · multimodal | Fixes mixed CLIP/SigLIP pooling text encoding | Multimodal robustness |
#52193 deserves a paragraph of its own. It is the key correctness fix for running MTP on Step-3.7-Flash at TP=8: when the draft and target models disagree on hidden dim, the old logic sized the workspace off a single dimension, and multi-GPU MTP / DFlash crashed at startup. Taking the max of the two makes it come up cleanly.
What makes bugs like this obnoxious is that they do not produce wrong answers — they stop you from starting at all, and the error message usually points at a completely unrelated shape assertion.
Also worth flagging: #52157 extends adaptive verification to variable-length trtllm-gen paths, echoing the main DFlash2 adaptive-verification line of work. The “fixed draft length” assumption is being retired path by path.
SGLang
- #36204 (quantization · docs) Ling-3.0-flash DSpark validated on H200 across all four quantization tiers: BF16 / FP8 / NVFP4 / INT4. The impact is straightforward — you no longer have to pick between quantization tiers, you just take the cheapest one.
- #35719 (bugfix · speculative decoding · quantization) AMD fix for Qwen3.5 MTP dropping fused shared-expert weights: on MI3xx, “MTP + quantization” was silently dropping weights. The keyword is silently — no error, you just quietly under-compute the shared expert.
- #35840 (PD disaggregation · quantization) PD regression test for inkling + mxfp8 KV: the combination of low-bit KV and PD disaggregation finally has CI watching it.
- #36035 / #36061 (quantization · multimodal) Diffusion models gain component-level quantization overrides plus mixed Comfy NVFP4/INT8 layers: quantization granularity for video-generation workloads is refined down to the component.
SGLang’s theme for the day is unmistakable: pushing from “does it run” to “does it run correctly at every quantization tier.” The component-level work on the diffusion side is especially notable — different components in image/video generation vary enormously in quantization sensitivity, and whole-model one-size-fits-all quantization has always been a compromise.
3. The One-Line Takeaway
No new releases today, but #53247 trades one environment variable for a ~3× decode kernel on 4090D/H20, and #52193 stops multi-GPU speculative decoding from dying at startup — on days with no tag, what actually pays is exactly this class of commit: no headline, but it changes what you can run.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。