★ Most Worth Your Attention Today
Not v0.28.0 itself — the three default values main changed starting 08-27.
Version numbers are the easiest thing to watch, but what actually changes your production behavior is usually defaults. This time there are three:
| Change | PR | New default | What you should worry about |
|---|---|---|---|
| Model Runner V2 on by default for all models | #53183 | Every model runs through MRv2 | Coverage expands from “some models” to “all models” — your regression scope must expand with it |
| Blackwell CUDA graph capture defaults to 1024 | #49390 | Capture batch-size ceiling of 1024 | Higher memory footprint; small-VRAM cards need capacity recalculated |
| max_num_batched_tokens 8192 → 16384 | #51726 | Per-batch token ceiling doubled | Single-GPU / small-VRAM setups can OOM outright |
Why I rank defaults above release benefits: you cannot enjoy a performance gain without upgrading, but you absorb a default change whether you notice it or not. The doubling of max_num_batched_tokens is the textbook case — it genuinely helps throughput, but the cost is that peak activation memory rises with it, and plenty of teams tuned their concurrency settings against the old number.
Recommended action: if you track main rather than stable tags, write
max_num_batched_tokensback into your config explicitly, today. An implicit default is only a convenience while it stays put; the moment it changes, it becomes a production change with no changelog.
1. AI Industry & Paper Highlights
Paper: Octo (RSS 2024) — A Small-Model Generalist Robot Foundation
Octo: An Open-Source Generalist Robot Policy, available at 27M / 93M parameters, pretrained on 800k Open X-Embodiment trajectories.
Two key designs:
- Readout Tokens: 8 read-only learnable tokens serve as a context summary, and the downstream head looks only at those 8. This is input/output decoupling in essence — the backbone can ingest observation sequences of any length while head complexity stays constant.
- Action Chunking: predict T_p = 64 steps in one pass, but execute only the first T_a = 16 steps before replanning (receding horizon). Plan long, act short, correct often.
Why it matters: with a generic observation serialization, it turns a policy that was “one body, one task” into one that adapts to a new body in five fine-tuning steps. That makes it a powerful open-source instrument for the VLA paradigm — at 27M parameters you can fine-tune on a single GPU, which is an entirely different accessibility tier from RT-1 or π0.
Operator: Flow Matching — Diffusion, Simplified
A continuous ODE velocity field. The core is strikingly compact:
$$\frac{dx}{dt} = v_\theta(x_t, t), \quad \mathcal{L} = \mathbb{E}\lVert v_\theta(x_t, t) - (x_1 - x_0) \rVert^2$$
The training objective is simply to regress the straight-line direction from noise to data.
Versus DDPM: 4–16 steps instead of 100, measured at 27× faster with a 2% quality drop. Validated on SD3 / π0 / Octo-Flow.
Why this deserves attention: the main obstacle for diffusion in robotics was never quality, it was step count. Action generation has to run at the control frequency, and a 100-step DDPM is simply not viable inside a 50 Hz control loop. Compressing to 4–16 steps is the change that makes the diffusion approach usable in embodied settings at all.
Performance: Distributed RL Training — RLHF, Engineered
Stack: Ray + vLLM/SGLang + DeepSpeed ZeRO-3 + LoRA broadcast.
Architecture: Actor / Rollout / Critic / Reference roles decoupled, each scaling independently; LoRA broadcast synchronizes in milliseconds, avoiding a full weight distribution every round.
Result: RLHF on a 70B model compressed from 7 days to 19 hours.
A few key details:
- GRPO: group-normalized advantage, drops the Critic and saves half the memory. In large-model RL the Critic is often the same size as the Actor, so removing it is a structural saving rather than a tweak.
- DeepSeek V4’s CANN-GRPO + FP8: RLHF on a 670B model goes from 7 days to 19 hours (3.55×).
Investment Map (Along the Embodied AI Theme)
| Name | Segment | Key move |
|---|---|---|
| XPeng Robotics | Whole machine | $900M first round at a $6.3B valuation |
| AgiBot | Whole machine | H1 shipments 9,700 units / 43% share, #1 globally |
| UBTech 09880.HK | Whole machine | Walker S2 placed with BYD / Foxconn / Airbus |
| Unitree 688836 | Whole machine | Pulled back post-listing, but still the A-share anchor |
| BYD 002594 | Vehicles + robotics | 150 units in training at the Changsha plant |
| Wujie Dynamics | Dexterous hands | ¥700M Series A plus a ¥700M single order |
| Youai Zhihe | Industrial mobile manipulation | ¥600M C+ round; semiconductors are 38% of revenue |
Current main-thread read: embodied AI is entering the “training era” — the RL training stack (OpenRLHF / veRL / AReaL), the VLA fine-tuning stack (Octo / OpenVLA), and the world-model stack (V-JEPA 2 / Genie / Cosmos) are converging into one. All three rest on the same foundation: distributed training capability. That, rather than model architecture, is what becomes the new moat.
2. vLLM & SGLang Community Tracking
vLLM v0.28.0 (08-26)
Performance · full-stack Kimi-K3 optimization:
- DCP (#50484)
- Fused FlashKDA (#50654 / #51311 / #52458)
- Merged all-gather, 1.5–3× (#51070)
- Shared-expert sharding, ~17 GiB saved per GPU (#50912)
Teams running K3 benefit immediately on upgrade.
New features · keeping speculative decoding from losing at high concurrency:
| PR | What it does |
|---|---|
| #51725 | Adaptive DSpark verification budget; K3 TTFT improves ~60% |
| #52816 | DFlash2 local-convolution candidate selector |
| #48341 | Asynchronous scheduling with automatic drafting |
Default changes (see the ★ section above): #53183 MRv2 on for all models, #49390 Blackwell CUDA graph capture defaults to 1024, #51726 max_num_batched_tokens 8192→16384.
Under the Hood: The Adaptive DSpark Verification Budget (#51725)
It turns “how many draft tokens to verify per step” from a fixed k into an adaptive budget allocated by live load and acceptance rate.
The core fact: with a fixed k, once the GPU saturates at high concurrency, verification costs more than it returns (33% slower at c=256). After adaptivity: ±3% at low concurrency, gains preserved at high concurrency, and K3 DSpark TTFT improves ~60%.
What it looks like in practice: once DSpark is on, you stop hand-tuning k — the framework sizes the budget off the acceptance rate. Prerequisite: it is experimental and needs v0.28.0+.
SGLang
| PR | What it does | Impact |
|---|---|---|
| #36288 | Mixed Chunk Prefill Base (1/N) — rebuilds the prefill / mixed-batching foundation | Groundwork for fine-grained prefill/decode decoupling |
| #35451 | Full prefill CUDA graph supports PP | Faster long-context prefill |
| #36160 | Unified PD control plane (mori) | PD enters the “elastic routing” stage |
| #34608 | Load-aware routing, publishing per-scheduler load to a dedicated socket | PD enters the “operational observability” stage |
| #36518 / #36608 | GLM-5.3-Flash reaches first-class tuning (cookbook plus AMD MI300X/325X/355X recipes) | Domestic model as a first-class citizen |
| — | LingBot-Video / cosmos3 diffusion fusion | Multimodal diffusion |
| #36547 | Fixes DeepSeek-V4 multi-stream QKV lifetime | Correctness |
| #36330 | Fixes AMD Qwen3.5 MTP | Correctness |
| #35275 | Fixes spec adaptive startup crash | Crash-at-startup class |
#36288 deserves an extra note. Mixed Chunk Prefill Base is labeled 1/N, which tells you this is a multi-stage foundation rebuild and today is only the first step. Its goal is fine-grained prefill/decode decoupling — the same problem vLLM is attacking through MRv2, while SGLang approaches it from the prefill side. Different paths, same target: prefill and decode have such different resource profiles that running them through one scheduler inevitably makes them drag on each other.
Standing Topics
PD disaggregation: vLLM v0.28.0’s tiered KV offloading is mature (disk #49644 + second tier #51007 + partial loading #50321 + metrics #48798), paired with E/P/D decoupling; SGLang main is building a unified control plane (#36160) plus load-aware routing (#34608).
Side by side: vLLM emphasizes tiering KV offload, SGLang emphasizes smarter request routing. The former asks “does the data fit,” the latter asks “where should this request go.”
PyTorch vs transformers: vLLM keeps migrating models onto the Transformers backend (MLA #48250 / hardware-agnostic model definitions #49458) for breadth, while decoupling by moving bitsandbytes out of tree (#43529); SGLang has no aggressive “off transformers” PRs. Conclusion unchanged — layered convergence, at the cost of having to schedule the breaking Transformers 5.15.0 upgrade.
Step adaptation: all upstream PRs remain open (vLLM #49642 / #53174, SGLang #32325 / #35206), with #53174 updated on 08-26 (Step-3.5 MTP + structured outputs). Meanwhile the official vLLM recipe now ships a Step-3.7-Flash MTP-3 example (on the stepfun37 image), runnable on 4 cards at NVFP4. The upstream v0.28.0 release notes list no Step models — production deployment still depends on StepFun images / forks.
3. The One-Line Takeaway
The v0.28.0 benefit list (K3 freeing ~17 GiB per GPU, DSpark TTFT improving ~60%) is the visible part; what actually bites is the three defaults quietly changed on main — above all max_num_batched_tokens doubling from 8192 to 16384, which can OOM a small-VRAM card without you touching a single config. If you track main, write it back into your config explicitly, now.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。