系列:每日AI热点

Daily AI Hotspot · 2026-08-28: Three Default Values Changed on main After v0.28.0, While SGLang Rebuilds Its Prefill Foundation

★ Most Worth Your Attention Today

Not v0.28.0 itself — the three default values main changed starting 08-27.

Version numbers are the easiest thing to watch, but what actually changes your production behavior is usually defaults. This time there are three:

ChangePRNew defaultWhat you should worry about
Model Runner V2 on by default for all models#53183Every model runs through MRv2Coverage expands from “some models” to “all models” — your regression scope must expand with it
Blackwell CUDA graph capture defaults to 1024#49390Capture batch-size ceiling of 1024Higher memory footprint; small-VRAM cards need capacity recalculated
max_num_batched_tokens 8192 → 16384#51726Per-batch token ceiling doubledSingle-GPU / small-VRAM setups can OOM outright

Why I rank defaults above release benefits: you cannot enjoy a performance gain without upgrading, but you absorb a default change whether you notice it or not. The doubling of max_num_batched_tokens is the textbook case — it genuinely helps throughput, but the cost is that peak activation memory rises with it, and plenty of teams tuned their concurrency settings against the old number.

Recommended action: if you track main rather than stable tags, write max_num_batched_tokens back into your config explicitly, today. An implicit default is only a convenience while it stays put; the moment it changes, it becomes a production change with no changelog.

1. AI Industry & Paper Highlights

Paper: Octo (RSS 2024) — A Small-Model Generalist Robot Foundation

Octo: An Open-Source Generalist Robot Policy, available at 27M / 93M parameters, pretrained on 800k Open X-Embodiment trajectories.

Two key designs:

Why it matters: with a generic observation serialization, it turns a policy that was “one body, one task” into one that adapts to a new body in five fine-tuning steps. That makes it a powerful open-source instrument for the VLA paradigm — at 27M parameters you can fine-tune on a single GPU, which is an entirely different accessibility tier from RT-1 or π0.

Operator: Flow Matching — Diffusion, Simplified

A continuous ODE velocity field. The core is strikingly compact:

$$\frac{dx}{dt} = v_\theta(x_t, t), \quad \mathcal{L} = \mathbb{E}\lVert v_\theta(x_t, t) - (x_1 - x_0) \rVert^2$$

The training objective is simply to regress the straight-line direction from noise to data.

Versus DDPM: 4–16 steps instead of 100, measured at 27× faster with a 2% quality drop. Validated on SD3 / π0 / Octo-Flow.

Why this deserves attention: the main obstacle for diffusion in robotics was never quality, it was step count. Action generation has to run at the control frequency, and a 100-step DDPM is simply not viable inside a 50 Hz control loop. Compressing to 4–16 steps is the change that makes the diffusion approach usable in embodied settings at all.

Performance: Distributed RL Training — RLHF, Engineered

Stack: Ray + vLLM/SGLang + DeepSpeed ZeRO-3 + LoRA broadcast.

Architecture: Actor / Rollout / Critic / Reference roles decoupled, each scaling independently; LoRA broadcast synchronizes in milliseconds, avoiding a full weight distribution every round.

Result: RLHF on a 70B model compressed from 7 days to 19 hours.

A few key details:

Investment Map (Along the Embodied AI Theme)

NameSegmentKey move
XPeng RoboticsWhole machine$900M first round at a $6.3B valuation
AgiBotWhole machineH1 shipments 9,700 units / 43% share, #1 globally
UBTech 09880.HKWhole machineWalker S2 placed with BYD / Foxconn / Airbus
Unitree 688836Whole machinePulled back post-listing, but still the A-share anchor
BYD 002594Vehicles + robotics150 units in training at the Changsha plant
Wujie DynamicsDexterous hands¥700M Series A plus a ¥700M single order
Youai ZhiheIndustrial mobile manipulation¥600M C+ round; semiconductors are 38% of revenue

Current main-thread read: embodied AI is entering the “training era” — the RL training stack (OpenRLHF / veRL / AReaL), the VLA fine-tuning stack (Octo / OpenVLA), and the world-model stack (V-JEPA 2 / Genie / Cosmos) are converging into one. All three rest on the same foundation: distributed training capability. That, rather than model architecture, is what becomes the new moat.

2. vLLM & SGLang Community Tracking

vLLM v0.28.0 (08-26)

Performance · full-stack Kimi-K3 optimization:

Teams running K3 benefit immediately on upgrade.

New features · keeping speculative decoding from losing at high concurrency:

PRWhat it does
#51725Adaptive DSpark verification budget; K3 TTFT improves ~60%
#52816DFlash2 local-convolution candidate selector
#48341Asynchronous scheduling with automatic drafting

Default changes (see the ★ section above): #53183 MRv2 on for all models, #49390 Blackwell CUDA graph capture defaults to 1024, #51726 max_num_batched_tokens 8192→16384.

Under the Hood: The Adaptive DSpark Verification Budget (#51725)

It turns “how many draft tokens to verify per step” from a fixed k into an adaptive budget allocated by live load and acceptance rate.

The core fact: with a fixed k, once the GPU saturates at high concurrency, verification costs more than it returns (33% slower at c=256). After adaptivity: ±3% at low concurrency, gains preserved at high concurrency, and K3 DSpark TTFT improves ~60%.

What it looks like in practice: once DSpark is on, you stop hand-tuning k — the framework sizes the budget off the acceptance rate. Prerequisite: it is experimental and needs v0.28.0+.

SGLang

PRWhat it doesImpact
#36288Mixed Chunk Prefill Base (1/N) — rebuilds the prefill / mixed-batching foundationGroundwork for fine-grained prefill/decode decoupling
#35451Full prefill CUDA graph supports PPFaster long-context prefill
#36160Unified PD control plane (mori)PD enters the “elastic routing” stage
#34608Load-aware routing, publishing per-scheduler load to a dedicated socketPD enters the “operational observability” stage
#36518 / #36608GLM-5.3-Flash reaches first-class tuning (cookbook plus AMD MI300X/325X/355X recipes)Domestic model as a first-class citizen
—LingBot-Video / cosmos3 diffusion fusionMultimodal diffusion
#36547Fixes DeepSeek-V4 multi-stream QKV lifetimeCorrectness
#36330Fixes AMD Qwen3.5 MTPCorrectness
#35275Fixes spec adaptive startup crashCrash-at-startup class

#36288 deserves an extra note. Mixed Chunk Prefill Base is labeled 1/N, which tells you this is a multi-stage foundation rebuild and today is only the first step. Its goal is fine-grained prefill/decode decoupling — the same problem vLLM is attacking through MRv2, while SGLang approaches it from the prefill side. Different paths, same target: prefill and decode have such different resource profiles that running them through one scheduler inevitably makes them drag on each other.

Standing Topics

PD disaggregation: vLLM v0.28.0’s tiered KV offloading is mature (disk #49644 + second tier #51007 + partial loading #50321 + metrics #48798), paired with E/P/D decoupling; SGLang main is building a unified control plane (#36160) plus load-aware routing (#34608).

Side by side: vLLM emphasizes tiering KV offload, SGLang emphasizes smarter request routing. The former asks “does the data fit,” the latter asks “where should this request go.”

PyTorch vs transformers: vLLM keeps migrating models onto the Transformers backend (MLA #48250 / hardware-agnostic model definitions #49458) for breadth, while decoupling by moving bitsandbytes out of tree (#43529); SGLang has no aggressive “off transformers” PRs. Conclusion unchanged — layered convergence, at the cost of having to schedule the breaking Transformers 5.15.0 upgrade.

Step adaptation: all upstream PRs remain open (vLLM #49642 / #53174, SGLang #32325 / #35206), with #53174 updated on 08-26 (Step-3.5 MTP + structured outputs). Meanwhile the official vLLM recipe now ships a Step-3.7-Flash MTP-3 example (on the stepfun37 image), runnable on 4 cards at NVFP4. The upstream v0.28.0 release notes list no Step models — production deployment still depends on StepFun images / forks.

3. The One-Line Takeaway

The v0.28.0 benefit list (K3 freeing ~17 GiB per GPU, DSpark TTFT improving ~60%) is the visible part; what actually bites is the three defaults quietly changed on main — above all max_num_batched_tokens doubling from 8192 to 16384, which can OOM a small-VRAM card without you touching a single config. If you track main, write it back into your config explicitly, now.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。