系列:每日AI热点

Daily AI Hotspot · 2026-08-29: SGLang Cuts Cold Start to 35.6s (2.38×), and vLLM v0.28.0's Three Breaking Changes

★ Most Worth Your Attention Today

SGLang v0.5.18: overlapping checkpoint paging with CUDA graph capture — engine startup for Qwen3-32B / H100 goes from 84.8s to 35.6s (2.38×).

Why this one over the larger architecture headlines: cold-start latency is the most under-measured and least optimized metric of the last two years.

Everyone’s attention sits on steady-state throughput and TTFT, but in real production cold start directly determines three things:

  1. Whether elastic scaling is actually elastic — if a scale-out takes 85 seconds to become ready, you have to pre-provision, and the elasticity is discounted.
  2. Your failure-recovery SLO — when an instance dies, its replacement is unavailable for 85 seconds.
  3. Research iteration speed — every parameter tweak costs you a minute and a half of waiting.

The technique is not exotic: overlap checkpoint paged loading with CUDA graph capture, two phases that used to run serially. The former is I/O-bound, the latter GPU-bound — naturally parallelizable. Nobody had simply done it.

That is where 2.38× comes from.

Worth reading alongside it: the same release unifies all compiled-kernel cache paths under SGLANG_CACHE_DIR (a breaking change). The upside is that cache management is finally predictable; the cost is that the first startup after upgrading needs a recompile — meaning your first cold start after the upgrade will be noticeably slower than 35.6s. Shipping both in one release reads like giving you 2.38× and then clawing part of it back with one recompilation.

1. AI Industry & Paper Highlights

Paper: RDT-1B (Tsinghua + ByteDance Seed, 2024) — A Diffusion Transformer in Action Space

Robot Diffusion Transformer: a 1.2B-parameter diffusion foundation model that generates action sequences for bimanual robot manipulation.

The core idea: treat an “action chunk” (the next several steps of joint angles / poses) as the diffusion variable.

Impact and bottleneck: it validated that “diffusion + foundation model” works for robot manipulation, directly inspiring π0 / OpenVLA-ACTION / TRACE. But the bottleneck is equally plain: ~50 sampling steps means ~50× the cost. That is precisely the reason work like Flow Matching — which compresses step count — exists.

Set it next to π0 and the two action-representation roadmaps in VLA become clear: RDT-1B denoises via diffusion, π0 flows via flow matching. The former has more expressive power at a high step cost; the latter is cheap per step with slightly compromised expressiveness. Flow matching currently wins on engineering, but diffusion’s ceiling is not necessarily lower.

Operator: KV Cache Quantization (KIVI, 2-bit)

Compress attention KV from FP16 down to 2-bit, cutting memory 4–8×, which is what makes long-context and on-device VLAs feasible — with almost no accuracy loss.

The crux is that K and V need different quantization strategies — KIVI’s central insight:

The reason: K’s per-channel magnitudes vary widely (some head dims are much larger) and must be handled in groups, whereas V is smoother along tokens. This asymmetric design is why accuracy barely moves.

It pairs best with PagedAttention — page granularity lines up exactly with grouped-quantization boundaries.

In use today: DeepSeek V4 / Qwen3 store KV in FP8 (same family of ideas, less extreme); vLLM / llama.cpp quantization paths already integrate the extreme 2-bit variant.

Performance: KV Cache Offloading and Tiered Storage (GPU → CPU → NVMe)

Representative work: FlexGen / vLLM CPU offload. Serves 10–100× longer contexts.

Industry Highlights

The sharpest numbers in UBTech’s set are 921 units (+1,947%). Going from three digits to nearly a thousand units, paired with Thinker-VLA at +176% inference efficiency / −60% storage, says this is no longer “can we demo it” but can we deliver it. Competition in humanoids is shifting from “can it be built” to “can it be shipped in volume, cheaply.”

2. vLLM & SGLang Community Tracking

Two major releases in one cycle: vLLM v0.28.0 (8/26, 584 commits / 270 contributors) and SGLang v0.5.18 (8/22, 710 PRs / 212 contributors).

vLLM v0.28.0

Performance:

New features:

Architecture: Model Runner V2 lands E/P/D disaggregation, weight offloading, and multi-layer MTP KV cache; tiered KV offloading gains a disk tier.

Breaking changes (read before upgrading):

ChangeImpact
max_num_batched_tokens 8192 → 16384Higher peak memory; small-VRAM setups need retesting
bitsandbytes moved out of core → externalThe quantization path needs a separate install
Transformers upgraded to 5.15.0Custom model definitions will very likely need edits

SGLang v0.5.18

CategoryWhat it doesNumbers
PerfCheckpoint paging overlapped with CUDA graph capture at startupQwen3-32B/H100 2.38× (35.6s vs 84.8s) ★
PerfTP LMHead switched to All-to-AllV4-Pro B200 decode LMHead 320μs → 169μs, TPOT 36.97 → 35.67ms
PerfFlashInfer MNNVL pure allreduceV4-Flash TP4/Blackwell, up to +6.9% at small batch
FeatureNVFP4 weights running on AMD (quark_mxfp4)GSM8K recovers to 97.5–100.2%
FeatureKimi K3 tuned for MI355XThroughput 1.37–1.77×, ITL 1.45–2.42×
Feature7 day-0 models (5 of them diffusion: SANA-Video, LTX-2.5, Cosmos3, LongCat-Image, etc.)Multimodal / video generation as a priority
BreakingAll compiled-kernel caches unified under SGLANG_CACHE_DIRFirst startup after upgrade requires recompilation

The TP LMHead switch to All-to-All is worth expanding on. Under TP, LMHead used to all-gather the per-device logits back together, even though only the top-k is ever needed. Switching to All-to-All transfers only what is necessary, hence 320μs → 169μs — nearly halved. The resulting TPOT 36.97 → 35.67ms looks modest, but it accumulates layer by layer at long context — save a little every step and long sequences save a meaningful chunk.

Five of the seven day-0 models are diffusion-class, and that ratio is itself the signal: SGLang is treating multimodal / video generation as a primary battlefield, not merely as an LLM serving framework with extras.

Under the Hood: The Adaptive Speculative Token Budget

Conventional speculative decoding fixes the draft count per step, so under shifting load it either fails to amortize verification cost or wastes compute.

The adaptive approach uses last round’s acceptance rate $a = \text{accepted} / \text{drafted}$ to size the next round’s budget:

Paired with DSpark confidence-scheduled verification: high-confidence drafts get verified first.

No model change, no operator change — scheduling alone buys DSpark TTFT improvements of ~55–65%. This is the textbook “zero-cost gain.”

Portability: the same idea applies to a robot VLA’s low-latency control (first-action latency). Anywhere your system has a “guess first, verify later” structure, there is room for an adaptive budget.

Standing Topics

PD disaggregation: vLLM’s V2 lands E/P/D disaggregation plus weight offloading plus tiered KV offloading (disk tier, external second tier via module_path, and a CPU layout decoupled from parallelism), so very long shared prefixes can be cached across restarts; SGLang contributes a pluggable communication backend for DCP plus FlashInfer MNNVL pure allreduce to cut transfer overhead.

Architecture evolution: three speculative decoding roadmaps in parallel (DFlash2 / DSpark / Eagle3 versus SGLang’s DSpark + KDA prefix cache); sparse attention becomes standard (V4 sparse MLA, K3 KDA); quantization sinks to 4-bit (NVFP4 / MXFP4 / FP8 KV); multimodal diffusion models become SGLang’s day-0 mainstay.

Pure PyTorch vs transformers: no explicit “off transformers” PR surfaced this cycle, but vLLM v0.28.0 sends clear signals in that direction (bitsandbytes externalized, Transformers at 5.15.0, runtime KV scale computation removed).

Step adaptation (real progress this period): Step-3.7-Flash (198B sparse MoE VLM, ~11B active, 256k context) now has first-class support in vLLM — official image vllm/vllm-openai:stepfun37 plus a dedicated reasoning/tool parser, expert-parallel, and FP8/NVFP4 (NVFP4 runs on 4 cards); SGLang has a dev image supporting multimodal and speculative decoding, though it is less mature. It also supports Transformers and llama.cpp, integrates with NVIDIA Nemo/NIM, and runs on Mac / DGX / AMD Ryzen AI Max+; three reasoning_effort levels, native MCP/Skills, API output at ¥8.1 per million tokens.

Note the change in wording. Previous days concluded “depends on prebuilt images, all upstream PRs open”; today the phrasing is “first-class support.” The difference is that an official image plus a dedicated parser plus expert-parallel means production-usable, not merely a community workaround. This is the first unambiguous step forward for Step adaptation after many days of stalemate.

3. The One-Line Takeaway

Two major releases landed in one cycle, but the story is not the versions themselves: SGLang cut a metric everyone ignored for two years (cold start) by 58% through “checkpoint paging × CUDA graph capture overlap,” and vLLM took 55–65% off TTFT via an adaptive speculative budget without touching model or operators — the two best gains of this cycle both came from the scheduling layer, not the compute layer.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。