系列:每日AI热点

Daily AI Hotspot · 2026-08-23: SGLang v0.5.18 Cuts Startup to 35.6 Seconds, While vLLM v0.28.0 Enters RC with DFlash2

Today is a release day: SGLang ships v0.5.18 (08-22, 710 PRs / 212 contributors), and vLLM pushes v0.28.0rc1 / rc2 into RC in the same week (08-20 / 08-21), with rc2 headlined by DFlash2 — which only landed in SGLang on 08-19. Both frameworks are doing the same thing: eliminating bubbles on the critical path. SGLang fills the I/O bubble at startup; vLLM fills the synchronization bubble in speculative decoding.

★ Most Worth Your Attention Today

SGLang #32017: overlapped checkpoint loading cuts Qwen3-32B startup from 84.8s to 35.6s (2.38×).

Start with what a traditional startup actually does. Launching an inference instance is two serial phases:

  1. Weight loading: copy the checkpoint from disk or object storage into GPU memory — pure I/O, with the GPU waiting.
  2. CUDA graph capture: record the kernel sequences for each shape — pure GPU compute plus host-side recording, with I/O idle.

These two wait on each other while using completely orthogonal resources: the GPU sits idle while weights copy, and the disk sits idle while graphs are captured. This is a textbook pipelining opportunity.

v0.5.18 pipelines chunked weight staging with progressive graph capture: weights move into memory chunk by chunk while graphs are captured incrementally for whatever shapes are already available. The gaps where the GPU would be waiting on I/O get filled with graph capture.

Measured (Qwen3-32B, H100):

Traditional serialOverlapped loadingChange
Startup time84.8 s35.6 s2.38×

The flag is opt-in: --startup-weight-load-mode overlap.

Who should care about this number: not services running long steady state, but three specific scenarios — restart, scale-out, and preemption recovery.

  • Elastic scale-out: spinning up new instances at peak. 2.38× means scale-out responds more than twice as fast.
  • Preemptible instances: after preemption, rebuild time counts directly against availability.
  • Rolling upgrades / canaries: each release restarts a batch of instances, and startup time multiplied by instance count is your total rollout duration.
  • Development: waiting 85 seconds versus 36 seconds after a one-line config change are two completely different working rhythms.

The flip side: a service running long steady state gains nothing — startup happens once. Don’t mistake this for a throughput metric.

Three changes of the same nature: this release carries two more items structurally identical to #32017 — TP LMHead switched to All-to-All, #32313 (LMHead on DS-V4-Pro B200 drops from 320 µs to 169 µs, eliminating a collective-communication bubble) and FlashInfer MNNVL pure allreduce, #30700 (+6.9% at small batch on Blackwell). Add DFlash2 from 08-19 (which removes a CPU-GPU sync bubble) and SGLang’s week looks remarkably unified: fill every spot on the critical path where one side is idle while the other waits.

1. About Today’s Paper-Side Digest

There is no AI paper / industry digest for this date. No paper-digest file for 2026-08-23 exists in this series’ daily digest archive, so this post has no “AI Industry & Paper Highlights” section — we do not fabricate one, and we do not substitute content from adjacent dates. This entry covers only the vLLM / SGLang engineering side, with the space going into community tracking and mechanism deep dives instead.

(For context: the 08-21 paper-side theme was ResNet skip connections and DeepNorm; digests from 08-25 onward are covered by other entries in this series.)

2. vLLM & SGLang Community Tracking

SGLang v0.5.18 (released 08-22, 710 PRs / 212 contributors)

SGLang’s “we’re dropping torchao” signal

#34304, removing torchao, deserves separate attention because it is the clearest architectural statement of the year so far. Combined with #32434’s unified kernel cache directory, the message is: the quantization path and the kernel stack are ours to own.

This matches the conclusion this series has tracked all along — model definition converges toward transformers for breadth, while the data plane and runtime converge toward in-house code for determinism; it is layered, not either-or. SGLang is simply tightening its grip on that data-plane layer.

The costs deserve equal weight:

vLLM v0.28.0 enters RC

Standing topics

PD disaggregation. SGLang v0.5.18 introduces a dedicated Disaggregation + PD section — a signal that this has moved from “experimental feature” to “subsystem with an owner”: NIXL bootstrap timeout #34692, Mooncake PP prefill #33807, unified-memory PD #33362, prefill skipping speculative scratch #34191. vLLM, meanwhile, continues hardening KV offloading and parallelism decoupling. PD on both sides has entered the “operable” phase: the question is no longer whether it runs, but whether it explodes at boundary conditions (timeouts, drivers, unified memory).

Architecture evolution. The main thread is DFlash2 landing across frameworks: SGLang #35371 (merged 08-19) → vLLM #52816 (into rc2). The core algorithms of speculative decoding are now synchronizing between the two frameworks on a timescale of days, which would have been unimaginable two years ago. On the other side, torch 2.13 is tracked by both frameworks, SGLang’s removal of torchao and unified kernel cache directory consolidate its in-house kernel stack, and MoE deferred finalize plus DSV4 mhc fusion are both on by default.

Step-series support (still entirely open). All five vLLM Step PRs remain open (#53174, updated 08-20 / #52115 / #49490 / #49642 / #40070), as do SGLang’s #35206 and #32325.

This cycle produced a rare official admission. Step’s own ModelScope deployment documentation states outright that “Full MTP3 not yet available in vLLM,” and notes that vLLM deployments can only set num_speculative_tokens: 1; by contrast, SGLang’s dev-step-3.7-flash with EAGLE MTP(3) is more mature.

This matters more than the status of any single PR. It converts “Step’s performance story depends on MTP” — previously an inference drawn from stalled PRs — into an officially acknowledged fact, complete with a verifiable configuration constraint: on vLLM, num_speculative_tokens can only be 1, which is to say no speculative speedup at all. If you need Step’s MTP gains today, the only route is SGLang.

In the StepFun-ai org, the only activity this cycle was a small Step-Realtime-CLI update on 08-21; the core model repos remain frozen (Step-3.7-Flash since 06-01, Step-3.5-Flash since 04-03).

3. The One-Line Takeaway

SGLang v0.5.18 and the vLLM v0.28.0 RC put two things on the table in the same week: “non-steady-state” metrics like startup time are finally being treated as first-class (Qwen3-32B 84.8s → 35.6s, directly benefiting scale-out, preemption recovery, and rolling releases), and DFlash2 took only two days to go from merging in SGLang to entering a vLLM RC — cross-framework propagation of core speculative-decoding algorithms is now fast enough to require day-level tracking; the two things actually worth doing today are adding --startup-weight-load-mode overlap to elastically scaled instances, and checking whether your Step deployment is stuck at num_speculative_tokens: 1 — if so, switch to SGLang.


Sources: SGLang v0.5.18 release (2026-08-22); SGLang PRs #32017 / #32313 / #30700 / #28836 / #34304 / #32434 / #34692 / #33807 / #33362 / #34191 / #35206 / #32325; vLLM v0.28.0rc1 / rc2 tags; vLLM PRs #52816 / #52560 / #50272 / #53204 / #50723 / #53174 / #52115 / #49490 / #49642 / #40070; Step-3.7-Flash-FP8 and Step-3.5-Flash-Int4 deployment docs (ModelScope); StepFun-ai org push timestamps.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。