系列:vLLM & SGLang Serving Notes

(5) The July 2026 Release Waves — vLLM 0.25→0.26 and SGLang 0.5.15→0.5.16

1. The Whole Timeline

July 2026 was the most change-dense month for these two engines in years: both nearly synchronized a generation of architecture turnover, then nearly released their next versions on the same day.

7/10 7/11-12 7/13-14 7/18-19 7/21-23 7/25 7/29-31 SGLang v0.5.15 · Spec V2 default +11% TPS vLLM v0.25.0 · MRv2 default + PagedAttention removed vLLM v0.25.1 (NVFP4 garbage fix) · SGLang v0.5.15.post1 · GLM-5.2 500 TPS blog SGLang main: #31468 DFlash drops host sync · #31487 less prefill graph pad Split work: SGLang polishes DSpark (#31986/#31985) · vLLM dense stability fixes (#48524/#49302…) Same-day release: vLLM 0.26.0 + SGLang 0.5.16 · sync-stall war ends Axis shifts: production stability · GLM-5.2 NVFP4+MTP+P/D landing

Fig 1: The two product lines moved in almost perfect lockstep through July — evidence they face the same bottlenecks.

2. vLLM Evolution Detail

2.1 v0.25.0 (7/11–7/12): the architecture turnover

This was the biggest leap, four headline changes:

ChangePRMeaning
Model Runner V2 becomes default for all dense models#39337async-first, zero CPU-GPU sync, overlap step N and N+1
PagedAttention removed#47361abstraction pushed down to the attention-backend kernel; the paging mechanism itself is kept
Transformers v4 deprecated#40389must migrate to v5
Transformers backend parity—any new architecture HF implements is served day-0 at full speed

Also in tow: unified streaming-parse engine (#46610), native DSpark speculative decoding for DeepSeek V4 Pro (8×B300 ≈ 250 tok/s, 12–42% above MTP), heterogeneous-vocab general speculation (#38174), thinking-budget-aware speculation (#34668), and the compile requirement raised to C++20.

Performance character: MRv2's gains show most in small/medium batch — exactly the regime real Agent traffic lives in. If you only run large-batch offline jobs, you'll feel it far less.

2.2 v0.25.1 (7/14): a two-commit must-have patch

Two days after release came a patch, because a silent correctness bug was found:

For anyone serving NVFP4 (Gemma4 / Qwen-family / GLM-5.2, etc.), 0.25.0 must upgrade to 0.25.1. This bug doesn't error or crash — it just quietly outputs garbage, so it's easy to miss in production.

2.3 7/18–7/25 main: stability wrap-up

0.25 was a big change, and a wave of edge-config fixes followed:

PRFixes
#48524DFlash layer sizing wrong when num_target_layers ≠ num_hidden_layers
#49302DSA crash under interruptible segmented CUDA Graph
#48843graph_pool_id not set before full CUDA Graph capture
#49306FA4 JIT warmup MLA fallback
#48860KV Connector delayed request double-counts prefix-cache metric
#49292Qwen3-VL M-RoPE on the Transformers backend
#49190Cosmos3 Edge video model
#48816GPTQ Qwen3.5 MTP weight-loading error when speculation on
#42569FA4 on SM100 (Blackwell) adds FP8 KV cache support
#48683ROCm up to AITER v0.1.16.post5
#45991XPU adds DeepSeek-V4 fuse_index_q SYCL path
#49244remove old partial-prefill params
#48914 / #49427FlashInfer up 0.6.15; restore dequant_cache OOB guard
vLLM main merged roughly 64 commits/day in June. That speed means: production must lock release tags, don't follow main; run a full regression before upgrading.

2.4 v0.26.0 (7/25)

2.5 7/29–7/31: shifting to production deployment

As the version cadence slowed, the focus turned to landing:

3. SGLang Evolution Detail

3.1 v0.5.15 (7/10): zero-overhead Spec V2

FeatureNote
Zero-overhead Spec V2 defaultdraft-extend made CUDA-graph-capturable, cut D2H/H2D, fused metadata → end-to-end +~11% TPS
IndexShare MTPdraft step reuses target model’s DSA top-k index, long-context draft cost down ~1.9×
Breakable CUDA Graph defaultgraph interruptible, balancing graph gains with scheduling flexibility
MLA context parallel decodingcontext-parallel decoding for DeepSeek-family MLA
FlashInfer all-to-all MoE routingMoE routing via FlashInfer
Native web searchbuilt-in retrieval tool

It also fixed the 0.5.5–0.5.12 multimodal path-traversal vulnerability (GHSA-qwrp-wghp-94q2).

3.2 v0.5.15.post1 (7/14): GLM-5.2 tuning

Coinciding with the lmsys official blog “GLM-5.2 NVFP4 500 TPS”:

Teams running GLM-5.2 should lock post1.

3.3 7/18–7/23 main: speculative decoding keeps getting polished

PRContent
#31468DFlash removes per-step host sync in spec decode — CPU leads GPU by a full step (pipeline); branch logic made CUDA-graph-capturable
#31487reduce prefill CUDA graph padding
#31986DSpark dense-draft per-layer ctx KV projection stacked into a single big GEMM
#31985fold draft embedding into the draft graph via forward_embed
#31682DP attention default breakable prefill cuda graph
#31981DSA draft-extend metadata kernel skips page-table columns beyond KV length
#30272SM120 DeepSeek V4 flashinfer_mxfp4 MoE + TP2 (consumer Blackwell)
#24013VLM cross-request ViT encoding batching (multimodal concurrency)
#31825ModelOpt FP4 path supports NVFP4_AWQ
#31762fix marlin_nvfp4 routed_scaling_factor
#30924unified per-token-group quant kernel
#30540new HPC-Ops attention backend
#31109remove QServe and FBGEMM FP8 quant (old paths migrate to ModelOpt / compressed-tensors)
#32047fix stale per-token-group quant callers; evict_from_tree_cache only evicts the KV gap
#27894 / #31835NIXL disagg functional tests into CI; PrefillDelayer delays until KV-budget admission
Read #31468, #31986, #31985 together and a very clear main line emerges: eliminate one by one every "ask the CPU every step" in speculative decoding, merge small kernels into big kernels, finally let the whole step into CUDA Graph. What's saved is accelerator idle time, not compute.

3.4 v0.5.16 (7/25): UnifiedRadixTree

UnifiedRadixTree becomes the default prefix cache — unifying previously scattered cache paths (ordinary prefix cache, hierarchical cache, PD-scenario cache) into one tree for hit and eviction logic.

3.5 7/29–7/31

4. Upgrade Checklist

From a month of traps, a practical checklist:

vLLM 0.24.x → 0.25.x/0.26.x

SGLang 0.5.x

5. What the Timeline Tells Us

  1. The two move in tight lockstep: 0.25.0 (7/11) vs 0.5.15 (7/10), 0.26.0 vs 0.5.16 (both 7/25). This means they face the same bottlenecks, not separate wheel-reinvention.
  2. A big release is always followed by a wave of fixes: a patch within a week of 0.25.0, then two weeks of dense edge-config fixes. Waiting 1–2 weeks after a new version before production is the rational choice.
  3. The gap between main and release is large: SGLang’s frontier optimizations (DSpark GEMM stacking, etc.) are all on main; stable tags lag. Want performance → follow main and benchmark yourself; want stability → lock the tag.
  4. The main axis is migrating: early July was “kill sync stalls”; after the same-day late-July release, the axis shifted to speculative decoding + prefix cache + day-one model support + production deployment.

The next post is the most practical one: which models are supported, and who to pick for what scenario.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。