系列:每日AI热点

Daily AI Hotspot · 2026-08-27: vLLM v0.28.0 Ships (584 Commits) — Full-Stack Kimi-K3 Performance, End-to-End DeepSeek V4 Sparse MLA, DFlash2 Reaches Stable

★ Most Worth Your Attention Today

vLLM v0.28.0 is officially out (2026-08-26, 584 commits / 270 contributors).

This is the unambiguous headline of the period. It is not a patch-up release — it pushes the two weightiest technical threads to production-ready in one shot.

Thread one: system-level performance for Kimi-K3. Four commits land as a coordinated combination:

PRWhat it doesEffect
#50484DCP decode-context parallelismScales long-context decode
#51070Merged all-gather1.5–3× kernel speedup
#50912Shared-expert sharding instead of replicationSaves ~17 GiB per GPU
#51725Adaptive DSpark budgetK3 DSpark TTFT improves ~60%

Individually each is “just an optimization”; stacked, they produce an order-of-magnitude improvement in K3 throughput and latency under long context plus high concurrency. #50912’s ~17 GiB per GPU deserves special attention — it changes K3’s deployability outright. The reclaimed memory can go to higher concurrency or longer contexts instead of being spent duplicating shared experts.

Thread two: DeepSeek V4 sparse MLA wired end-to-end (#51538), covering all three modes — plain / MTP / DSpark. Together with DFlash2 (#52816) reaching stable, vLLM’s speculative decoding stack converges in this release.

Why #51725 is worth understanding on its own: DSpark ranks prefixes by confidence cumprod, applies EMA smoothing, and verifies via varlen CUDA graph. The core insight is that with a fixed draft length k, once the GPU saturates at high concurrency, verification costs more than it returns (measured 08-13: at c=256, a fixed draft of 7 was 33% slower). #51725 turns this into a per-request adaptive budget, and K3’s DSpark TTFT improves by roughly 60%. Upgrade to v0.28.0 and you get it — no workload changes required.

1. AI Industry & Paper Highlights

Paper: DALL·E 2 (Ramesh et al., OpenAI, 2022)

Hierarchical text-conditional image generation. Its crucial contribution is splitting “semantic alignment” from “pixel reconstruction” into two stages, bridged by the CLIP latent space:

  1. Prior: maps the CLIP text embedding $z_t$ to an image embedding $z_i$. This stage handles semantics only.
  2. Decoder (GLIDE + CFG): diffuses an image from $z_i$. This stage handles pixel quality only.
  3. Upsamplers extend to $1024^2$.

Additional capabilities: image variations and latent-space interpolation.

Impact: it validated the “shared latent space + conditional diffusion” approach, directly inspiring Stable Diffusion. Looking back, this two-stage decoupling is now the default architecture for essentially every text-to-image system.

Operator: MTP (Multi-Token Prediction, DeepSeek V3/V4)

Stack k lightweight prediction heads on top of the backbone Transformer, predicting the next k tokens simultaneously, with loss weight λ = 0.3.

At inference those heads are naturally the draft model, with an extremely high acceptance rate, cutting 10–20% of latency overall. In use today: DeepSeek V3 / V4, Qwen3. In embodied settings: roughly 15% higher real-time control frequency for VLAs.

MTP’s elegance is that something you trained as a side effect turns out to be useful at inference. No separate draft model, no extra alignment training — the only cost is the modest compute for k heads during training. It is the best value-for-effort design in this period.

Performance: Long-Context Streaming and Chunking Strategy

A three-part combination:

Result: 128K context memory drops from 80 GB to 12 GB. Already deployed in vLLM / SGLang / DeepSeek V4.

The Attention Sink piece is easy to overlook but critical: a large share of attention mass concentrates on a handful of tokens at the very start of the sequence, and sliding them out of the window visibly degrades model behavior. Pinning the first 4 tokens costs essentially nothing and buys stability.

For on-device VLAs this is what makes real-time video streaming feasible at all.

Industry Highlights

  1. NVIDIA: Q2 revenue $96.2B, up 106% year over year; Q3 guidance $108B; the CFO issued a rare long-range outlook of another 70% growth by fiscal 2028. → Compute demand is real and still accelerating.
  2. Zhipu open-sources GLM-5.3-Flash: the first model trained on a cluster of 100,000 domestically produced chips, priced at just 1/40 of Opus 4.8. → The domestic full-stack compute picture is filling in fast.
  3. OpenAI’s in-house inference chip Jalapeño appears at Hot Chips: 1.5–1.9× the throughput per watt of NVIDIA’s GB300. → The inference silicon landscape is fragmenting.
  4. The Road Traffic Safety Law draft amendment adds a dedicated “automated driving” chapter for the first time: violations while the system is engaged are borne by the automaker. → The liability boundary for L3+ commercialization is finally defined.
  5. Alibaba open-sources Qwen3.8-Flash: 125B MoE with only 6B active, 8×+ faster at 1M context, 90% lower training cost. → Chinese AI is shifting from a scale race to an efficiency race.
  6. Meta launches the MTIA 400 generative AI accelerator: 3nm, dual chiplets, 12 petaFLOPS. → In-house silicon has moved from optional to mandatory.

Embodied AI Dispatches

CompanyReadNotes
Unitree (688836)NeutralDown from ¥1,100 to ¥586 seven days after listing; valuation resetting; the sector’s valuation anchor
AgiBot (Hong Kong IPO)PositiveTarget valuation HK$40–50B; #1 global shipments in 2025 (39%); proxy 688585
XPeng RoboticsPositiveOver $900M first round at $6.3B; Tencent and Alibaba participating; proxy 09868.HK
Tiangong UltraPositive8.64 s in the 100 m, beating the human record; motion-control supply chain
Xiaomi (01810.HK)PositiveThree Xuanjie chips launched at once, fully in-house; on-device AI + autonomous driving + robotics in concert

The industry item most worth remembering today is NVIDIA’s “another 70% by fiscal 2028.” A CFO putting a quantified three-year number on the record is exceptionally rare in semiconductors — it usually means order visibility already extends that far, rather than optimism about demand.

2. vLLM & SGLang Community Tracking

vLLM v0.28.0 (08-26)

Release: 584 commits / 270 contributors.

Performance (Kimi-K3 system-level; see the table at the top).

New features:

Architecture evolution / breaking changes (read this before you upgrade):

⚠️ Deployment impact: this is an environment-level and dependency-level change. bitsandbytes no longer ships with the core, so the quantization path needs a separate install; Transformers 5.15.0 is a major-version jump that will very likely break custom model definitions. Schedule it as “rebuild image → smoke test → accuracy regression → canary.” Do not run pip install -U vllm in place.

SGLang (still v0.5.18, 08-22, 710 PRs)

No new tag in the last 72 hours, but main saw heavy commits on 08-26 / 08-27.

PRWhat it doesImpact
#33561Support for Ling-3.0-flash (BailingMoeV3, Baichuan’s new flagship MoE)Domestic MoE reaches first-class support
#31626Beam search supportFills in the non-sampling decode path
#36233CUDA 13.4 container probing Rubin (sm_107)Groundwork for next-generation hardware
#36397custom all-reduce v2 tuningCommunication kernel polishing
#35379Generalized hybrid SWA MTP draft-pool routingSpeculative decoding on hybrid architectures
#33871Cut idle DP work in breakable prefill CUDA graphRemoves spin
#35640Coordinate FullCG prefill across DP-attn ranksPD + speculative polish

Standing Topic: Two PD Disaggregation Roadmaps Are Diverging

DimensionvLLM (v0.28.0)SGLang (main)
This periodE/P/D disaggregation, tiered KV offload to disk, canonical CPU layoutCross-DP-attn FullCG prefill coordination, splitting custom all-reduce into push/pull planes, configurable weight-cache daemon path
Technical bentResource decoupling + external storageProtocol / topology alignment + operational observability

The split is clear: vLLM is asking “how far away can the KV live,” while SGLang is asking “how should a request be routed.” The former expands the capacity envelope; the latter improves scheduling intelligence. Long term these converge — once capacity is tiered, routing has to take into account which tier the data sits in.

PyTorch vs transformers: vLLM v0.28.0 upgrades to Transformers 5.15.0 (dual-backend convergence, chasing day-N breadth); SGLang keeps its in-house runtime (Rust server, self-managed kernel stack). No “off transformers” PRs this period. Conclusion unchanged: model definitions converge on transformers for breadth, while the data plane and runtime converge on in-house code for determinism — layered, not either/or.

Step adaptation: no new models for days on end (Step-3.7-Flash 06-01 / 328★, Step-3.5-Flash 04-03 / 2070★, vllm fork 05-28). The Step MTP PRs — vLLM #49642 / #49490 / #52115 / #53174 and SGLang #32325 (verified still open, untouched since 2026-07-24) / #35206 — are all open, and neither the v0.28.0 nor the v0.5.18 official notes contain a single Step entry.

Conclusion: Step can still only get speedups from prebuilt images — vllm/vllm-openai:stepfun37 plus an MTP config on the vLLM side, dev-step-3.7-flash plus EAGLE on the SGLang side. Full MTP support has not landed upstream, and upstream maintenance investment lags.

3. The One-Line Takeaway

v0.28.0 is the release most worth scheduling an upgrade for this quarter — the K3 stack frees ~17 GiB per GPU, DSV4 sparse MLA is end-to-end, DFlash2 is stable, and most of the gains (adaptive DSpark cutting TTFT by ~60%) need no changes to your code; but treat it as an environment-level change: bitsandbytes has moved out of tree, Transformers has jumped to 5.15.0, and a pip install -U in place will very likely earn you a phone call at 3 a.m.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。