★ Most Worth Your Attention Today
vLLM v0.28.0 is officially out (2026-08-26, 584 commits / 270 contributors).
This is the unambiguous headline of the period. It is not a patch-up release — it pushes the two weightiest technical threads to production-ready in one shot.
Thread one: system-level performance for Kimi-K3. Four commits land as a coordinated combination:
| PR | What it does | Effect |
|---|---|---|
| #50484 | DCP decode-context parallelism | Scales long-context decode |
| #51070 | Merged all-gather | 1.5–3× kernel speedup |
| #50912 | Shared-expert sharding instead of replication | Saves ~17 GiB per GPU |
| #51725 | Adaptive DSpark budget | K3 DSpark TTFT improves ~60% |
Individually each is “just an optimization”; stacked, they produce an order-of-magnitude improvement in K3 throughput and latency under long context plus high concurrency. #50912’s ~17 GiB per GPU deserves special attention — it changes K3’s deployability outright. The reclaimed memory can go to higher concurrency or longer contexts instead of being spent duplicating shared experts.
Thread two: DeepSeek V4 sparse MLA wired end-to-end (#51538), covering all three modes — plain / MTP / DSpark. Together with DFlash2 (#52816) reaching stable, vLLM’s speculative decoding stack converges in this release.
Why #51725 is worth understanding on its own: DSpark ranks prefixes by confidence
cumprod, applies EMA smoothing, and verifies via varlen CUDA graph. The core insight is that with a fixed draft length k, once the GPU saturates at high concurrency, verification costs more than it returns (measured 08-13: at c=256, a fixed draft of 7 was 33% slower). #51725 turns this into a per-request adaptive budget, and K3’s DSpark TTFT improves by roughly 60%. Upgrade to v0.28.0 and you get it — no workload changes required.
1. AI Industry & Paper Highlights
Paper: DALL·E 2 (Ramesh et al., OpenAI, 2022)
Hierarchical text-conditional image generation. Its crucial contribution is splitting “semantic alignment” from “pixel reconstruction” into two stages, bridged by the CLIP latent space:
- Prior: maps the CLIP text embedding $z_t$ to an image embedding $z_i$. This stage handles semantics only.
- Decoder (GLIDE + CFG): diffuses an image from $z_i$. This stage handles pixel quality only.
- Upsamplers extend to $1024^2$.
Additional capabilities: image variations and latent-space interpolation.
Impact: it validated the “shared latent space + conditional diffusion” approach, directly inspiring Stable Diffusion. Looking back, this two-stage decoupling is now the default architecture for essentially every text-to-image system.
Operator: MTP (Multi-Token Prediction, DeepSeek V3/V4)
Stack k lightweight prediction heads on top of the backbone Transformer, predicting the next k tokens simultaneously, with loss weight λ = 0.3.
At inference those heads are naturally the draft model, with an extremely high acceptance rate, cutting 10–20% of latency overall. In use today: DeepSeek V3 / V4, Qwen3. In embodied settings: roughly 15% higher real-time control frequency for VLAs.
MTP’s elegance is that something you trained as a side effect turns out to be useful at inference. No separate draft model, no extra alignment training — the only cost is the modest compute for k heads during training. It is the best value-for-effort design in this period.
Performance: Long-Context Streaming and Chunking Strategy
A three-part combination:
- Chunked Prefill: online softmax makes chunking mathematically equivalent to a full prefill.
- Streaming KV: keep only the recent window, discard the middle.
- Attention Sink: always retain the first 4 tokens.
Result: 128K context memory drops from 80 GB to 12 GB. Already deployed in vLLM / SGLang / DeepSeek V4.
The Attention Sink piece is easy to overlook but critical: a large share of attention mass concentrates on a handful of tokens at the very start of the sequence, and sliding them out of the window visibly degrades model behavior. Pinning the first 4 tokens costs essentially nothing and buys stability.
For on-device VLAs this is what makes real-time video streaming feasible at all.
Industry Highlights
- NVIDIA: Q2 revenue $96.2B, up 106% year over year; Q3 guidance $108B; the CFO issued a rare long-range outlook of another 70% growth by fiscal 2028. → Compute demand is real and still accelerating.
- Zhipu open-sources GLM-5.3-Flash: the first model trained on a cluster of 100,000 domestically produced chips, priced at just 1/40 of Opus 4.8. → The domestic full-stack compute picture is filling in fast.
- OpenAI’s in-house inference chip Jalapeño appears at Hot Chips: 1.5–1.9× the throughput per watt of NVIDIA’s GB300. → The inference silicon landscape is fragmenting.
- The Road Traffic Safety Law draft amendment adds a dedicated “automated driving” chapter for the first time: violations while the system is engaged are borne by the automaker. → The liability boundary for L3+ commercialization is finally defined.
- Alibaba open-sources Qwen3.8-Flash: 125B MoE with only 6B active, 8×+ faster at 1M context, 90% lower training cost. → Chinese AI is shifting from a scale race to an efficiency race.
- Meta launches the MTIA 400 generative AI accelerator: 3nm, dual chiplets, 12 petaFLOPS. → In-house silicon has moved from optional to mandatory.
Embodied AI Dispatches
| Company | Read | Notes |
|---|---|---|
| Unitree (688836) | Neutral | Down from ¥1,100 to ¥586 seven days after listing; valuation resetting; the sector’s valuation anchor |
| AgiBot (Hong Kong IPO) | Positive | Target valuation HK$40–50B; #1 global shipments in 2025 (39%); proxy 688585 |
| XPeng Robotics | Positive | Over $900M first round at $6.3B; Tencent and Alibaba participating; proxy 09868.HK |
| Tiangong Ultra | Positive | 8.64 s in the 100 m, beating the human record; motion-control supply chain |
| Xiaomi (01810.HK) | Positive | Three Xuanjie chips launched at once, fully in-house; on-device AI + autonomous driving + robotics in concert |
The industry item most worth remembering today is NVIDIA’s “another 70% by fiscal 2028.” A CFO putting a quantified three-year number on the record is exceptionally rare in semiconductors — it usually means order visibility already extends that far, rather than optimism about demand.
2. vLLM & SGLang Community Tracking
vLLM v0.28.0 (08-26)
Release: 584 commits / 270 contributors.
Performance (Kimi-K3 system-level; see the table at the top).
New features:
- DeepSeek V4 sparse MLA end-to-end (#51538): plain / MTP / DSpark, all three modes.
- AMD Quark NVFP4 (#47972) plus ROCm gfx11 / gfx950 (#47017 / #52212). → The DSV4 inference stack converges and AMD parts become usable.
Architecture evolution / breaking changes (read this before you upgrade):
- MRv2 matures: E/P/D disaggregation (#38390), weight offloading (#51413), multi-layer MTP KV (#50062).
- Tiered KV offloading to disk (#49644).
- bitsandbytes moved out to an out-of-tree plugin (#43529).
- Transformers upgraded to 5.15.0 (#51668).
⚠️ Deployment impact: this is an environment-level and dependency-level change.
bitsandbytesno longer ships with the core, so the quantization path needs a separate install; Transformers 5.15.0 is a major-version jump that will very likely break custom model definitions. Schedule it as “rebuild image → smoke test → accuracy regression → canary.” Do not runpip install -U vllmin place.
SGLang (still v0.5.18, 08-22, 710 PRs)
No new tag in the last 72 hours, but main saw heavy commits on 08-26 / 08-27.
| PR | What it does | Impact |
|---|---|---|
| #33561 | Support for Ling-3.0-flash (BailingMoeV3, Baichuan’s new flagship MoE) | Domestic MoE reaches first-class support |
| #31626 | Beam search support | Fills in the non-sampling decode path |
| #36233 | CUDA 13.4 container probing Rubin (sm_107) | Groundwork for next-generation hardware |
| #36397 | custom all-reduce v2 tuning | Communication kernel polishing |
| #35379 | Generalized hybrid SWA MTP draft-pool routing | Speculative decoding on hybrid architectures |
| #33871 | Cut idle DP work in breakable prefill CUDA graph | Removes spin |
| #35640 | Coordinate FullCG prefill across DP-attn ranks | PD + speculative polish |
Standing Topic: Two PD Disaggregation Roadmaps Are Diverging
| Dimension | vLLM (v0.28.0) | SGLang (main) |
|---|---|---|
| This period | E/P/D disaggregation, tiered KV offload to disk, canonical CPU layout | Cross-DP-attn FullCG prefill coordination, splitting custom all-reduce into push/pull planes, configurable weight-cache daemon path |
| Technical bent | Resource decoupling + external storage | Protocol / topology alignment + operational observability |
The split is clear: vLLM is asking “how far away can the KV live,” while SGLang is asking “how should a request be routed.” The former expands the capacity envelope; the latter improves scheduling intelligence. Long term these converge — once capacity is tiered, routing has to take into account which tier the data sits in.
PyTorch vs transformers: vLLM v0.28.0 upgrades to Transformers 5.15.0 (dual-backend convergence, chasing day-N breadth); SGLang keeps its in-house runtime (Rust server, self-managed kernel stack). No “off transformers” PRs this period. Conclusion unchanged: model definitions converge on transformers for breadth, while the data plane and runtime converge on in-house code for determinism — layered, not either/or.
Step adaptation: no new models for days on end (Step-3.7-Flash 06-01 / 328★, Step-3.5-Flash 04-03 / 2070★, vllm fork 05-28). The Step MTP PRs — vLLM #49642 / #49490 / #52115 / #53174 and SGLang #32325 (verified still open, untouched since 2026-07-24) / #35206 — are all open, and neither the v0.28.0 nor the v0.5.18 official notes contain a single Step entry.
Conclusion: Step can still only get speedups from prebuilt images — vllm/vllm-openai:stepfun37 plus an MTP config on the vLLM side, dev-step-3.7-flash plus EAGLE on the SGLang side. Full MTP support has not landed upstream, and upstream maintenance investment lags.
3. The One-Line Takeaway
v0.28.0 is the release most worth scheduling an upgrade for this quarter — the K3 stack frees ~17 GiB per GPU, DSV4 sparse MLA is end-to-end, DFlash2 is stable, and most of the gains (adaptive DSpark cutting TTFT by ~60%) need no changes to your code; but treat it as an environment-level change: bitsandbytes has moved out of tree, Transformers has jumped to 5.15.0, and a pip install -U in place will very likely earn you a phone call at 3 a.m.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。