系列:每日AI热点

Daily AI Hotspot · 2026-09-12: SGLang Lands Day-0 DeepSeek-V4.1 Support (Bounded Recompute for Less Cache — 8×H200 Prefill 1.56×, pass@1 Lossless), vLLM Holds at v0.29.0

★ Most Worth Your Attention Today

SGLang + Miles land Day-0 support for DeepSeek-V4.1 — “bounded recompute for less cache” pushes 8×H200 prefill to 1.56× with AIME 2026 pass@1 lossless (453/480), the only big update in the last 72h with both concrete benchmark numbers and source-level reading; vLLM holds at v0.29.0 (Blackwell end-to-end latency −33.6%, Hopper low-latency GEMM +12.9~25.2%) and drops 10 deprecated architectures.

Three layers of fact:

  1. SGLang + Miles shipped Day-0 DeepSeek-V4.1 support on 09-10: two schemes — encoder-side bounded replay / decoder-side tail-only — with 8×H200 prefill 1.56×, 4×GB300 1.37×, and AIME 2026 pass@1 lossless (453/480). This is not an “announcement” but a real update with measured pass rates and source paths.
  2. vLLM holds at v0.29.0 (09-09, no new tag in 72h) but delivers genuine performance dividends this cycle: generational kernel retuning brings Hopper low-latency GEMM +12.9~25.2% and Blackwell autotuning end-to-end latency −33.6%; at the same time it removes 10 deprecated architectures (Arctic / Chameleon / MPT) plus the old api_server path — a breaking change that requires migration first.
  3. SGLang + Mooncake raise the KV-pooling bar to a trillion tokens per day: >90% hit rate, batch read <50ms, 3× single-machine efficiency / 30× total capacity — PD disaggregation moves from a paper config to a production line.

Actionable conclusion: this is a boundary tradeoff of “what to cache vs what to recompute,” isomorphic to my own OpenInfer / Qwen3-4B DFlash speculative decoding — use bounded recompute to trade for less cache and higher hit. SGLang users can turn on V4.1 support directly to test prefill; vLLM users should validate parity on staging before upgrading (dropping 10 architectures is a hard breaking change).

Worth saying separately: the lead switches from vLLM (which shipped two consecutive releases, 0907/0910) to SGLang this round — Day-0 DeepSeek-V4.1 support of the “race the release window” kind is one of the moats of an inference framework. But vLLM’s Blackwell −33.6% / Hopper +12.9~25.2% are real, on-the-ground deployment dividends; don’t fixate only on release cadence.

2. vLLM & SGLang Community Tracking

Version status: vLLM v0.29.0 (formal, 09-09) remains the latest stable, no new tag in 72h; SGLang v0.5.19 (09-05) remains the latest stable, but on 09-10 it gained Day-0 DeepSeek-V4.1 support via Miles (not a tag — an ecosystem adaptation).

vLLM v0.29.0

New features / architecture evolution:

Breaking changes:

ChangeImpact
Removed 10 deprecated architectures (Arctic / Chameleon / MPT)Models depending on these must migrate first or fail to launch
Old api_server path removedLaunch scripts move to vllm serve
CUDA Graph VRAM profiling folded into KV auto-sizingAfter upgrade, re-check VRAM headroom and graph capture

SGLang v0.5.19 + DeepSeek-V4.1 Day-0 (09-10)

New features / major adaptation:

Production practice: SGLang + Mooncake at a trillion tokens per day — Prefill with HiCache on, Mooncake Store pooling node DRAM, KV hit rate >90%, batch read <50ms, 3× single-machine efficiency / 30× total capacity.

Under the Hood: DeepSeek-V4.1’s “Bounded Recompute for Less Cache” (SGLang implementation)

DeepSeek-V4.1 is a 92-layer hybrid of gated linear + full attention alternating. SGLang does not cache each layer’s window KV; instead it rebuilds via “bounded replay”:

Mechanism: ① decoder-side tail-only — layer 20 is the final compressed-KV source; layers 21–39 reuse its compressed KV + indexer keys, and each chunk computes only the last ≤128 tokens; ② encoder-side bounded replay — on a cache hit, recompute the prefix’s last 128 tokens to rebuild the window KV. Both bound the local-attention truncation at the rebuild boundary (a bounded approximation — must be measured in practice), both opt-in, and encoder replay does not support speculative decoding.

Insight: trade “bounded compute” for “less cache footprint + higher hit” → prefill 1.56× / 1.37×. This shares roots with speculative decoding’s “draft prior + verify” — use a little controlled recompute to save cache globally. It is isomorphic to my own OpenInfer / Qwen3-4B DFlash “lightweight draft + verify-accept” tradeoff — the lessons of cache boundaries are directly borrowable.

My read: V4.1 changes the answer to “what to cache” from “full window KV” to “compressed KV + bounded tail recompute,” in essence pushing the Pareto frontier of attention caching one step further. The decisive factor in inference frameworks increasingly lands on “KV boundary management” rather than simply stacking operators.

Standing Topics

PD disaggregation: SGLang + Mooncake trillion-token production practice (Prefill with HiCache on, hit >90%, read <50ms); vLLM MultiConnector supports XpYd (P2P MooncakeConnector + shared pool MooncakeStoreConnector, “read first-come-first-served, write broadcast”); llm-d DisaggregatedSet unifies orchestration of NIXL / Mooncake / MoRIIO.

Architecture evolution: V4.1 SWA bounded replay; SGLang DeepEP v2 + W4A8 MoE + LayerNorm SP; vLLM MRV2 default + batch-sharded sampling + CUDA graph memory folded into auto-sizing + per-request speculative metrics.

PyTorch vs transformers: no new “leave transformers” PR this cycle; the signal leans reverse — vLLM moves FlexOlmo / Olmo3 / Hunyuan back to the Transformers backend; an enterprise benchmark dropped transformers+accelerate for vLLM+SGLang on Qwen3.6-35B-A3B (P99 3.2s→412ms, root cause static-tuple deep copy of past_key_values). Conclusion: the in-house stack wins on KV management and batch scheduling.

Step adaptation: no new adaptation PR in 72h; step-3.7-flash (198B / 11B sparse MoE, 256K, native multimodal + tool calling) self-hosted on vLLM stepfun37 (MTP) with maturity > SGLang dev- image (EAGLE), NVFP4 about 4 cards.

3. AI Papers & Industry Hotspots

Highlight (one line)

Meta’s SAM turns “segmentation” from one-model-per-class into a promptable general capability, proving vision can have foundation models; it is exactly the visual front-end for robots to “see the world” — complementing today’s “act” embodied-industry dynamics.

Paper core (SAM / Meta AI, arXiv:2304.02643, ICCV 2023)

Operator explainer: QK-Norm

Performance optimization: ZeRO / FSDP (sharded data parallel)

Industry hotspots (embodied-intelligence companies / chain speed)

⚠️ Industry dynamics do not constitute investment advice.

4. The One-Line Takeaway

SGLang + Miles land Day-0 DeepSeek-V4.1 (bounded recompute for less cache, 8×H200 prefill 1.56×, pass@1 lossless) — the only big update this round with both benchmarks and source-level reading; vLLM holds at v0.29.0 yet delivers the deployment dividends of Blackwell −33.6% / Hopper +12.9~25.2% and drops 10 deprecated architectures (hard breaking); SGLang + Mooncake push KV pooling to a trillion tokens per day at >90% hit; on the research side SAM turns segmentation into a promptable general interface as the robot visual front-end, QK-Norm cures loss spikes, ZeRO/FSDP trade communication for VRAM to support 100B VLA training, while UBTECH’s ¥50M export, Zhiyuan’s 44% share, Unitree’s volume-up-profit-down, and DeepSeek V4.1-Flash cutting KV to 1/4 each tag the “see + act” embodied thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。