★ Most Worth Your Attention Today
SGLang + Miles land Day-0 support for DeepSeek-V4.1 — “bounded recompute for less cache” pushes 8×H200 prefill to 1.56× with AIME 2026 pass@1 lossless (453/480), the only big update in the last 72h with both concrete benchmark numbers and source-level reading; vLLM holds at v0.29.0 (Blackwell end-to-end latency −33.6%, Hopper low-latency GEMM +12.9~25.2%) and drops 10 deprecated architectures.
Three layers of fact:
- SGLang + Miles shipped Day-0 DeepSeek-V4.1 support on 09-10: two schemes — encoder-side bounded replay / decoder-side tail-only — with 8×H200 prefill 1.56×, 4×GB300 1.37×, and AIME 2026 pass@1 lossless (453/480). This is not an “announcement” but a real update with measured pass rates and source paths.
- vLLM holds at v0.29.0 (09-09, no new tag in 72h) but delivers genuine performance dividends this cycle: generational kernel retuning brings Hopper low-latency GEMM +12.9~25.2% and Blackwell autotuning end-to-end latency −33.6%; at the same time it removes 10 deprecated architectures (Arctic / Chameleon / MPT) plus the old api_server path — a breaking change that requires migration first.
- SGLang + Mooncake raise the KV-pooling bar to a trillion tokens per day: >90% hit rate, batch read <50ms, 3× single-machine efficiency / 30× total capacity — PD disaggregation moves from a paper config to a production line.
Actionable conclusion: this is a boundary tradeoff of “what to cache vs what to recompute,” isomorphic to my own OpenInfer / Qwen3-4B DFlash speculative decoding — use bounded recompute to trade for less cache and higher hit. SGLang users can turn on V4.1 support directly to test prefill; vLLM users should validate parity on staging before upgrading (dropping 10 architectures is a hard breaking change).
Worth saying separately: the lead switches from vLLM (which shipped two consecutive releases, 0907/0910) to SGLang this round — Day-0 DeepSeek-V4.1 support of the “race the release window” kind is one of the moats of an inference framework. But vLLM’s Blackwell −33.6% / Hopper +12.9~25.2% are real, on-the-ground deployment dividends; don’t fixate only on release cadence.
2. vLLM & SGLang Community Tracking
Version status: vLLM v0.29.0 (formal, 09-09) remains the latest stable, no new tag in 72h; SGLang v0.5.19 (09-05) remains the latest stable, but on 09-10 it gained Day-0 DeepSeek-V4.1 support via Miles (not a tag — an ecosystem adaptation).
vLLM v0.29.0
New features / architecture evolution:
- Model Runner V2 full default (closed out): CUDA Graph VRAM profiling folded into KV cache auto-sizing → upgrades must re-check model compatibility / graph capture / VRAM headroom.
- batch-sharded sampling: per-step logits VRAM drops to 1/TP, raising the VRAM ceiling for high-concurrency sampling.
- Generational kernel retuning: Hopper low-latency GEMM +12.9~25.2%, Blackwell autotuning end-to-end latency −33.6%.
- Mamba prefix cache keeps the prefill checkpoint, TTFT improves 9~25%; Kimi-K3 latent tail folded into Mamba metadata.
- New models: Hy4-preview (Tencent 770B / 49B MoE), Qwen3.8-Flash-Next, and others.
Breaking changes:
| Change | Impact |
|---|---|
| Removed 10 deprecated architectures (Arctic / Chameleon / MPT) | Models depending on these must migrate first or fail to launch |
Old api_server path removed | Launch scripts move to vllm serve |
| CUDA Graph VRAM profiling folded into KV auto-sizing | After upgrade, re-check VRAM headroom and graph capture |
SGLang v0.5.19 + DeepSeek-V4.1 Day-0 (09-10)
New features / major adaptation:
- DeepSeek-V4.1 Day-0 support (LMSYS 09-10): encoder-side bounded replay / decoder-side tail-only; 8×H200 prefill 1.56×, 4×GB300 1.37×, AIME 2026 pass@1 lossless (453/480).
- Lean attention on by default for MI300X / MI355X: long / heterogeneous decode throughput ≤1.52×, ITL down 3.62×; GLM-5.2 8×MI355X TPOT 23→8ms.
- Unified radix tree becomes the default prefix cache (Breaking); DeepEP v2 ElasticBuffer (
--moe-a2a-backend deepep_v2) supports cross-node MoE decode into CUDA graph. --enable-layernorm-sp(Qwen3-8B prefill H100 −3.5% / B200 −5.6%); Hopper W4A8 MoE pushes DeepSeek-V4-Flash output throughput +12%.
Production practice: SGLang + Mooncake at a trillion tokens per day — Prefill with HiCache on, Mooncake Store pooling node DRAM, KV hit rate >90%, batch read <50ms, 3× single-machine efficiency / 30× total capacity.
Under the Hood: DeepSeek-V4.1’s “Bounded Recompute for Less Cache” (SGLang implementation)
DeepSeek-V4.1 is a 92-layer hybrid of gated linear + full attention alternating. SGLang does not cache each layer’s window KV; instead it rebuilds via “bounded replay”:
Mechanism: ① decoder-side tail-only — layer 20 is the final compressed-KV source; layers 21–39 reuse its compressed KV + indexer keys, and each chunk computes only the last ≤128 tokens; ② encoder-side bounded replay — on a cache hit, recompute the prefix’s last 128 tokens to rebuild the window KV. Both bound the local-attention truncation at the rebuild boundary (a bounded approximation — must be measured in practice), both opt-in, and encoder replay does not support speculative decoding.
Insight: trade “bounded compute” for “less cache footprint + higher hit” → prefill 1.56× / 1.37×. This shares roots with speculative decoding’s “draft prior + verify” — use a little controlled recompute to save cache globally. It is isomorphic to my own OpenInfer / Qwen3-4B DFlash “lightweight draft + verify-accept” tradeoff — the lessons of cache boundaries are directly borrowable.
My read: V4.1 changes the answer to “what to cache” from “full window KV” to “compressed KV + bounded tail recompute,” in essence pushing the Pareto frontier of attention caching one step further. The decisive factor in inference frameworks increasingly lands on “KV boundary management” rather than simply stacking operators.
Standing Topics
PD disaggregation: SGLang + Mooncake trillion-token production practice (Prefill with HiCache on, hit >90%, read <50ms); vLLM MultiConnector supports XpYd (P2P MooncakeConnector + shared pool MooncakeStoreConnector, “read first-come-first-served, write broadcast”); llm-d DisaggregatedSet unifies orchestration of NIXL / Mooncake / MoRIIO.
Architecture evolution: V4.1 SWA bounded replay; SGLang DeepEP v2 + W4A8 MoE + LayerNorm SP; vLLM MRV2 default + batch-sharded sampling + CUDA graph memory folded into auto-sizing + per-request speculative metrics.
PyTorch vs transformers: no new “leave transformers” PR this cycle; the signal leans reverse — vLLM moves FlexOlmo / Olmo3 / Hunyuan back to the Transformers backend; an enterprise benchmark dropped transformers+accelerate for vLLM+SGLang on Qwen3.6-35B-A3B (P99 3.2s→412ms, root cause static-tuple deep copy of past_key_values). Conclusion: the in-house stack wins on KV management and batch scheduling.
Step adaptation: no new adaptation PR in 72h; step-3.7-flash (198B / 11B sparse MoE, 256K, native multimodal + tool calling) self-hosted on vLLM stepfun37 (MTP) with maturity > SGLang dev- image (EAGLE), NVFP4 about 4 cards.
3. AI Papers & Industry Hotspots
Highlight (one line)
Meta’s SAM turns “segmentation” from one-model-per-class into a promptable general capability, proving vision can have foundation models; it is exactly the visual front-end for robots to “see the world” — complementing today’s “act” embodied-industry dynamics.
Paper core (SAM / Meta AI, arXiv:2304.02643, ICCV 2023)
- One-line positioning: give a point / box / mask / text prompt and zero-shot segment the target without retraining — turning segmentation into a “promptable” general interface.
- Core idea ① (three-part decoupling): Image Encoder (ViT-H/14, ~632M, runs once per image) + Prompt Encoder (light) + Mask Decoder (2-layer Transformer, ~4M, millisecond) — the expensive part runs once, the interactive part is cheap to poke.
- Core idea ② (ambiguity-friendly): output 3 masks at once (whole / part / subpart), training back-propagates the min-loss;
L = 20·L_focal + L_dice + L_mse. - Core idea ③ (data engine): SA-1B = 11M images / 1.1B masks (~400× larger than the previous biggest segmentation set), three-round iteration (human → model-assisted → auto-generated + spot-checked).
- Impact: the dawn of visual foundation models; for embodied = the visual front-end of VLA, zero-shot segmenting targets for the policy network to pick grasp points.
Operator explainer: QK-Norm
- Positioning: apply RMSNorm to Q / K per head before attention (
q̂=γ⊙q/RMS(q)), fixing ‖q̂‖ → bounded logits → unsaturated softmax → cures large-model loss spikes. - Order and cost: Norm before RoPE, fp32 inside Norm; only +1–2% step time, trading purely for stability.
- In use: Qwen3, Gemma 2/3, OLMo 2, ViT-22B (the SAM-backbone ViT family); DeepSeek MLA does not use it.
Performance optimization: ZeRO / FSDP (sharded data parallel)
- Positioning: fit a 100B model into VRAM. VRAM budget: Adam training takes 16Ψ per parameter (fp16 weight 2 + grad 2 + fp32 master 4 + Adam m 4 + v 4).
- Mechanism: ZeRO-1/2/3 shard optimizer / gradient / parameter respectively, each card ≈ 16Ψ/N (Z3); FSDP = PyTorch-native ZeRO-3.
- Essence: trade communication for VRAM, the 3D-parallel foundation for 100B models. Embodied link = lower VLA training cost, quantization / distillation = lower edge cost.
Industry hotspots (embodied-intelligence companies / chain speed)
- [Embodied] UBTECH (09880.HK): another ¥50M overseas order (multiple Walker S2), China’s humanoid export accelerating → bullish for whole machines and the “brain”, maps to the Huaxia robot ETF (562500).
- [Embodied] Zhiyuan Robotics / Shangwei New Material (688585·A-share): YTD shipments >8400 units, ~44% global share, #1 → bullish for the A-share Zhiyuan-mapping platform.
- [Embodied] Unitree (688836·A-share): ~5900 units shipped, ~31% share, but Q1 deducted non-net profit −52.55% YoY — “volume up, profit down” warns of a price war; watch gross margin.
- [Chain] XPeng high-end humanoid line, Tesla Optimus bulk orders, Zhiyuan AGILE 2.0 end-to-end sense-control model landing → bullish for actuators (Sanhua 002050 / Top Group 601689), reducers (Lead 688017), robot ETF (562500).
- [General] DeepSeek V4.1-Flash cuts KV Cache to 1/4; GPT-6 Astra training breaks 100k cards; Kimi K2.8; Amap ABot-Earth 0.7 world model; the Bund Conference wind shifts from “dialogue” to “getting work done”.
⚠️ Industry dynamics do not constitute investment advice.
4. The One-Line Takeaway
SGLang + Miles land Day-0 DeepSeek-V4.1 (bounded recompute for less cache, 8×H200 prefill 1.56×, pass@1 lossless) — the only big update this round with both benchmarks and source-level reading; vLLM holds at v0.29.0 yet delivers the deployment dividends of Blackwell −33.6% / Hopper +12.9~25.2% and drops 10 deprecated architectures (hard breaking); SGLang + Mooncake push KV pooling to a trillion tokens per day at >90% hit; on the research side SAM turns segmentation into a promptable general interface as the robot visual front-end, QK-Norm cures loss spikes, ZeRO/FSDP trade communication for VRAM to support 100B VLA training, while UBTECH’s ¥50M export, Zhiyuan’s 44% share, Unitree’s volume-up-profit-down, and DeepSeek V4.1-Flash cutting KV to 1/4 each tag the “see + act” embodied thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。