Today brings a rare piece of symmetric evidence: the same optimization idea — MoE shared-expert fusion — produced double-digit gains on 08-20 on both the NVIDIA side (vLLM #53040) and the AMD side (SGLang #32340). When two vendors, two frameworks, and two kernel stacks produce gain curves of the same shape, you can be fairly confident the win comes from a structural cause — eliminating fixed overhead — rather than from one vendor’s proprietary kernel magic. The second thread is observability: vLLM #48915 writes per-request speculative decoding acceptance rates into the OpenAI-compatible response body.
★ Most Worth Your Attention Today
Shared-expert fusion: vLLM #53040 (NVIDIA) and SGLang #32340 (AMD) both land on the same day.
Start with what vLLM #53040 does. MoE models in the DeepSeek V4 family run two expert paths per layer: routed experts (FP4) and shared experts (FP8, which every token passes through). Before fusion, that is four kernel launches:
- routed FP4 MegaMoE
- shared FP8 gate/up
- shared down
- add (residual)
Between those four launches, intermediate results are repeatedly written back to HBM and read again. After fusion, a single SM100 persistent MegaMoE kernel schedules everything: the shared L1 segment, routed dispatch plus MMA, the shared L2 segment, and FP32 accumulation, writing the BF16 output exactly once.
Measured numbers (B200/GB200, DSv4, 128/256 outputs):
| Metric | Before | After | Change |
|---|---|---|---|
| BS=1 output throughput | 130.02 tok/s | 149.50 tok/s | +14.98% |
| BS=1 TPOT | 7.653 ms | 6.643 ms | −13.19% |
| Throughput at concurrency 64 | — | — | +9.44% |
| TTFT, 128-token short input | — | — | +4.93% (worse) |
How to read this table: the bottleneck in decode is bandwidth and launch overhead, not compute. That is why small batches gain the most — the TPOT improvement at BS=1 (−13.19%) is clearly larger than the throughput improvement at concurrency 64 (+9.44%), because the larger the batch, the more thinly launch overhead is amortized and the smaller the share left to save. And TTFT gets 4.93% worse because on a very short prefill the persistent kernel’s startup cost cannot be amortized — the same mechanism seen from the other side, not a bug.
On the AMD side, #32340 fixes two assumptions inherited from the V3 era (bf16 correction bias, and top-6 not being a power of two), delivering +8–11% output throughput at low concurrency on MI355X / TP4 with GSM8K accuracy preserved. The gain curve has the same shape as on the NVIDIA side.
Direct deployment impact: both paths ship behind a flag and are off by default.
- B200/GB200 running DSv4: add
--moe-backend deep_gemm_mega_moefor a double-digit win.- MI355X running V3/V4: upgrade to a build containing #32340. Leave them off and you are simply giving up 10%+. But mind the TTFT regression: if your workload is short-input, low-output, and sensitive to first-token latency, benchmark before enabling.
1. AI Industry & Paper Highlights
Paper: ResNet (Deep Residual Learning for Image Recognition)
- One-line positioning: skip connections plus residual learning pushed network depth from 22 layers to 152+, solving vanishing gradients and degradation, with >200K citations.
- Three core ideas: (1) learn the residual
F(x) = H(x) − xrather thanH(x)directly — asF → 0the network degrades to an identity mapping, which guarantees that “deeper cannot be worse”; (2) skip connections cost zero extra parameters; (3) 1×1 bottleneck dimensionality reduction makes a 152-layer network cheaper to compute than VGG-19. - Why it matters: it spawned ResNeXt, SENet, EfficientNet, and ConvNeXt. More importantly, the skip-connection idea crossed domains — residual connections in Transformers, U-Nets in diffusion, and vision encoders in VLAs all carry its imprint. Without residual skip connections there would be no trainable deep vision-action models today. In a year when everyone talks about MoE and sparse attention, revisiting ResNet has a specific value: it demonstrates that one sufficiently simple structural change raises the ceiling more effectively than stacking parameters.
Operator: DeepNorm (stable deep training in GLM)
- Positioning: lets a Transformer train to 1000 layers by amplifying the residual path with α > 1 so gradients stay stable at depth.
- Key formula:
DeepNorm(x) = LayerNorm(α·x + SubLayer(x)), whereα = (2N)^(1/4), withβ = (3M²)^(−1/4)scaling initialization. - Shipping models: GLM-130B / GLM-4 / GLM-5 (Zhipu, standard on 100+ layer models); DeepSeek V4 (a kindred residual-amplification idea); Qwen3 (Pre-LN route, but the warmup strategy echoes it).
- Three key points: (1) zero extra parameters — only the two α / β scaling factors; (2) Pre-LN models do not need it — don’t copy it blindly; (3) β scales only V/O and the FFN’s second linear, not Q/K — get this detail wrong and training collapses outright.
Performance: gradient checkpointing (activation recomputation)
- Positioning: activations occupy 60–80% of training memory; this trades compute for memory — the forward pass saves activations for only some layers, and the backward pass recomputes the rest.
- Mechanism: save a checkpoint every √N layers; memory drops from
O(N·d²)toO(√N·d²), i.e. 1/√N, at a cost of +33% compute. - Representative work: PyTorch
torch.utils.checkpoint(7B goes from 8×A100 to 4×A100); Megatron Selective Recomputation (70% memory saved for only 10–15% more compute, standard for GPT-3 175B); DeepSeek V4 (gradient checkpointing + FP8 + ZeRO-3, 284B trained on 256×H200). - Why it matters for embodied AI: cutting memory by √N directly lowers the barrier to training large models — VLA 7B fine-tuning becomes feasible on a single device, which is decisive for small teams doing domain-specific robotics fine-tuning.
Industry roundup
- Unitree Robotics (688836): three days after listing — up +460% on day one to close at 845 CNY, trading at 675.89 CNY (−1.62%) today, +348% cumulative over three days, with 2.195B CNY of net main-capital inflow, first in the entire market. Positive for supply chain names 688017 / 002050 / 601689.
- UBTech (09880.HK): showed a full-stack industrial solution at WRC and announced the first super-factory with 10,000-unit annual capacity, built with Siemens and entering production in August, with +176% on-device VLA inference. At roughly 43B HKD market cap versus Unitree’s 340B CNY, there is room for valuation repair.
- Beijing Humanoid Innovation Center: at WRC it released the Pelican-Unify unified embodied world model plus the Tiangong Omni robot, with 20+ companies from 6 countries signing onto the “Kaiwu ecosystem.”
- GLM 5.3 released (8/18): 83 on LMArena with a 1M context — a direct product of the DeepNorm operator above. Chinese open-source models now account for >63.5% of global API call volume across DeepSeek and Qwen.
- Nanjing University’s “Guangyuxin”: the world’s first ultra-low-power intelligent vision sensor chip, converting optical signals directly into tokens without analog-to-digital conversion, improving energy efficiency by >10×. If it reaches production, it could disrupt the entire on-device visual perception pipeline.
- Industry overall: according to MIIT, 2025 robotics industry revenue passed 300B CNY (+20% CAGR over 5 years), reaching 165.5B CNY (+24.5%) in H1 2026; global humanoid shipments of 19,100 units, with China at 97%.
2. vLLM & SGLang Community Tracking
No new version tags from either project in the last 72 hours (vLLM v0.27.1 / SGLang v0.5.17).
vLLM
- #48915 (new feature, key): adds
--per-request-spec-decode-metrics summary|detailed, so the response’smetrics.speculative_decodingreturns per-requestmean_acceptance_length,acceptance_histogram, anddraft_acceptance_rate.This is underrated. Previously, when speculative decoding underperformed, you got only a service-wide average and could not tell whether “every request has a low acceptance rate” or “one traffic class is dragging the average down.” With a per-request histogram you can bucket by traffic type — for instance, long-output agent traffic and short-output classification traffic may differ by a factor of two in acceptance rate, a fact that was previously invisible. This is the dividing line between suspecting an acceptance-rate drop and measuring one.
- #52998 (default change): TP CUDA groups now enable FlashInfer all-reduce by default;
VLLM_ALLREDUCE_USE_FLASHINFER=0opts out, andVLLM_BATCH_INVARIANT=1forces exclusion.Anyone tracking main must check which communication backend the startup log actually selected — the danger of a default change is that you assume nothing changed.
- #52466: the Mooncake Store consumer gains
save_decode_cache, so decode-phase KV also lands in the shared store; #51777: nixl upgraded to 1.3.2; #51362: fixes sparse Mamba block table saving. Multi-turn and agent workloads can now resume across instances without recomputation.
SGLang
- #27770 (PD disaggregation): decode-side radix cache for SWA hybrid models (a unified radix tree) — full-attention KV is reused across turns while SWA transfers only the fresh window, cutting redundant P/D KV transfers. Note it is mutually exclusive with HiCache / Mamba / DSA.
- #33370 (PD routing, large change): in-process KV Indexer plus Router integration (42 files, +6605). The path runs worker ZMQ → bridge → gRPC → indexer →
MatchExternalKvPrefix→ sgl-router, removing the Redis / Dragonfly external dependency. One design point worth stealing: an indexer failure costs only cache affinity, not availability — the degradation path is explicit rather than a total outage. - #35496 (speculative decoding): the DFlash2 selector now supports a quantized target lm_head (switching to
quant_method.applywith padded tails masked to −inf, avoiding both theCHECK_INPUTfailure in FlashInfer’s radix top-k and hangs at TP>1).Impact: a 32GB card running Qwen3.8-27B quantized weights can finally turn on DFlash2. But you need a build from main — it is not in a release yet.
- #32340 (AMD performance): see ★ above. Also: #29525, the DeepEPv2 ElasticBuffer A2A backend, was merged on 08-19 and reverted by #35568 just 2h20m later — a good illustration of the risk density in communication-layer changes.
Standing topic: Step-series support (first substantive progress of this series)
vLLM #53174 (opened 08-20T23:13, still open) fixes both the Step-3.5 MTP startup crash and the reasoning parser — the first substantive move on the Step front since this tracking began. More importantly, it supplies the first quantitative evidence:
With only the MTP fix and not the parser fix, Step-3.5-Flash-FP8 under
tool_choice: requiredreturned a valid tool call on just 18 of 50 requests (36%), with 30 requests producingno_tool_call_empty.
What that number means: the problem with Step models on vLLM is not only visible failures like “crashes at startup,” but also silent functional gaps where it runs yet 64% of requests ignore the required format. Anyone focused solely on the MTP draft head would miss that half entirely.
The coupling cost is also articulated clearly for the first time: Transformers v5 requires layer_types to have length equal to num_hidden_layers, whereas Step-3.5-Flash is 48 vs 45 (three extra MTP tail layers). The earlier #38247 simply truncated to pass validation, which made the MTP layer index layer_types[num_hidden_layers+i] raise IndexError at startup. #53174 changes it so that truncation applies only during super().__init__() validation, after which 45+3 is restored.
All other Step PRs remain open (vLLM #52115 / #49490 / #40070 / #49642; SGLang #35206 / #32325). In the StepFun-ai org, only Step-Realtime-CLI saw a push on 08-20; Step-3.7-Flash has been frozen since 06-01, Step-3.5-Flash since 04-03, and the official vLLM fork since 05-28, with no new open-source models or official deployment guide updates this window.
Methodological takeaway: model definition converges on transformers for breadth, while the data plane converges on in-house code for determinism — that “layered, not either-or” call is right. But the seams between layers (config validation, vocabulary padding, layer-type tables) are precisely where failures concentrate. Step’s crash did not happen inside either layer; it happened at the seam.
3. The One-Line Takeaway
Shared-expert fusion was independently validated on NVIDIA and AMD on the same day (+14.98% throughput and −13.19% TPOT at BS=1; +8–11% on MI355X), and the matching gain-curve shapes tell you the win comes from eliminating fixed launch and HBM round-trip overhead — structural, not accidental — so anyone running DSv4 on B200/GB200 should add --moe-backend deep_gemm_mega_moe today; and with vLLM #48915 putting per-request acceptance rates into the response body, speculative decoding finally moves from “service-wide average guesswork” to “attribute it per traffic bucket” — those two together are the week’s genuinely actionable gains.
Sources: vLLM v0.27.1 release (2026-08-11); vLLM PRs #53040 / #48915 / #52998 / #52466 / #51777 / #51362 / #53174 / #38247; SGLang PRs #27770 / #33370 / #35496 / #32340 / #29525 / #35568; SGLang v0.5.17 release (2026-08-08); GitHub commits API (vllm-project/vllm, sgl-project/sglang, since 2026-08-19T01:00:00Z); StepFun-ai org push timestamps.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。