系列:每日AI热点

Daily AI Hotspot · 2026-08-21: Shared-Expert Fusion Lands on NVIDIA and AMD the Same Day, +14.98% Throughput at BS=1

Today brings a rare piece of symmetric evidence: the same optimization idea — MoE shared-expert fusion — produced double-digit gains on 08-20 on both the NVIDIA side (vLLM #53040) and the AMD side (SGLang #32340). When two vendors, two frameworks, and two kernel stacks produce gain curves of the same shape, you can be fairly confident the win comes from a structural cause — eliminating fixed overhead — rather than from one vendor’s proprietary kernel magic. The second thread is observability: vLLM #48915 writes per-request speculative decoding acceptance rates into the OpenAI-compatible response body.

★ Most Worth Your Attention Today

Shared-expert fusion: vLLM #53040 (NVIDIA) and SGLang #32340 (AMD) both land on the same day.

Start with what vLLM #53040 does. MoE models in the DeepSeek V4 family run two expert paths per layer: routed experts (FP4) and shared experts (FP8, which every token passes through). Before fusion, that is four kernel launches:

  1. routed FP4 MegaMoE
  2. shared FP8 gate/up
  3. shared down
  4. add (residual)

Between those four launches, intermediate results are repeatedly written back to HBM and read again. After fusion, a single SM100 persistent MegaMoE kernel schedules everything: the shared L1 segment, routed dispatch plus MMA, the shared L2 segment, and FP32 accumulation, writing the BF16 output exactly once.

Measured numbers (B200/GB200, DSv4, 128/256 outputs):

MetricBeforeAfterChange
BS=1 output throughput130.02 tok/s149.50 tok/s+14.98%
BS=1 TPOT7.653 ms6.643 ms−13.19%
Throughput at concurrency 64——+9.44%
TTFT, 128-token short input——+4.93% (worse)

How to read this table: the bottleneck in decode is bandwidth and launch overhead, not compute. That is why small batches gain the most — the TPOT improvement at BS=1 (−13.19%) is clearly larger than the throughput improvement at concurrency 64 (+9.44%), because the larger the batch, the more thinly launch overhead is amortized and the smaller the share left to save. And TTFT gets 4.93% worse because on a very short prefill the persistent kernel’s startup cost cannot be amortized — the same mechanism seen from the other side, not a bug.

On the AMD side, #32340 fixes two assumptions inherited from the V3 era (bf16 correction bias, and top-6 not being a power of two), delivering +8–11% output throughput at low concurrency on MI355X / TP4 with GSM8K accuracy preserved. The gain curve has the same shape as on the NVIDIA side.

Direct deployment impact: both paths ship behind a flag and are off by default.

  • B200/GB200 running DSv4: add --moe-backend deep_gemm_mega_moe for a double-digit win.
  • MI355X running V3/V4: upgrade to a build containing #32340. Leave them off and you are simply giving up 10%+. But mind the TTFT regression: if your workload is short-input, low-output, and sensitive to first-token latency, benchmark before enabling.

1. AI Industry & Paper Highlights

Paper: ResNet (Deep Residual Learning for Image Recognition)

Operator: DeepNorm (stable deep training in GLM)

Performance: gradient checkpointing (activation recomputation)

Industry roundup

2. vLLM & SGLang Community Tracking

No new version tags from either project in the last 72 hours (vLLM v0.27.1 / SGLang v0.5.17).

vLLM

SGLang

Standing topic: Step-series support (first substantive progress of this series)

vLLM #53174 (opened 08-20T23:13, still open) fixes both the Step-3.5 MTP startup crash and the reasoning parser — the first substantive move on the Step front since this tracking began. More importantly, it supplies the first quantitative evidence:

With only the MTP fix and not the parser fix, Step-3.5-Flash-FP8 under tool_choice: required returned a valid tool call on just 18 of 50 requests (36%), with 30 requests producing no_tool_call_empty.

What that number means: the problem with Step models on vLLM is not only visible failures like “crashes at startup,” but also silent functional gaps where it runs yet 64% of requests ignore the required format. Anyone focused solely on the MTP draft head would miss that half entirely.

The coupling cost is also articulated clearly for the first time: Transformers v5 requires layer_types to have length equal to num_hidden_layers, whereas Step-3.5-Flash is 48 vs 45 (three extra MTP tail layers). The earlier #38247 simply truncated to pass validation, which made the MTP layer index layer_types[num_hidden_layers+i] raise IndexError at startup. #53174 changes it so that truncation applies only during super().__init__() validation, after which 45+3 is restored.

All other Step PRs remain open (vLLM #52115 / #49490 / #40070 / #49642; SGLang #35206 / #32325). In the StepFun-ai org, only Step-Realtime-CLI saw a push on 08-20; Step-3.7-Flash has been frozen since 06-01, Step-3.5-Flash since 04-03, and the official vLLM fork since 05-28, with no new open-source models or official deployment guide updates this window.

Methodological takeaway: model definition converges on transformers for breadth, while the data plane converges on in-house code for determinism — that “layered, not either-or” call is right. But the seams between layers (config validation, vocabulary padding, layer-type tables) are precisely where failures concentrate. Step’s crash did not happen inside either layer; it happened at the seam.

3. The One-Line Takeaway

Shared-expert fusion was independently validated on NVIDIA and AMD on the same day (+14.98% throughput and −13.19% TPOT at BS=1; +8–11% on MI355X), and the matching gain-curve shapes tell you the win comes from eliminating fixed launch and HBM round-trip overhead — structural, not accidental — so anyone running DSv4 on B200/GB200 should add --moe-backend deep_gemm_mega_moe today; and with vLLM #48915 putting per-request acceptance rates into the response body, speculative decoding finally moves from “service-wide average guesswork” to “attribute it per traffic bucket” — those two together are the week’s genuinely actionable gains.


Sources: vLLM v0.27.1 release (2026-08-11); vLLM PRs #53040 / #48915 / #52998 / #52466 / #51777 / #51362 / #53174 / #38247; SGLang PRs #27770 / #33370 / #35496 / #32340 / #29525 / #35568; SGLang v0.5.17 release (2026-08-08); GitHub commits API (vllm-project/vllm, sgl-project/sglang, since 2026-08-19T01:00:00Z); StepFun-ai org push timestamps.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。