系列:每日AI热点

Daily AI Hotspot · 2026-08-26: Adaptive DSpark Delivers +33.6% Throughput at c=256, While Step's MTP Acceptance Rate Silently Collapses from 97% to 4%

★ Most Worth Your Attention Today

vLLM #52783: adaptive DSpark sizes the draft length by confidence on SM100 sparse MLA — +33.6% end-to-end throughput at concurrency c=256.

The number only lands properly when you read it backwards: with adaptivity turned off, a fixed draft length of 7 at high concurrency is actually slower than not using speculative decoding at all.

The reason is not mysterious. Verification cost for a fixed-k draft is a rigid expense, while its payoff depends on how much of the draft gets accepted. At low concurrency the GPU has slack, verification cost gets amortized, and fixed k pays off. Once concurrency saturates the GPU, verification flips from “amortized” to “net loss” — you spend more compute than you save.

What #52783 does:

  1. Rank prefixes by confidence cumprod — verify the high-confidence draft segments first, spending budget where acceptance is likeliest.
  2. EMA smoothing (α = 0.8) — keeps a single-step jitter from swinging the budget wildly.
  3. Varlen CUDA graph — variable-length batches still get graph capture instead of falling back to per-request launches.
  4. On SM100 sparse MLA, merge the sparse projection with the draft graph, eliminating a D2H copy.

Measured: ±3% at low concurrency (essentially neutral), gains preserved at high concurrency, +33.6% end-to-end throughput at c=256.

How to enable it: set VLLM_BATCH_INVARIANT, or turn on the adaptive budget.

⚠️ The second big item today is bad news. The Step adaptation story finally produced hard evidence. vLLM #38494 documents that Step-3.5-Flash’s MTP, because it unconditionally shares lm_head, sees acceptance rate collapse from ~97% to ~4% — silently, with no error raised.

If you run Step multi-GPU MTP on upstream main, diff your accept_length right now. That number is not printed in your logs; it just makes your service inexplicably slow.

1. AI Industry & Paper Highlights

Paper: Mobile ALOHA (CoRL 2024, Stanford)

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation (Fu, Zhao, Finn). One-line positioning: low-cost mobile bimanual teleoperation plus co-training, turning complex mobile manipulation into a reproducible engineering recipe.

Three core ideas:

  1. Add a mobile base to the ALOHA bimanual system, total cost $32K; 16-DoF whole-body control at 50 Hz; only 50 demonstrations per task.
  2. Co-training: mix a small amount of mobile demonstration data (16-DoF) with a large volume of static ALOHA data (14-DoF) to train an ACT behavior-cloning model, using pad base = 0 to align the dimensionality gap.
  3. Exploit the shared structure of manipulation skills for cross-scene transfer: success rate 50% → 90% — wiping wine stains 100%, calling an elevator 95%, stowing a pot 85%.

Why it matters: it shows that low-cost hardware plus few demonstrations can accomplish complex tasks, and that co-training is a data-efficient paradigm. It is fully open source and directly pushed the democratization of mobile manipulation. Mapped to industrial deployment, the takeaway is practical: a little in-scene data plus general pretraining buys you scene transfer.

Operator: Load Balancing Auxiliary Loss (Aux Loss)

The standard defense against MoE expert collapse.

$$L_{balance} = \alpha \cdot N \cdot \sum (f_i \cdot P_i)$$

where $f_i$ is the fraction of tokens routed to expert $i$, $P_i$ is the mean routing probability, and α = 0.01.

The key trap: α too large (>0.1) over-constrains routing quality, forcing uniform allocation and destroying expert specialization; α too small (<0.001) and expert collapse happens anyway. Start at 0.01 and keep monitoring the variance of expert utilization.

Performance: 3D Parallelism (DP + TP + PP)

The infrastructure that makes 100B-scale, 1000-GPU training possible.

Representative work: DeepSeek V4 (671B / 37B MoE) trained on 2,048 H800s with 3D parallelism + EP + FP8 + ZeRO-3 at a cost of $5.576M, with no loss spikes. Qwen3 uses Megatron-Core’s four-way TP+PP+DP+EP; GLM-5 uses DeepNorm plus 3D parallelism to train 1,000 layers stably across a thousand cards.

Traps: the PP bubble — too few micro-batches leaves GPUs idle; and TP groups larger than 8, where all-reduce communication cancels out the gains from splitting.

Industry Snapshot

  1. XPeng Robotics (private): $900M first round at a $6.3B valuation, a record single round for Chinese embodied AI; IRON enters mass production by year end. Positive for whole-machine and component chains (688017 / 002050 / 601689).
  2. AgiBot / Zhiyuan Robotics (Hong Kong IPO filing; A-share proxy 688585): the mass-production Lingxi X2 took gold in the 100 m obstacle race with zero modification; shipments of 9,700 units (43% share) — surpassing Unitree for the global #1 spot for the first time.
  3. Unitree Robotics-W (688836 STAR): down 45% four days after listing, closing at ¥613 (+1.69%), margin balance ¥1.419B; last place in the Games’ 100 m, with the company saying its energy is going into mass production. Neutral near-term volatility; the long game is production volume.
  4. Silicon: OpenAI’s in-house Jalapeño chip beats NVIDIA’s GB300 on power (TDP 700W vs 1400W); NVIDIA’s Vera Rubin NVL72 delivers 30× throughput; Claude hits a 27% hit rate on protein design; Zhipu open-sources GLM-5.3.

The AgiBot numbers deserve their own look: 9,700 units / 43% share, passing Unitree for the first time. This is the first time the embodied AI industry has seen an unambiguous reshuffling of the share leaderboard — and the trigger was “a mass-production unit competing with zero modification.” Productization, not lab metrics, is becoming the dividing line.

2. vLLM & SGLang Community Tracking

Version status: no new tags from any of the three. vLLM stable v0.27.1 (08-11) / pre-release v0.28.0rc2 (08-21); SGLang v0.5.18 (08-22); no new Step models (the org’s core repos are frozen, Step-Realtime-CLI last touched 08-21).

vLLM

PRWhat it doesImpact
#52783Adaptive DSpark (SM100 sparse MLA)+33.6% at c=256; no loss at high concurrency ★
#53649Blackwell batch-invariant persistent matmul auto-tuning25.2% latency reduction (= 1.336×)
#52388Kimi-K3 Mamba metadata-preparation kernel6.6–7.6× faster
#52242DSpark gains logprobs (per-verified-token probability)Observability for speculative decoding
#53615Transformers backend migration wrap-upMore models route through HF transformers v5
#38494Step-3.5-Flash MTP hard evidenceAcceptance 97% → 4%, silent degradation ⚠️

Do not be fooled by the #53649 title: it says 33.6%, but the actual figure is a 25.2% latency reduction (i.e. a 1.336× speedup). These two numbers get conflated constantly — check before you cite.

SGLang

Standing Topic: Layered Convergence Confirmed Again — By Two Opposite Moves

Both projects took a step today, in opposite directions, and together they confirm the same conclusion:

The costs are symmetric too: the former must rebuild on every transformers upgrade, the latter forfeits the upstream path. So “pure PyTorch vs transformers” was never either/or — model definitions converge on transformers for breadth, while the data plane and runtime converge on in-house code for determinism.

Step adaptation: no new official PRs; the most substantive signal is the #38494 quantified evidence above. All upstream MTP PRs remain open, and the community can only get MTP speedups from the prebuilt stepfun37 image.

3. The One-Line Takeaway

Wiring adaptive scheduling into speculative decoding is the best value-per-effort change of the day — gains like +33.6% at c=256 need no model change at all; but go check your accept_length while you are at it, because Step’s 97%→4% silent degradation is a reminder that the most dangerous failure mode in speculative decoding was never an error message, it was “looks like it’s running, but it is accepting nothing.”


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。