★ Most Worth Your Attention Today
vLLM #52783: adaptive DSpark sizes the draft length by confidence on SM100 sparse MLA — +33.6% end-to-end throughput at concurrency c=256.
The number only lands properly when you read it backwards: with adaptivity turned off, a fixed draft length of 7 at high concurrency is actually slower than not using speculative decoding at all.
The reason is not mysterious. Verification cost for a fixed-k draft is a rigid expense, while its payoff depends on how much of the draft gets accepted. At low concurrency the GPU has slack, verification cost gets amortized, and fixed k pays off. Once concurrency saturates the GPU, verification flips from “amortized” to “net loss” — you spend more compute than you save.
What #52783 does:
- Rank prefixes by confidence
cumprod— verify the high-confidence draft segments first, spending budget where acceptance is likeliest. - EMA smoothing (α = 0.8) — keeps a single-step jitter from swinging the budget wildly.
- Varlen CUDA graph — variable-length batches still get graph capture instead of falling back to per-request launches.
- On SM100 sparse MLA, merge the sparse projection with the draft graph, eliminating a D2H copy.
Measured: ±3% at low concurrency (essentially neutral), gains preserved at high concurrency, +33.6% end-to-end throughput at c=256.
How to enable it: set VLLM_BATCH_INVARIANT, or turn on the adaptive budget.
⚠️ The second big item today is bad news. The Step adaptation story finally produced hard evidence. vLLM #38494 documents that Step-3.5-Flash’s MTP, because it unconditionally shares lm_head, sees acceptance rate collapse from ~97% to ~4% — silently, with no error raised.
If you run Step multi-GPU MTP on upstream main, diff your accept_length right now. That number is not printed in your logs; it just makes your service inexplicably slow.
1. AI Industry & Paper Highlights
Paper: Mobile ALOHA (CoRL 2024, Stanford)
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation (Fu, Zhao, Finn). One-line positioning: low-cost mobile bimanual teleoperation plus co-training, turning complex mobile manipulation into a reproducible engineering recipe.
Three core ideas:
- Add a mobile base to the ALOHA bimanual system, total cost $32K; 16-DoF whole-body control at 50 Hz; only 50 demonstrations per task.
- Co-training: mix a small amount of mobile demonstration data (16-DoF) with a large volume of static ALOHA data (14-DoF) to train an ACT behavior-cloning model, using pad base = 0 to align the dimensionality gap.
- Exploit the shared structure of manipulation skills for cross-scene transfer: success rate 50% → 90% — wiping wine stains 100%, calling an elevator 95%, stowing a pot 85%.
Why it matters: it shows that low-cost hardware plus few demonstrations can accomplish complex tasks, and that co-training is a data-efficient paradigm. It is fully open source and directly pushed the democratization of mobile manipulation. Mapped to industrial deployment, the takeaway is practical: a little in-scene data plus general pretraining buys you scene transfer.
Operator: Load Balancing Auxiliary Loss (Aux Loss)
The standard defense against MoE expert collapse.
$$L_{balance} = \alpha \cdot N \cdot \sum (f_i \cdot P_i)$$
where $f_i$ is the fraction of tokens routed to expert $i$, $P_i$ is the mean routing probability, and α = 0.01.
- Qwen3-MoE: classic aux loss plus an expert capacity limit — two layers of protection.
- DeepSeek V4: an innovative auxiliary-loss-free approach — a dynamic bias term balances load without polluting the main gradient; routing activation moves from Sigmoid to Softplus; stable across the entire 671B / 37B training run.
The key trap: α too large (>0.1) over-constrains routing quality, forcing uniform allocation and destroying expert specialization; α too small (<0.001) and expert collapse happens anyway. Start at 0.01 and keep monitoring the variance of expert utilization.
Performance: 3D Parallelism (DP + TP + PP)
The infrastructure that makes 100B-scale, 1000-GPU training possible.
- DP: a full model replica per device, different data, gradients synchronized.
- TP: intra-layer weight matrices split, results stitched back via all-reduce.
- PP: layers cut into stages, micro-batches pipelined to fill the gaps.
Representative work: DeepSeek V4 (671B / 37B MoE) trained on 2,048 H800s with 3D parallelism + EP + FP8 + ZeRO-3 at a cost of $5.576M, with no loss spikes. Qwen3 uses Megatron-Core’s four-way TP+PP+DP+EP; GLM-5 uses DeepNorm plus 3D parallelism to train 1,000 layers stably across a thousand cards.
Traps: the PP bubble — too few micro-batches leaves GPUs idle; and TP groups larger than 8, where all-reduce communication cancels out the gains from splitting.
Industry Snapshot
- XPeng Robotics (private): $900M first round at a $6.3B valuation, a record single round for Chinese embodied AI; IRON enters mass production by year end. Positive for whole-machine and component chains (688017 / 002050 / 601689).
- AgiBot / Zhiyuan Robotics (Hong Kong IPO filing; A-share proxy 688585): the mass-production Lingxi X2 took gold in the 100 m obstacle race with zero modification; shipments of 9,700 units (43% share) — surpassing Unitree for the global #1 spot for the first time.
- Unitree Robotics-W (688836 STAR): down 45% four days after listing, closing at ¥613 (+1.69%), margin balance ¥1.419B; last place in the Games’ 100 m, with the company saying its energy is going into mass production. Neutral near-term volatility; the long game is production volume.
- Silicon: OpenAI’s in-house Jalapeño chip beats NVIDIA’s GB300 on power (TDP 700W vs 1400W); NVIDIA’s Vera Rubin NVL72 delivers 30× throughput; Claude hits a 27% hit rate on protein design; Zhipu open-sources GLM-5.3.
The AgiBot numbers deserve their own look: 9,700 units / 43% share, passing Unitree for the first time. This is the first time the embodied AI industry has seen an unambiguous reshuffling of the share leaderboard — and the trigger was “a mass-production unit competing with zero modification.” Productization, not lab metrics, is becoming the dividing line.
2. vLLM & SGLang Community Tracking
Version status: no new tags from any of the three. vLLM stable v0.27.1 (08-11) / pre-release v0.28.0rc2 (08-21); SGLang v0.5.18 (08-22); no new Step models (the org’s core repos are frozen, Step-Realtime-CLI last touched 08-21).
vLLM
| PR | What it does | Impact |
|---|---|---|
| #52783 | Adaptive DSpark (SM100 sparse MLA) | +33.6% at c=256; no loss at high concurrency ★ |
| #53649 | Blackwell batch-invariant persistent matmul auto-tuning | 25.2% latency reduction (= 1.336×) |
| #52388 | Kimi-K3 Mamba metadata-preparation kernel | 6.6–7.6× faster |
| #52242 | DSpark gains logprobs (per-verified-token probability) | Observability for speculative decoding |
| #53615 | Transformers backend migration wrap-up | More models route through HF transformers v5 |
| #38494 | Step-3.5-Flash MTP hard evidence | Acceptance 97% → 4%, silent degradation ⚠️ |
Do not be fooled by the #53649 title: it says 33.6%, but the actual figure is a 25.2% latency reduction (i.e. a 1.336× speedup). These two numbers get conflated constantly — check before you cite.
SGLang
- #35314 SSD Expert Pack inference (paradigm-level): when expert weights exceed VRAM + host RAM, spill them to SSD and prefetch layer by layer. This is opt-in and limited in scope, but paradigm-level in direction — MoE parameter growth has outrun memory capacity, and extending the storage hierarchy one level down is an inevitability.
- #35505 Shared-expert fusion on GB200: P99 ITL −52.6%, with the largest gains at low concurrency. Low-concurrency optimization is usually neglected, yet it is exactly what governs perceived latency in interactive workloads.
- #36186 Nemotron 3.5 speculative decoding comparison table: quantifies the DSpark / EAGLE / MTP trade-offs into a lookup table. Documentation that “gives you the choice rather than the conclusion” is far more useful than simply recommending one algorithm.
- #36232 HiCache refactor: L2/L3 host pools plus a UnifiedRadixCache consolidation.
- Assorted PD disaggregation edge-case bugfixes.
Standing Topic: Layered Convergence Confirmed Again — By Two Opposite Moves
Both projects took a step today, in opposite directions, and together they confirm the same conclusion:
- vLLM #53615 migrates to HF transformers v5 — chasing day-N breadth with zero adaptation cost.
- SGLang #35314 manages its own SSD weight stack — chasing determinism.
The costs are symmetric too: the former must rebuild on every transformers upgrade, the latter forfeits the upstream path. So “pure PyTorch vs transformers” was never either/or — model definitions converge on transformers for breadth, while the data plane and runtime converge on in-house code for determinism.
Step adaptation: no new official PRs; the most substantive signal is the #38494 quantified evidence above. All upstream MTP PRs remain open, and the community can only get MTP speedups from the prebuilt stepfun37 image.
3. The One-Line Takeaway
Wiring adaptive scheduling into speculative decoding is the best value-per-effort change of the day — gains like +33.6% at c=256 need no model change at all; but go check your accept_length while you are at it, because Step’s 97%→4% silent degradation is a reminder that the most dangerous failure mode in speculative decoding was never an error message, it was “looks like it’s running, but it is accepting nothing.”
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。