系列:每日AI热点

Daily AI Hotspot · 2026-09-02: vLLM's Adaptive Speculative Budget Cuts DSpark TTFT 55–65% at Zero Model Cost; RT-X Confirms Robots Obey a Data Scaling Law

★ Most Worth Your Attention Today

vLLM 0.28’s adaptive speculative token budget: a pure scheduling change, zero model cost, DSpark TTFT improved 55–65%.

This is the one item that should go straight onto your to-do list — because it costs you no model cost at all, only changes scheduling logic, yet cuts DSpark’s TTFT by more than half.

Three layers of fact:

  1. Mechanism: instead of hard-coding a fixed budget per speculative step, it allocates on demand — fewer drafts for easy prefixes, more for hard ones.
  2. Gain: DSpark TTFT improves 55–65%, and introduces no extra model weights and no extra VRAM.
  3. Comparison: alongside vLLM 0.28’s EAGLE/MTP, this “adaptive budget” route is the only one of the three speculative-decoding roads with zero incremental cost.

Actionable conclusion: if you already run DSpark, upgrade to vLLM 0.28 and turn on the adaptive budget — the TTFT layer gets a 55%+ improvement almost for free; no model change, no hardware change.

Worth emphasizing: speculative decoding’s dividend is spilling from “text generation” into broader autoregressive scenarios (yesterday’s Step-Audio RTF was already a signal). The adaptive budget hands the “whether to speculate, and how much” decision to the runtime itself — a key step in productizing this route, turning acceleration from “tuning art” into “default behavior.”

2. vLLM & SGLang Community Tracking

Version status: vLLM v0.28.0 (8/26); SGLang v0.5.18 (8/22). No new release in the last 72 hours, but the community dug out actionable detail on both engines.

vLLM v0.28.0

New features:

Breaking / default changes: max_num_batched_tokens 8192 → 16384; single-GPU / small-VRAM setups can OOM — re-measure capacity before upgrading.

SGLang v0.5.18

Performance:

Under the Hood: Why the Adaptive Budget Is “Zero Cost”

Moving the speculative budget from “fixed” to “adaptive” eliminates the redundancy you used to reserve for the worst case every time.

The cost of a fixed budget: easy prefixes (“hi”, “continue”) are speculated at the hard-prefix budget, wasting draft compute; hard prefixes may be under-speculated and miss on hit rate. The adaptive budget allocates by actual difficulty on the fly, essentially handing “speculative efficiency” to the runtime for online optimization — which is exactly why it lands a 55–65% TTFT improvement via scheduling alone, at zero model cost.

My read: this route pays off most for online scenarios where first-token latency is sensitive but request-difficulty distribution is very uneven (agents, streaming completions). It asks for no model swap and no extra GPU — the highest cost-performance tier of acceleration you can get.

Standing Topics

PD disaggregation: Dynamo 1.0 / llm-d / Mooncake / NIXL keep evolving; on the SGLang side it is DCP decode-context parallelism, FlashInfer all-to-all MoE routing, and DeepSeek-V4 FlashMLA sparse prefill on by default.

Architecture evolution: three speculative-decoding roads (EAGLE / MTP / adaptive budget — vLLM’s adaptive budget improves DSpark TTFT by 55–65%), sparse MLA, FP8/INT4/KV quantization, and multimodality.

Pure PyTorch vs transformers: no new “off transformers, in-house stack” PR this cycle; the dependency-decoupling trend continues (vLLM moves Transformers to 5.15.0 + externalizes bitsandbytes).

Step adaptation:

3. AI Papers & Industry Hotspots

Highlight (one line)

Open X-Embodiment (RT-X) uses 22 robot types and 1M+ episodes of cross-embodiment data plus embodiment-embedding conditioning to prove that robots obey a “data scaling law” too — the technical bedrock of the “embodied-intelligence engine” investment thesis.

Paper core (Open X-Embodiment / RT-X)

Operator explainer: Expert Parallelism (EP)

Performance optimization: Adaptive Early-exit

Industry hotspots (embodied-intelligence companies / chain speed)

4. The One-Line Takeaway

vLLM 0.28 lands sparsity attention end-to-end and uses a pure-scheduling adaptive speculative budget to cut DSpark TTFT by a free 55–65% — the acceleration most worth deploying this cycle; on the research side RT-X uses 22 robot types to confirm a cross-embodiment data scaling law, pulling the “embodied-intelligence engine” from thematic narrative to an investment thesis with a technical bedrock — while Unitree’s ~50% drawdown, AgiBot’s HK-IPO sprint, and Horizon’s mass production are each tagging that thesis with a price.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。