系列:每日AI热点

Daily AI Hotspot · 2026-08-13: Speculative Decoding Starts Doing the Math — Fixed Draft Count Loses 33% at High Concurrency

Two threads run through today. On the research side it is “world models + long-context attention”: Dreamer V3 matches specialized SOTA on 150+ tasks with a single fixed set of hyperparameters, and DeepSeek’s CSA compresses the KV cache for a 1M-token context down to 10% of V3.2’s. On the engineering side there is a signal worth more attention: for the first time, speculative decoding has been publicly quantified as something that can lose you money.

★ Most Worth Your Attention Today

vLLM #47808: DSpark confidence-scheduled verification (merged to main 08-12).

Almost every piece of deployment folklore about speculative decoding rests on one unstated assumption: a larger draft budget k is better, or at worst “enabling it costs you nothing.” This PR demolishes that assumption with measurements — at concurrency c=256 with the GPU already saturated, DSpark with a fixed budget of 7 drafts is 33% slower than running with no speculation at all.

The mechanism is simple but easy to miss. The cost of verifying k drafts grows linearly in k, while the number of accepted tokens is capped by the acceptance rate and grows ever more slowly. At low concurrency the GPU has slack, verification overhead hides in idle time, and “bigger k is better” appears true. Once concurrency saturates the GPU, verification cost lands directly on the critical path, and every extra draft you verify steals compute from real decoding.

ScenarioFixed k=7Adaptive (confidence-based)
Low concurrency (c≤64)within ±3% of adaptivebaseline
High concurrency (c=256)33% slower than no speculationdeliberately drops accept length from ~3.9 to 3.5 to protect the win

Three implementation details are worth stealing: a Triton kernel ranks positions by the cumulative product (cumprod) of per-position confidence and picks the prefix that maximizes accepted-tokens-per-millisecond; a per-request EMA (α=0.8) smooths step-to-step jitter; and the decode path moves to varlen CUDA graphs so variable-length windows stop costing launch overhead.

Deployment impact: if you run DSpark at high concurrency, you must enable adaptive verification budgets — otherwise simply “turning on speculation” is actively slowing you down. The reverse also holds: don’t rush to set k=1. At low concurrency the two approaches differ by only ±3%. The real dividing line is whether the GPU is saturated, not the absolute value of k.

1. AI Industry & Paper Highlights

Paper: Dreamer V3 — world-model RL goes cross-domain

Operator: CSA (Compressed Sparse Attention in DeepSeek V4)

Performance: knowledge distillation

A teacher’s soft labels supervise a student, cutting parameters by 10–50x for a 1–3% accuracy loss. Representative work: DeepSeek R1 distillation (671B→7B, MATH-500 58.8%→92.8%), Meta’s Muse Glimmer (logit distillation, SWE-Bench 76.0 plus DFlash at 3.1x), and DistilBERT (−40% parameters / 60% faster / 97% retained). The implication for embodied AI is direct: distillation is the critical path for on-device VLA deployment — distill a cloud-scale model into a lightweight version for Jetson, which is exactly what GR00T’s “three-computer architecture” does.

Industry roundup

Company / eventKey numbersRead
Unitree Robotics (688836)IPO 8,288x oversubscribed, DeepSeek allocated 141M CNY, raising 6.099B CNYThe pricing anchor for the first humanoid stock on the A-share market
AgiBot (maps to 688585)WITA-Omni tops DailyOmni at 85.21Omni-modal VLA validated
Huilun Tech (incubated by GAC)100M+ CNY round; GoMate Mini deployed in 50 units, 63,000 km accumulatedAutomotive-grade supply chain transfer
Industry overallH1 2026 global humanoid shipments 19,100 units (+272%), Chinese vendors at 97%, AgiBot + Unitree at 75% combinedGoldman Sachs forecasts 76,000 units in 2027 and 502,000 by 2032
Foundation modelsMeta open-sources Muse Glimmer (30B agent model, SWE-Bench 76.0); Grok 4.6 launches (Elo above GPT-5.6); NVIDIA open-sources Alpamayo 2 Super (34B autonomous-driving model)Open agent and driving models scaling in parallel

2. vLLM & SGLang Community Tracking

vLLM

SGLang

Standing topic: Step-series support (8th consecutive day without progress)

Two MTP PRs remain open and unmerged: vLLM #49642 (draft, last activity 08-06, step3.7-flash bounds check) and SGLang #32325 (last activity 07-24, NVFP4 MTP shared-head). In the StepFun-ai org, Step-3.7-Flash has been frozen since 06-01 and Step-3.5-Flash since 04-03; the most recent org activity is Step-Realtime-CLI (08-10).

For contrast, Ling 3.0 Flash’s MTP was merged in the same window — the gap is upstream maintenance investment, not model capability. The standing conclusion is unchanged: Step’s performance story leans heavily on its bundled MTP draft heads, and MTP is precisely the least stable part.

3. The One-Line Takeaway

Speculative decoding’s narrative is shifting from “how do we raise throughput” to “how do we actually account for this” — vLLM #47808 supplies the first public negative-return data point (at c=256, fixed k is 33% slower than no speculation), while SGLang #27689 and vLLM #51738 are busy removing synchronization overhead from the verification path; if you run DSpark at high concurrency and have never measured a “speculation off” control group, today is the day to establish that baseline.


Sources: vLLM v0.27.1 release; vLLM PRs #50424 / #47808 / #51726 / #51738 / #51311 / #51255 / #47017 / #49642; SGLang PRs #33807 / #27689 / #34443 / #34524 / #33997 / #34642 / #32921 / #32325; StepFun-ai org.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。