Two threads run through today. On the research side it is “world models + long-context attention”: Dreamer V3 matches specialized SOTA on 150+ tasks with a single fixed set of hyperparameters, and DeepSeek’s CSA compresses the KV cache for a 1M-token context down to 10% of V3.2’s. On the engineering side there is a signal worth more attention: for the first time, speculative decoding has been publicly quantified as something that can lose you money.
★ Most Worth Your Attention Today
vLLM #47808: DSpark confidence-scheduled verification (merged to main 08-12).
Almost every piece of deployment folklore about speculative decoding rests on one unstated assumption: a larger draft budget k is better, or at worst “enabling it costs you nothing.” This PR demolishes that assumption with measurements — at concurrency c=256 with the GPU already saturated, DSpark with a fixed budget of 7 drafts is 33% slower than running with no speculation at all.
The mechanism is simple but easy to miss. The cost of verifying k drafts grows linearly in k, while the number of accepted tokens is capped by the acceptance rate and grows ever more slowly. At low concurrency the GPU has slack, verification overhead hides in idle time, and “bigger k is better” appears true. Once concurrency saturates the GPU, verification cost lands directly on the critical path, and every extra draft you verify steals compute from real decoding.
| Scenario | Fixed k=7 | Adaptive (confidence-based) |
|---|---|---|
| Low concurrency (c≤64) | within ±3% of adaptive | baseline |
| High concurrency (c=256) | 33% slower than no speculation | deliberately drops accept length from ~3.9 to 3.5 to protect the win |
Three implementation details are worth stealing: a Triton kernel ranks positions by the cumulative product (cumprod) of per-position confidence and picks the prefix that maximizes accepted-tokens-per-millisecond; a per-request EMA (α=0.8) smooths step-to-step jitter; and the decode path moves to varlen CUDA graphs so variable-length windows stop costing launch overhead.
Deployment impact: if you run DSpark at high concurrency, you must enable adaptive verification budgets — otherwise simply “turning on speculation” is actively slowing you down. The reverse also holds: don’t rush to set k=1. At low concurrency the two approaches differ by only ±3%. The real dividing line is whether the GPU is saturated, not the absolute value of k.
1. AI Industry & Paper Highlights
Paper: Dreamer V3 — world-model RL goes cross-domain
- One-line positioning: the first world-model RL algorithm to match or beat specialized SOTA on 150+ tasks using a single fixed set of hyperparameters.
- Three core ideas: (1) learn environment dynamics with an RSSM (deterministic
h_tplus discrete stochasticz_t); (2) use a symlog transform to unify reward scales across domains — this is what makes “one hyperparameter set for everything” possible, because the real difficulty in cross-domain tuning is reward magnitudes differing by orders of magnitude; (3) roll the policy out in imagination, consuming zero real interaction. - Why it matters: the bottleneck in embodied AI is data. The “one billion hour data gap” the industry keeps citing cannot be closed by collecting on real robots. Synthetic data from world models is the most promising path we have, and Dreamer V3 shows the approach generalizes at the algorithm level rather than being a one-off tuned inside a single environment.
Operator: CSA (Compressed Sparse Attention in DeepSeek V4)
- Positioning: compresses O(L²) attention to sub-quadratic. At 1M tokens, V4 uses only 27% of the FLOPs and 10% of the KV cache of V3.2.
- Shipping models: DeepSeek V4-Pro (61 layers interleaving CSA and HCA, k=1024) and V4-Flash (43 layers, k=512).
- Three key points: (1) KV compression at m=4 → top-k sparsity via an FP4 Lightning Indexer → plus a +128-token sliding window; (2) HCA (dense, m=128) interleaves with CSA to balance efficiency and expressiveness; (3) FP4 coarse filtering reaches 99.7% recall. That last number is the one worth copying: it validates that “low-precision coarse filter + high-precision exact compute” holds for attention, not just for quantization folklore.
Performance: knowledge distillation
A teacher’s soft labels supervise a student, cutting parameters by 10–50x for a 1–3% accuracy loss. Representative work: DeepSeek R1 distillation (671B→7B, MATH-500 58.8%→92.8%), Meta’s Muse Glimmer (logit distillation, SWE-Bench 76.0 plus DFlash at 3.1x), and DistilBERT (−40% parameters / 60% faster / 97% retained). The implication for embodied AI is direct: distillation is the critical path for on-device VLA deployment — distill a cloud-scale model into a lightweight version for Jetson, which is exactly what GR00T’s “three-computer architecture” does.
Industry roundup
| Company / event | Key numbers | Read |
|---|---|---|
| Unitree Robotics (688836) | IPO 8,288x oversubscribed, DeepSeek allocated 141M CNY, raising 6.099B CNY | The pricing anchor for the first humanoid stock on the A-share market |
| AgiBot (maps to 688585) | WITA-Omni tops DailyOmni at 85.21 | Omni-modal VLA validated |
| Huilun Tech (incubated by GAC) | 100M+ CNY round; GoMate Mini deployed in 50 units, 63,000 km accumulated | Automotive-grade supply chain transfer |
| Industry overall | H1 2026 global humanoid shipments 19,100 units (+272%), Chinese vendors at 97%, AgiBot + Unitree at 75% combined | Goldman Sachs forecasts 76,000 units in 2027 and 502,000 by 2032 |
| Foundation models | Meta open-sources Muse Glimmer (30B agent model, SWE-Bench 76.0); Grok 4.6 launches (Elo above GPT-5.6); NVIDIA open-sources Alpamayo 2 Super (34B autonomous-driving model) | Open agent and driving models scaling in parallel |
2. vLLM & SGLang Community Tracking
vLLM
- Release: v0.27.1 (08-11), a patch release adding support for quantized DSpark Markov heads (#50424 — the draft head’s
markov_w2can now run at W4A16). Anyone running DSpark can upgrade and cut cost: draft heads used to be stored in floating point, so quantizing saves both memory and bandwidth. - Speculative decoding economics: #47808, confidence-scheduled verification (see ★ above).
- Config change: #51726 raises the default
_max_num_batched_tokensfrom 8192 to 16384, doubling the per-batch token cap. Throughput goes up, but so do memory and scheduling pressure — you must re-tunemax_num_seqsafter upgrading, or long-input traffic will hit the memory wall. - Performance / models: #51738 eliminates another GPU↔CPU sync; #51311 Kimi K3 Flash KDA out kernel (1.1–1.4x on prefill); #51255 native Dots3 NOTE multimodal support; #47017 DeepSeek-V4 on ROCm gfx11; #51668 Transformers bumped to 5.15.0.
SGLang
- PD disaggregation: #33807 adds pipeline-parallel prefill plus a Mooncake staging buffer — another step toward production-grade long-context PD disaggregation.
- Speculative decoding performance: #27689 removes a blocking D2H sync from the spec-decode plan in the FlashInfer MLA backend, cutting per-step latency. This is the same theme as vLLM #51738: both projects are systematically cleaning synchronization points off the critical path this cycle.
- Correctness bug fixes: #34443 fixes a
num_splitscrash in DSA prefill-CP speculative decoding; #34524 fixes the causal default for DFlash sliding-window attention. Anyone running DSA/DFlash with CP must upgrade and re-validate — both are correctness defects that manifest as crashes or silent wrong answers. - Architecture / multimodal: #33997 upgrades FlashInfer to 0.6.17 and cleans up K3 temporary workarounds, but #33623’s K3 MLA gate fusion was reverted the same day by #34642 — K3 fusion kernels are still churning, so don’t build on them yet. Diffusion-side commits are dense (#32921 native SANA-Video T2V, MiniMax H3 LoRA, LTX-2 / ERNIE-Image / Krea-2 / Cosmos3).
Standing topic: Step-series support (8th consecutive day without progress)
Two MTP PRs remain open and unmerged: vLLM #49642 (draft, last activity 08-06, step3.7-flash bounds check) and SGLang #32325 (last activity 07-24, NVFP4 MTP shared-head). In the StepFun-ai org, Step-3.7-Flash has been frozen since 06-01 and Step-3.5-Flash since 04-03; the most recent org activity is Step-Realtime-CLI (08-10).
For contrast, Ling 3.0 Flash’s MTP was merged in the same window — the gap is upstream maintenance investment, not model capability. The standing conclusion is unchanged: Step’s performance story leans heavily on its bundled MTP draft heads, and MTP is precisely the least stable part.
3. The One-Line Takeaway
Speculative decoding’s narrative is shifting from “how do we raise throughput” to “how do we actually account for this” — vLLM #47808 supplies the first public negative-return data point (at c=256, fixed k is 33% slower than no speculation), while SGLang #27689 and vLLM #51738 are busy removing synchronization overhead from the verification path; if you run DSpark at high concurrency and have never measured a “speculation off” control group, today is the day to establish that baseline.
Sources: vLLM v0.27.1 release; vLLM PRs #50424 / #47808 / #51726 / #51738 / #51311 / #51255 / #47017 / #49642; SGLang PRs #33807 / #27689 / #34443 / #34524 / #33997 / #34642 / #32921 / #32325; StepFun-ai org.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。