★ Most Worth Your Attention Today
vLLM 0.28’s adaptive speculative token budget: a pure scheduling change, zero model cost, DSpark TTFT improved 55–65%.
This is the one item that should go straight onto your to-do list — because it costs you no model cost at all, only changes scheduling logic, yet cuts DSpark’s TTFT by more than half.
Three layers of fact:
- Mechanism: instead of hard-coding a fixed budget per speculative step, it allocates on demand — fewer drafts for easy prefixes, more for hard ones.
- Gain: DSpark TTFT improves 55–65%, and introduces no extra model weights and no extra VRAM.
- Comparison: alongside vLLM 0.28’s EAGLE/MTP, this “adaptive budget” route is the only one of the three speculative-decoding roads with zero incremental cost.
Actionable conclusion: if you already run DSpark, upgrade to vLLM 0.28 and turn on the adaptive budget — the TTFT layer gets a 55%+ improvement almost for free; no model change, no hardware change.
Worth emphasizing: speculative decoding’s dividend is spilling from “text generation” into broader autoregressive scenarios (yesterday’s Step-Audio RTF was already a signal). The adaptive budget hands the “whether to speculate, and how much” decision to the runtime itself — a key step in productizing this route, turning acceleration from “tuning art” into “default behavior.”
2. vLLM & SGLang Community Tracking
Version status: vLLM v0.28.0 (8/26); SGLang v0.5.18 (8/22). No new release in the last 72 hours, but the community dug out actionable detail on both engines.
vLLM v0.28.0
New features:
- Sparsity attention end-to-end (DeepSeek-V4 / Kimi-K3) — long context is no longer a brute-force VRAM fight.
- Adaptive speculative token budget → DSpark TTFT improved 55–65% (zero model cost, see ★).
- Model Runner V2: E/P/D disaggregation + weight offloading, pulling “scheduling” and “compute” further apart.
- Tiered KV offload adds a disk tier — long-context agent cold sessions no longer occupy HBM permanently.
max_num_batched_tokensdoubled to 16384.
Breaking / default changes: max_num_batched_tokens 8192 → 16384; single-GPU / small-VRAM setups can OOM — re-measure capacity before upgrading.
SGLang v0.5.18
Performance:
- Cold-start weight staging: Qwen3-32B bring-up 84.8s → 35.6s (~2.4× faster).
- DSpark speculative: 383.7 tok/s @ V4-Pro TP8.
- PD disaggregation production-ready: 5×TP deployments.
Under the Hood: Why the Adaptive Budget Is “Zero Cost”
Moving the speculative budget from “fixed” to “adaptive” eliminates the redundancy you used to reserve for the worst case every time.
The cost of a fixed budget: easy prefixes (“hi”, “continue”) are speculated at the hard-prefix budget, wasting draft compute; hard prefixes may be under-speculated and miss on hit rate. The adaptive budget allocates by actual difficulty on the fly, essentially handing “speculative efficiency” to the runtime for online optimization — which is exactly why it lands a 55–65% TTFT improvement via scheduling alone, at zero model cost.
My read: this route pays off most for online scenarios where first-token latency is sensitive but request-difficulty distribution is very uneven (agents, streaming completions). It asks for no model swap and no extra GPU — the highest cost-performance tier of acceleration you can get.
Standing Topics
PD disaggregation: Dynamo 1.0 / llm-d / Mooncake / NIXL keep evolving; on the SGLang side it is DCP decode-context parallelism, FlashInfer all-to-all MoE routing, and DeepSeek-V4 FlashMLA sparse prefill on by default.
Architecture evolution: three speculative-decoding roads (EAGLE / MTP / adaptive budget — vLLM’s adaptive budget improves DSpark TTFT by 55–65%), sparse MLA, FP8/INT4/KV quantization, and multimodality.
Pure PyTorch vs transformers: no new “off transformers, in-house stack” PR this cycle; the dependency-decoupling trend continues (vLLM moves Transformers to 5.15.0 + externalizes bitsandbytes).
Step adaptation:
- Step-3.7-Flash open weights + official vLLM / SGLang images are out; NVFP4 runs on 4 GPUs; MTP / EAGLE speculative; hit price ¥0.27/M (the Flash price anchor has come down).
- For specifics, defer to the official StepFun repositories.
3. AI Papers & Industry Hotspots
Highlight (one line)
Open X-Embodiment (RT-X) uses 22 robot types and 1M+ episodes of cross-embodiment data plus embodiment-embedding conditioning to prove that robots obey a “data scaling law” too — the technical bedrock of the “embodied-intelligence engine” investment thesis.
Paper core (Open X-Embodiment / RT-X)
- One-line positioning: the first open-source effort at large-scale cross-embodiment joint training that yields a transferable “robot foundation model.”
- Core idea (3 lines): ① unify 22 robot types into the OXE dataset; ② tag each sample with an embodiment embedding rₑ as a condition so one set of weights shares skills across embodiments; ③ RT-1-X modulates with FiLM, RT-2-X splices rₑ as a token into the VLA.
- Impact: data heterogeneity + conditioning = cross-embodiment generalization (unseen-task generalization 2–3×), directly supporting re-rating of whole-machine makers and component suppliers.
Operator explainer: Expert Parallelism (EP)
- Positioning: MoE experts sliced across GPUs, tokens routed via all-to-all, solving “can’t fit on one card / uneven compute.”
- In production: DeepSeek V4 (cross-node EP + aux-loss-free bias), Qwen3-MoE, MiniMax.
- Key point: vs naive implementation, memory sharing + parallelism shifts the bottleneck to all-to-all bandwidth; the pitfalls are load imbalance and communication overhead.
Performance optimization: Adaptive Early-exit
- Positioning: simple samples exit at an intermediate layer once the head’s confidence > τ, no need to run all L layers.
- Representative work: FastBERT / DeeBERT / CALM; gains 1.5–3× compute / latency.
- Embodied link: VLA early-exit on simple observations → lower robot response latency.
Industry hotspots (embodied-intelligence companies / chain speed)
- Unitree (688836): listed on the STAR Market 8/19; 9/2 close 547 (-4.21%), ~50% drawdown from its high; strategic placement includes DeepSeek (locked 36 months); H1 revenue 1.152B (+48.5%). Maps to: A-share humanoid valuation anchor.
- AgiBot / Zhiyuan (controls Shangwei New Material 688585): sprinting toward a HK IPO at a 400–500B HKD valuation; H1 shipments ~8,000–9,000 units. Maps to: the 688585 capitalization platform.
- Horizon Robotics (09660.HK): Journey 6M in mass production (20+ carmakers / 70+ models / 128TOPS urban NOA). Maps to: dual smart-driving + robotics thesis.
- UBTech (09880.HK): H1 revenue 1.27B (+104.2%), full-size humanoid 590M (+1445%), U1 orders 13,300 units.
- General: the Flash-model battlefield (DeepSeek V4 Flash / Zhipu GLM-5.3-Flash / Alibaba Qwen3.8-Flash / Tencent Hy4); smart-driving VLA + world models (Horizon 6M / Momenta R7 / WorldVLA).
4. The One-Line Takeaway
vLLM 0.28 lands sparsity attention end-to-end and uses a pure-scheduling adaptive speculative budget to cut DSpark TTFT by a free 55–65% — the acceleration most worth deploying this cycle; on the research side RT-X uses 22 robot types to confirm a cross-embodiment data scaling law, pulling the “embodied-intelligence engine” from thematic narrative to an investment thesis with a technical bedrock — while Unitree’s ~50% drawdown, AgiBot’s HK-IPO sprint, and Horizon’s mass production are each tagging that thesis with a price.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。