Both sides of the fence are chasing the same theme today — making speculative decoding finer-grained — but from opposite directions. SGLang’s DFlash2 improves draft quality (draft tokens stop being blind to one another), while vLLM defuses a production landmine (#52836 reverts a workspace-reuse optimization that could silently produce wrong results). On the research side there is DeepMind’s Genie, which turns a single image into a playable environment with actions discovered entirely through self-supervision. On the industry side, Unitree lists this week: the A-share market gets its first humanoid robotics stock.
★ Most Worth Your Attention Today
SGLang #35371: DFlash2 — local grouped depthwise convolution plus a candidate selector (merged 08-19, +929 / −61).
Start with the problem. Draft heads in the DFlash family carry an unavoidable constraint: positions inside a draft block are generated in parallel, so position i cannot see the draft token that position i−1 just produced. When you guess the 5th token, you do not yet know what you guessed for the 4th. For an autoregressive model this is a serious structural handicap — the longer the draft, the more the tail is pure blind guessing.
DFlash2 patches the hole with two complementary changes.
Change one: grouped dynamic depthwise convolution. A depthwise convolution along the position dimension is added inside the draft block, letting each proposal position look directly at representations from earlier positions:
out[i, c] = Σ_t (base[t, c] + δ[i, t, g(c)]) · x[i−t, c]
base is a learnable base kernel; δ is a position-dependent dynamic increment shared within channel groups g(c); taps at block boundaries are zeroed so nothing leaks across blocks. The benefit is direct: a proposal position gains access to prior context without paying for another backbone pass. The cost is essentially one cheap convolution, and the return is a real improvement in draft quality.
Change two: the candidate selector. The conventional approach takes an argmax at every slot, pinning the draft path to a single sequence. DFlash2 instead keeps top-K at every slot and then scores transitions with
edge(p→c) = ⟨A[p] ⊙ project(h), B[c]⟩ + unary[c]
walking an optimal path over the candidate graph from the verified anchor. When sampling at T>0 it returns K candidates via the inverse CDF for lossless verification — the crucial detail, because it means raising the sampling temperature no longer costs you verification correctness.
Both changes are folded into the draft CUDA graph at decode time, so no extra launch overhead is introduced.
Why this matters more than a percentage: two years of speculative-decoding optimization have gone almost entirely into the verification side — verify less, stop early, hide the syncs. DFlash2 is a rare case of attention returning to the draft side. If draft quality itself does not improve, every verification-side optimization is just refilling a leaking bucket. The two ideas here are also portable: “parallel-generated drafts can’t see each other” is a problem in every blockwise draft head, not something specific to SGLang.
1. AI Industry & Paper Highlights
Paper: Genie (DeepMind, 11B parameters, arXiv:2402.15391)
- One-line positioning: a foundation world model that generates an interactive environment from a single image, trained without any human action labels.
- Three-part structure: (1) a Video Tokenizer (VQ-VAE) compresses frames into discrete tokens; (2) a Latent Action Model discovers actions self-supervised — the cleverest step in the paper, since it requires no annotated actions in the data and instead infers a discrete action space from differences between adjacent frames; (3) a Dynamics Model, a causal Transformer, predicts the next frame.
- Training data: 30,000 hours of unlabeled gameplay video.
- Why it matters: the expensive part of robot Sim2Real is not the algorithm, it is the environment. Real-world collection is slow, wears out hardware, and is hard to reproduce; hand-built simulators demand substantial modeling labor. Genie offers a third path: generate a cheap, infinitely parallelizable, playable environment straight from one image. For embodied AI that reframes rollout cost from “robot-hours” to “GPU-hours.” Genie 2 has already extended the approach to general video domains.
Operator: MQA (Multi-Query Attention, Shazeer 2019)
- Positioning: all query heads share a single K/V pair, compressing the inference KV cache by h× (h = number of heads).
- Shipping models: PaLM, LLaMA-2, Falcon. On-device VLA deployment benefits most.
- Key number: LLaMA-2-7B’s KV cache drops from roughly 1.0 MB/token to 0.13 MB/token.
- The caveat: this is not free. MQA trades KV head diversity for memory, and quality loss on small models must be measured, not assumed safe just because large models degrade gracefully. On-device VLAs are precisely the sub-7B models where this warning applies most.
Performance: pruning
- Positioning: remove redundant weights, neurons, or attention heads to shrink model size and compute directly.
- Representative work: SparseGPT (second-order approximation, 30–50% sparsity with near-zero loss), LLM-Pruner (structured pruning plus LoRA recovery), Sheared LLaMA (treating pruning itself as a training objective).
- Gains: LLaMA-2-7B at 50% sparsity degrades perplexity by only 7%; OpenVLA 7B with INT4 plus pruning runs on a Jetson Orin Nano.
- Relation to distillation and quantization: the three are distinct levers for on-device deployment — quantization changes numeric representation, pruning changes structural sparsity, distillation changes the source of knowledge. The OpenVLA result is a combination of all three, and is currently the most realistic recipe for embodied edge deployment.
Industry roundup
- Unitree Robotics (688836): the first humanoid robotics stock on the A-share market, expected to list this week, at an issue price of 150.8 CNY and a 61B CNY valuation, with DeepSeek, Tencent, and the social security fund in the strategic allocation. Maps to robotics ETFs, Leader Harmonious Drive Systems, and Sanhua Intelligent Controls.
- AgiBot: expected to list in Hong Kong in August at a 40–50B HKD valuation, having shipped 8,400 units in H1, overtaking Unitree. Maps to Shangwei New Materials (688585).
- Symbiotic Zhixing: released a Unitree G1 driving-a-kart demo using DPC end-to-end whole-body control. Its significance is directional rather than commercial — it tests whether the VLA route holds up on unstructured tasks requiring continuous whole-body coordination.
- Beite Technology (603009): core lead-screw supplier for Tesla Optimus, with orders booked through September. A useful observation point for how the overseas leader’s production ramp transmits into the A-share supply chain.
- Industry overall: H1 global humanoid shipments of 19,100 units (+272%), China at 97%, with a duopoly taking shape.
2. vLLM & SGLang Community Tracking
Neither project shipped a release this cycle; both mainlines continue with high commit volume.
vLLM
- #52836 (bugfix, important): reverts “DSv4 eager workspace reuse” (#49236). The cause is buffer aliasing across CUDA streams causing silent correctness problems in production — note the word “silent.” This class of defect does not crash and does not log; it just occasionally returns a wrong answer. The revert restores allocator-managed buffers.
Impact: any team running DSv4 that depends on this reuse path should check its version immediately. Debugging a silent correctness defect costs far more than the memory reuse ever saved.
- #51875 (core, important): the prefix cache’s
NONE_HASHnow defaults to a fixed seed. Previously, sharing a prefix cache across nodes required manually pinning every instance’sPYTHONHASHSEEDto the same value; otherwise hashes disagreed and the cache missed forever. After this change, cross-node, object-store, and Mooncake connectors hit the shared cache out of the box.Impact: operational complexity for PD disaggregation, KV tiering, and multi-instance deployments drops a full notch. This is the kind of change that produces no percentage point but saves a lot of people a late night.
- #52512: GLM-5.2 no longer misuses dense MHA. #51368: fixes the dummy load in the DSv4 mHC broadcast buffer. #50493: Kimi-K3 gains DCP partial prefix hits.
- #52539: fused GDN MTP supports Qwen head ratios. #52681: FlashInfer upgraded to 0.6.17. Plus a batch of ROCm / MI325X quantization and kernel enablement.
SGLang
- #35371: DFlash2 (see ★ above).
- #35375 (memory): lends idle fragments of the CUDA graph pool to EAGLE’s full-vocabulary verification probabilities (flag:
SGLANG_ENABLE_GRAPH_POOL_BORROW=1), cutting the memory footprint of speculative decoding. The idea is neat: EAGLE verification needs a temporary probability buffer spanning the whole vocabulary, which happens to match the shape that idle graph-pool memory can cover. Anyone memory-constrained who wants EAGLE on should try this flag. - #35214: DSV4 mhc post-pre fusion enabled by default. #35162: adds a
deepseek_v4_flash_w8a8_8p_in32k_out1k_50mslow-latency recipe — the recipe name itself encodes its constraints (8 devices, 32K input / 1K output, 50 ms target). #35224: docs add a DSV4 low-latency PD-disaggregation recipe. - #35049: deferred KV release for aborts during transfer. #35396 / #35286: lower-bound assertions for SWA eviction in PD decode preallocation. #35220: folds DSA PD + MTP + CP layerwise tests into the B300 baseline suite — getting the most complex combination paths into routine testing is the precondition for anyone daring to run these configurations in production.
Standing topic: Step-series support (still no substantive progress)
All three Step MTP PRs remain open and unmerged: vLLM #49490 (Step-3.7-Flash MTP on MRv2, still open as of 08-12), vLLM #40070 (Step-3.5 MTP layer type, open 08-08), and SGLang #32325 (Step-3.7-Flash-NVFP4 BF16 MTP shared head, open 07-24).
The most recent push in the StepFun-ai org is Step-Realtime-CLI (08-10); the main model repos remain frozen at Step-3.7-Flash 06-01 / Step-3.5-Flash 04-03, with no new open-source model in August. Public activity has shifted toward AI terminals (STEPX / Step AOS) and commercialization coverage.
The standing conclusion is unchanged: MTP is the lifeline of Step’s performance story, and also its least stable component. The model’s selling point depends on upstream merges, and upstream investment is clearly not on that side.
3. The One-Line Takeaway
Speculative decoding’s optimization center of gravity is shifting from “how much can we save on verification” to “how good are the drafts we generate” — DFlash2 puts the structural flaw of parallel drafts that cannot see each other squarely on the table with one depthwise convolution and one top-K candidate selector, while vLLM #52836’s revert is a reminder from the other side: an optimization that saves memory by aliasing across streams is paid for in silently wrong answers. The two things actually worth doing today are — if you run DSv4, check #52836; if you share KV across instances, upgrade past #51875.
Sources: SGLang PRs #35371 / #35375 / #35214 / #35162 / #35224 / #35049 / #35396 / #35286 / #35220 / #32325; vLLM PRs #52836 / #51875 / #52512 / #51368 / #50493 / #52539 / #52681 / #49490 / #40070; StepFun-ai org push timestamps; StepFun commercialization coverage (NetEase / Lanjinger, 2026-08-15).
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。