系列:每日AI热点

Daily AI Hotspot · 2026-08-19: DFlash2 Gives the Draft Head a Local Conv and a Candidate Selector — Speculative Decoding Moves from Acceptance Rate to Draft Quality

Both sides of the fence are chasing the same theme today — making speculative decoding finer-grained — but from opposite directions. SGLang’s DFlash2 improves draft quality (draft tokens stop being blind to one another), while vLLM defuses a production landmine (#52836 reverts a workspace-reuse optimization that could silently produce wrong results). On the research side there is DeepMind’s Genie, which turns a single image into a playable environment with actions discovered entirely through self-supervision. On the industry side, Unitree lists this week: the A-share market gets its first humanoid robotics stock.

★ Most Worth Your Attention Today

SGLang #35371: DFlash2 — local grouped depthwise convolution plus a candidate selector (merged 08-19, +929 / −61).

Start with the problem. Draft heads in the DFlash family carry an unavoidable constraint: positions inside a draft block are generated in parallel, so position i cannot see the draft token that position i−1 just produced. When you guess the 5th token, you do not yet know what you guessed for the 4th. For an autoregressive model this is a serious structural handicap — the longer the draft, the more the tail is pure blind guessing.

DFlash2 patches the hole with two complementary changes.

Change one: grouped dynamic depthwise convolution. A depthwise convolution along the position dimension is added inside the draft block, letting each proposal position look directly at representations from earlier positions:

out[i, c] = Σ_t (base[t, c] + δ[i, t, g(c)]) · x[i−t, c]

base is a learnable base kernel; δ is a position-dependent dynamic increment shared within channel groups g(c); taps at block boundaries are zeroed so nothing leaks across blocks. The benefit is direct: a proposal position gains access to prior context without paying for another backbone pass. The cost is essentially one cheap convolution, and the return is a real improvement in draft quality.

Change two: the candidate selector. The conventional approach takes an argmax at every slot, pinning the draft path to a single sequence. DFlash2 instead keeps top-K at every slot and then scores transitions with

edge(p→c) = ⟨A[p] ⊙ project(h), B[c]⟩ + unary[c]

walking an optimal path over the candidate graph from the verified anchor. When sampling at T>0 it returns K candidates via the inverse CDF for lossless verification — the crucial detail, because it means raising the sampling temperature no longer costs you verification correctness.

Both changes are folded into the draft CUDA graph at decode time, so no extra launch overhead is introduced.

Why this matters more than a percentage: two years of speculative-decoding optimization have gone almost entirely into the verification side — verify less, stop early, hide the syncs. DFlash2 is a rare case of attention returning to the draft side. If draft quality itself does not improve, every verification-side optimization is just refilling a leaking bucket. The two ideas here are also portable: “parallel-generated drafts can’t see each other” is a problem in every blockwise draft head, not something specific to SGLang.

1. AI Industry & Paper Highlights

Paper: Genie (DeepMind, 11B parameters, arXiv:2402.15391)

Operator: MQA (Multi-Query Attention, Shazeer 2019)

Performance: pruning

Industry roundup

2. vLLM & SGLang Community Tracking

Neither project shipped a release this cycle; both mainlines continue with high commit volume.

vLLM

SGLang

Standing topic: Step-series support (still no substantive progress)

All three Step MTP PRs remain open and unmerged: vLLM #49490 (Step-3.7-Flash MTP on MRv2, still open as of 08-12), vLLM #40070 (Step-3.5 MTP layer type, open 08-08), and SGLang #32325 (Step-3.7-Flash-NVFP4 BF16 MTP shared head, open 07-24).

The most recent push in the StepFun-ai org is Step-Realtime-CLI (08-10); the main model repos remain frozen at Step-3.7-Flash 06-01 / Step-3.5-Flash 04-03, with no new open-source model in August. Public activity has shifted toward AI terminals (STEPX / Step AOS) and commercialization coverage.

The standing conclusion is unchanged: MTP is the lifeline of Step’s performance story, and also its least stable component. The model’s selling point depends on upstream merges, and upstream investment is clearly not on that side.

3. The One-Line Takeaway

Speculative decoding’s optimization center of gravity is shifting from “how much can we save on verification” to “how good are the drafts we generate” — DFlash2 puts the structural flaw of parallel drafts that cannot see each other squarely on the table with one depthwise convolution and one top-K candidate selector, while vLLM #52836’s revert is a reminder from the other side: an optimization that saves memory by aliasing across streams is paid for in silently wrong answers. The two things actually worth doing today are — if you run DSv4, check #52836; if you share KV across instances, upgrade past #51875.


Sources: SGLang PRs #35371 / #35375 / #35214 / #35162 / #35224 / #35049 / #35396 / #35286 / #35220 / #32325; vLLM PRs #52836 / #51875 / #52512 / #51368 / #50493 / #52539 / #52681 / #49490 / #40070; StepFun-ai org push timestamps; StepFun commercialization coverage (NetEase / Lanjinger, 2026-08-15).

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。