系列:Speculative Decoding Notes

DFlash Deep Dive: Block-Diffusion Drafting, KV Injection, and a Bubble-Free Pipeline

1. Background and origin

DFlash was systematically introduced by Z Lab + SGLang + Modal in a blog post on 2026-06-15, with the goal of pushing speculative decoding on the Qwen3 family up to 6×. It breaks the “serial drafting” ceiling from Ep2 because it reworks both the drafting paradigm and the execution engine.

2. Core 1: Block Diffusion Drafter

Classic drafters are autoregressive and serial — they guess one token at a time. DFlash switches to a diffusion paradigm:

This single move sidesteps Ep2’s bottlenecks (a) serial drafting and (c) in-block sequential dependency — drafting goes from “serial token chaining” to “parallel block emission”.

3. Core 2: Target hidden-state conditioning + KV injection

Parallelism alone is not enough; the acceptance length α is what drives speedup. DFlash’s key trick is to inject the target model’s intermediate hidden states into the drafter’s KV projections (across layers):

This is where DFlash’s high α comes from, and a major reason it is ~2.5× faster than EAGLE-3.

4. Core 3: Spec V2 engine + overlap scheduler (execution layer)

Most SD work only changes the algorithm; DFlash also reworks the execution engine:

This is what separates DFlash from “pure-algorithm” approaches: it changes both the drafting paradigm and the execution engine.

5. Measured results

MetricValue
Qwen3-8B peak speedup6×
vs EAGLE-3~2.5× faster
On Blackwellup to 15×
overlap scheduler gain+33% throughput

6. Summary and next

DFlash’s triple move — block-diffusion parallel drafting + KV injection for higher acceptance + bubble-free pipeline — lifts SD from 2–3× to 6×. Next we look at the alternative route DSpark: instead of diffusion, it uses semi-autoregressive drafting + a Markov head + a confidence scheduler to take a different path.

This is Ep3 of the “Speculative Decoding Notes” series. Prior: The Autoregressive Drafter’s Ceiling; head-to-head: DFlash vs DSpark; next: DSpark Deep Dive.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。