系列:Speculative Decoding Notes

Head-to-Head: Where DFlash and DSpark Actually Differ

1. Common ground: both are speculative decoding

DFlash and DSpark both build on the speculative decoding (SD) paradigm: a small draft model emits blocks in parallel, the target large model verifies in parallel, and rejection sampling preserves the distribution (zero quality loss). Neither is a “new model” — both are decoding / serving-layer acceleration schemes.

Below we pull them apart dimension by dimension.

2. Dimension-by-dimension

DimensionDFlashDSpark
OriginZ Lab + SGLang + Modal (2026-06-15 blog)DeepSeek + Peking University (2026-06-27, open source, MIT, deepseek-ai/DeepSpec)
Draft paradigmBlock Diffusion: one forward pass predicts a whole block of masked future tokens in parallelSemi-Autoregressive + Markov head: block-level parallel emission, Markov head injects intra-block sequential dependency
Conditioning / orderingTarget hidden-state conditioning + KV injection (injects target features across layers, keeps acceptance high)Markov head injects intra-block sequential dependency
Verify schedulingFixed block size (default 16), standard parallel verificationConfidence scheduler dynamically adjusts verify length (verify more when confident, less when not)
Execution optimizationSpec V2 engine + overlap scheduler: eliminates host-device sync idle (+33% more)Algorithm-layer focus; execution relies on the host engine (vLLM / SGLang)
Published resultsQwen3-8B up to 6× (≈2.5× faster than EAGLE-3); 15× on BlackwellV4 speedup 57–85%, throughput +400%
EcosystemSGLang primary; vLLM PR #16818 in progress; multiple Qwen3 drafter tiersvLLM PR #46995 merging; also supports SGLang / OpenInfer

3. The one-line difference

4. Why they are orthogonal and stackable

DFlash changes “how the draft is produced” and “how it is executed”; DSpark changes “how intra-block order is built” and “how many tokens to verify”. They operate at different layers, hence orthogonal: OpenInfer can mount both paths on Qwen3-4B simultaneously — one side is DFlash’s block-diffusion drafter + KV injection + overlap scheduler, the other is DSpark’s semi-AR drafter + Markov head + confidence scheduler.

5. A common misconception

DFlash and DSpark are orthogonal and stackable, but they do not “run two drafters in parallel at the same time”. The more likely picture is: a single shared verify pipeline, with only the draft stage swapped between the two schemes. Definitive wording should follow the OpenInfer code / official blog.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。