1. Common ground: both are speculative decoding
DFlash and DSpark both build on the speculative decoding (SD) paradigm: a small draft model emits blocks in parallel, the target large model verifies in parallel, and rejection sampling preserves the distribution (zero quality loss). Neither is a “new model” — both are decoding / serving-layer acceleration schemes.
Below we pull them apart dimension by dimension.
2. Dimension-by-dimension
| Dimension | DFlash | DSpark |
|---|---|---|
| Origin | Z Lab + SGLang + Modal (2026-06-15 blog) | DeepSeek + Peking University (2026-06-27, open source, MIT, deepseek-ai/DeepSpec) |
| Draft paradigm | Block Diffusion: one forward pass predicts a whole block of masked future tokens in parallel | Semi-Autoregressive + Markov head: block-level parallel emission, Markov head injects intra-block sequential dependency |
| Conditioning / ordering | Target hidden-state conditioning + KV injection (injects target features across layers, keeps acceptance high) | Markov head injects intra-block sequential dependency |
| Verify scheduling | Fixed block size (default 16), standard parallel verification | Confidence scheduler dynamically adjusts verify length (verify more when confident, less when not) |
| Execution optimization | Spec V2 engine + overlap scheduler: eliminates host-device sync idle (+33% more) | Algorithm-layer focus; execution relies on the host engine (vLLM / SGLang) |
| Published results | Qwen3-8B up to 6× (≈2.5× faster than EAGLE-3); 15× on Blackwell | V4 speedup 57–85%, throughput +400% |
| Ecosystem | SGLang primary; vLLM PR #16818 in progress; multiple Qwen3 drafter tiers | vLLM PR #46995 merging; also supports SGLang / OpenInfer |
3. The one-line difference
- DFlash reinvents how the draft is generated (autoregressive → block diffusion) + the execution engine (overlap scheduler that removes host-device sync idle).
- DSpark reinvents intra-block ordering (semi-autoregressive + Markov head) + dynamic verification scheduling (confidence scheduler).
4. Why they are orthogonal and stackable
DFlash changes “how the draft is produced” and “how it is executed”; DSpark changes “how intra-block order is built” and “how many tokens to verify”. They operate at different layers, hence orthogonal: OpenInfer can mount both paths on Qwen3-4B simultaneously — one side is DFlash’s block-diffusion drafter + KV injection + overlap scheduler, the other is DSpark’s semi-AR drafter + Markov head + confidence scheduler.
5. A common misconception
DFlash and DSpark are orthogonal and stackable, but they do not “run two drafters in parallel at the same time”. The more likely picture is: a single shared verify pipeline, with only the draft stage swapped between the two schemes. Definitive wording should follow the OpenInfer code / official blog.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。