系列:Speculative Decoding Notes

DFlash 2 and the Combine Question: Block-Diffusion Upgrade, but Why Don't Frameworks Just Stack DSpark Scheduling On Top?

The one-line thread: DFlash 2 changes how the draft is produced (post-hoc path selection + physical information flow); DSpark changes how much to verify (confidence scheduling). The drafter is a mutually-exclusive config knob in the frameworks; only the scheduler is orthogonally composable — and DSpark’s scheduler depends on its own confidence-head signal, so reusing it on a DFlash drafter still needs a confidence proxy plus a benchmark. That is why it is not shipped as a named combo yet.

This post follows ep3 (DFlash) / ep4 (DSpark) / ep5 (head-to-head) / ep7 (MTP & confidence heads). Quick DFlash 2 recap, then the meta-question.

1. What DFlash 2 actually changes (recap)

DFlash 1 had two known flaws: block-internal incoherence (per-position independent Top-1 → “我 我 想”) and tail decay (block-internal attention down to 8% in shallow layers). DFlash 2 patches both:

PatchWhat it doesCost
Path selector (2M params)LM Head already emits the full vocab → keep Top-16 candidates for free; a light scorer computes “adjacency” in parallel; greedy backtracking picks one coherent path (lookup-only, 0.6% latency); rejection sampling preserved → provably lossless+2M
Local conv (16.5M / +3%)Insert two-tap dynamic depthwise conv (current + left token) around each Attn/FFN sublayer; forces left→right flow; block-internal attention 9.4%→0.5%, equivalent to deepening to 15 layers+16.5M

Measured (MindStudio, single H200, Qwen3.8-27B): acceptance length 5.46 on GSM8K vs 5.02 (built-in MTP head) and 4.36 (community DSpark drafter); 3.1–3.4× at concurrency 1, still >1.0× at concurrency 32, and ahead of MTP and the DSpark drafter at every concurrency and task.

2. The meta-question: why not just combine them?

The intuition is fair — “DFlash 2 drafter + DSpark-style confidence-scheduled verification, one owns the draft, the other owns the verify, different layers, why not stack?” Three layers of the answer:

2.1 The drafter is a mutually-exclusive config, not a stackable module

A single SD request runs one drafter. Frameworks pick it with one method field:

DFlash is block diffusion (one forward, whole block); DSpark is semi-AR + Markov head (block-parallel, intra-block serial). Different checkpoints and forward logic — you cannot be “block-diffusion” and “semi-AR” in the same forward. So “combining drafters” is physically impossible; you pick DFlash 2 or DSpark.

2.2 What is truly composable is the verify/scheduling policy — and it is already partly unified

DSpark’s confidence scheduler (verify more when confident, less when unsure) lives in the verify stage, decoupled from the drafter. In principle it is model-agnostic and can sit on top of any drafter. In practice:

So “DFlash 2 drafter + adaptive verify budget” is already the closest real shape — it just usually goes by “DFlash + adaptive K”, not “DFlash × DSpark”.

2.3 So why doesn’t anyone bolt DSpark’s scheduler onto DFlash directly? The confidence signal source

DSpark’s scheduler consumes the per-token confidence its own Markov / Confidence Head produced at draft time. Swap in a DFlash 2 drafter and that signal source changes:

To reuse DSpark’s scheduler verbatim on a DFlash drafter you first need a confidence proxy (mapping path-selector score / target accept probability into what the scheduler expects), then publish a benchmark proving the combo wins. That glue + eval is exactly what is “not yet packaged as a named combo” — not impossible, just missing a contributor who PRs it and benchmarks it.

2.4 Two more real constraints

3. Conclusion: what the “best combo” actually looks like in today’s frameworks

ComboFeasible?Status
DFlash 2 drafter + fixed-K verify✅ shippedSGLang (source) / vLLM (PR #52816)
DSpark drafter + confidence scheduler✅ shippedvLLM #46995 / SGLang #30261
DFlash 2 drafter + adaptive verify budget✅ effectively feasiblevLLM #48692 / SGLang adaptive budget, just not called a “combo”
DFlash 2 drafter + DSpark scheduler (zero-change)❌ impossibleconfidence signal source mismatch
DFlash 2 + DSpark dual drafters in parallel❌ impossibledrafters are mutually exclusive

So your line “not a replacement, an overlay” is right — with one precise footnote: the overlay is at the scheduler layer, not the drafter layer; and moving DSpark’s scheduler onto DFlash “as-is” is blocked by a missing confidence proxy plus a missing benchmark. For OpenInfer the pragmatic landing spot is the first row: DFlash 2 drafter + the frameworks’ existing adaptive verify budget, letting “how much to verify” ride on the already-shipped confidence scheduler rather than rewiring DSpark’s Markov head.

4. Thread back to earlier posts

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。