The main thread in one line: MTP teaches the model during training to predict several tokens in one forward pass; EAGLE-3 guesses autoregressively with a specially trained Draft Head; DFlash guesses a whole block in parallel in one shot; DSpark = DFlash parallel backbone + lightweight Markov Head + Confidence Head + load-aware scheduling. All three are “self-drafting, no separate small LM,” they just differ in draft topology.
First, correcting the most easily confused point: DSpark is not MTP, and it is not simply “one draft model generating multiple tokens autoregressively.” The drafting core of DSpark is “DFlash’s block-parallel backbone + a lightweight Markov Head,” plus a Confidence Head — that is, “parallel generation + lightweight sequential dependency + confidence scheduling.” Both the official paper and the vLLM Speculators documentation describe it this way. The rest of this post takes the whole chain apart.
1. What Exactly Is an MTP Head
MTP = Multi-Token Prediction. An ordinary LM Head only produces the next token:
A B C D -> Transformer -> LM Head -> P(E | ABCD) -> E
MTP attaches several extra prediction heads / MTP modules to the original model, so one forward yields a string of future tokens:
A B C D -> Transformer
|- LM Head -> E
|- MTP-1 Head -> F
|- MTP-2 Head -> G
|- MTP-3 Head -> H
So one forward gives E F G H. The core idea is: explicitly train the model during training to predict multiple future tokens (the vLLM Speculators definition of MTP is also “finetune the model’s native multi-token prediction head”).
Note: an MTP Head is not “four completely independent Linear layers.” Modern MTP lets later prediction positions use earlier hidden/token information, and the exact MTP module structure differs across models.
2. Why MTP Can Be Used for Speculative Decoding
After MTP predicts E F G H, the target / verifier can verify the whole chain in parallel:
A B C D -> E accepted
F accepted
G accepted
H rejected
The final result accepts E F G, discards H, and lets the target correct after G. So MTP is itself a ready-made speculative decoding drafter — this is where the common DeepSeek pattern of “Target Model with MTP modules attached, verified in parallel” comes from.
3. How Is EAGLE-3 Different from MTP
This is the most important layer. EAGLE-3 is a specifically trained Draft Head / Speculator, not a few MTP Heads bolted onto the original model:
Target Model -> hidden states -> Eagle3 Draft Head -> autoregressively generate t1->t2->t3->t4 -> Target Verify
EAGLE-3’s draft tokens have clear sequential dependency (the vLLM docs describe it as “autoregressively predict draft tokens using Llama-style draft layers”). MTP, by contrast, is native to the model, a multi-token head embedded during training. Both can serve as drafters, but their origin and draft topology differ.
4. What Is DFlash Then
DFlash goes in a completely different direction — one forward pass predicts an entire block in parallel:
Eagle3 (serial): DFlash (parallel):
t1 Input |- t1
| |- t2
t2 |- t3
| |- t4
t3 |- t5
| |- t6
t4
DFlash’s advantage is that it is very fast; the cost is insufficient dependency between tokens, so the acceptance rate drops the further back you go — this is acceptance decay / suffix decay (the DSpark paper states explicitly that the main weakness of a pure parallel drafter is insufficient in-block token dependency).
5. How Does DSpark Solve It
The body of DSpark is still DFlash’s parallel generation, but with a lightweight Markov Head on top so that the k-th token is aware of the previous token:
Anchor -> DFlash Parallel Backbone -> t1 t2 t3 ... tN
|
Markov Head
|
adds local token-to-token dependency
That is, “parallel backbone + very light local sequential dependency,” which is where the paper title Semi-Autoregressive Generation comes from.
6. Why the Markov Head Helps
With pure parallelism, each t_k only depends on the context; after adding the Markov Head, t_k also depends on the previous token t_{k-1}. In the official implementation the Markov Head is a low-rank logit bias:
B = W1 @ W2 (default rank = 256)
It adds a bias to the current draft logits based on the previous token. The cost is almost zero, yet it patches the weakness of a parallel draft tail being incoherent and easily rejected.
7. Only Then Comes the Confidence Head
The Confidence Head is a lightweight head:
hidden state -> Linear -> sigmoid -> c_k
where:
c_k = P(token k is accepted | all previous tokens are accepted)
This is very important — it is not “how probable this token is by itself,” but “if everything before is correct, what is the conditional probability that this token is also correct.”
8. Is “High Confidence” the Same as “High Probability”
Yes, but it must be stated precisely: “high confidence” here means the acceptance probability predicted by the Confidence Head is high. For example c1=0.95, c2=0.90, c3=0.85, c4=0.40 corresponds to
P(t1 accepted | previous correct)
P(t2 accepted | t1 correct)
P(t3 accepted | t1,t2 correct)
P(t4 accepted | t1,t2,t3 correct)
and not to P(t1), P(t2), P(t3), P(t4), the tokens’ own generation probabilities. So the strict name is conditional acceptance probability.
9. Why It Is Called Prefix Survival Probability
What we really care about is “can this whole prefix survive.” Multiply the conditional probabilities:
Prefix 1: 0.95
Prefix 2: 0.95 x 0.90 = 0.855
Prefix 3: 0.95 x 0.90 x 0.85 = 0.727
Prefix 4: 0.95 x 0.90 x 0.85 x 0.40 = 0.291
The definition given in the DSpark paper is a_{r,j} = prod_{i=1}^{j} c_{r,i} — the prefix survival probability is the product of the conditional probabilities at each position.
10. The Complete DSpark Structure
Putting it all together, the end-to-end flow is: the target first generates the anchor, then DSpark produces E F G H plus confidence in parallel, the scheduler truncates it to a prefix, and the target verifies in parallel.
11. How DSpark Relates to EAGLE-3 / DFlash / MTP
Put the four in one picture: EAGLE-3 = shallow autoregressive drafter; DFlash = parallel drafter; MTP = native multi-token head; DSpark = DFlash + Markov Head + Confidence Head + Hardware-aware Scheduler. So the relationship between DSpark and EAGLE-3 is not “DSpark is an improved EAGLE-3,” but “a sequential head and confidence scheduling stacked on top of a parallel draft (DFlash).” The official Speculators docs explicitly define DSpark as “extends DFlash with Markov head + confidence head.”
<text x="250" y="222">- autoregressive per-position feature prediction</text>
<text x="250" y="244">- dynamic draft tree plus tree attention</text>
<text x="250" y="266">- multi-layer fusion plus Training-Time Test</text>
<text x="250" y="288">- up to 6.5x speedup</text>
<text x="250" y="310">- founder of the self-drafting school</text>
<text x="470" y="222">- D heads attached to the main model (pretrained)</text>
<text x="470" y="244">- each head predicts offset d+1</text>
<text x="470" y="266">- used as draft head at inference (zero cost)</text>
<text x="470" y="288">- reuses main model weights</text>
<text x="470" y="310">- replaced the V4 MTP-1 baseline</text>
12. Does DSpark’s Draft “Generate Multiple Tokens at Once”
Yes. But it should be seen in two layers:
- DFlash Backbone: one forward pass produces the candidate tokens of the whole block in parallel (
t1..t6). - Markov Head: uses
previous token -> logit biasto correct each position, adding inter-token dependency.
So DSpark is “parallel backbone + lightweight sequential dependency” — neither pure parallel nor pure autoregressive, which is exactly Semi-Autoregressive Generation.
13. The Easiest Comparison Table to Remember
| Method | How the draft is produced | Inter-token dependency | Tokens per forward | Main issue / advantage |
|---|---|---|---|---|
| Plain decode | Target | strong | 1 token | slow |
| EAGLE-3 | autoregressive Draft Head | strong | multiple steps | draft is serial |
| DFlash | block parallel | weak | multiple tokens | suffix acceptance decay |
| DSpark | DFlash + Markov | medium | multiple tokens | balances speed and acceptance |
| MTP | native MTP Head | depends on the MTP structure | multiple tokens | needs native model support / training |
One line to close: EAGLE-3 guesses one by one but guesses accurately; DFlash guesses a batch in one shot but the later ones drift; DSpark guesses a batch in one shot while using a Markov Head so later guesses refer to the previous token, then uses a Confidence Head to judge how far this chain is worth verifying. And the precise meaning of “high confidence = high probability” is: the probability that this position is accepted by the target conditioned on the prefix being correct; multiplying these conditional probabilities gives the prefix survival probability, and the scheduler decides the verification length from it. On DeepSeek-V4 production traffic the paper reports a per-user generation speed improvement of about 60%-85% over the production MTP-1 baseline.
This is Ep7 of the Speculative Decoding Notes series. Previously: DFlash deep dive, DSpark deep dive, DFlash vs DSpark, EAGLE-3 deep dive; back to the series index.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。