系列:Speculative Decoding Notes

MTP Heads and Confidence Heads: Linking Three Self-Drafting Routes into One Chain

The main thread in one line: MTP teaches the model during training to predict several tokens in one forward pass; EAGLE-3 guesses autoregressively with a specially trained Draft Head; DFlash guesses a whole block in parallel in one shot; DSpark = DFlash parallel backbone + lightweight Markov Head + Confidence Head + load-aware scheduling. All three are “self-drafting, no separate small LM,” they just differ in draft topology.

First, correcting the most easily confused point: DSpark is not MTP, and it is not simply “one draft model generating multiple tokens autoregressively.” The drafting core of DSpark is “DFlash’s block-parallel backbone + a lightweight Markov Head,” plus a Confidence Head — that is, “parallel generation + lightweight sequential dependency + confidence scheduling.” Both the official paper and the vLLM Speculators documentation describe it this way. The rest of this post takes the whole chain apart.

1. What Exactly Is an MTP Head

MTP = Multi-Token Prediction. An ordinary LM Head only produces the next token:

A B C D -> Transformer -> LM Head -> P(E | ABCD) -> E

MTP attaches several extra prediction heads / MTP modules to the original model, so one forward yields a string of future tokens:

A B C D -> Transformer
              |- LM Head      -> E
              |- MTP-1 Head   -> F
              |- MTP-2 Head   -> G
              |- MTP-3 Head   -> H

So one forward gives E F G H. The core idea is: explicitly train the model during training to predict multiple future tokens (the vLLM Speculators definition of MTP is also “finetune the model’s native multi-token prediction head”).

MTP Head structure (DeepSeek-V3 multi-token prediction module) Main Transformer hidden state plus the token embedding shifted one step forward, passed through an MTP Transformer Block, RMSNorm, and a shared LM Head to predict one extra future token; D modules can be stacked for multi-token supervision. MTP Head structure (DeepSeek-V3 multi-token prediction module) Main Transformer shared vocab / LM Head h_t hidden at pos t Emb(t+1) token shifted by one MTP module (d=1) MTP Transformer Block RMSNorm LM Head (shared) predict t+2

D MTP modules can be stacked: MTP1 to t+2, MTP2 to t+3, …, MTP-D to t+D+1 (multi-token supervision in training) At inference: MTP1 reuses the main model hidden state and predicts 1 extra token as a speculative draft head (near zero cost)

Note: an MTP Head is not “four completely independent Linear layers.” Modern MTP lets later prediction positions use earlier hidden/token information, and the exact MTP module structure differs across models.

2. Why MTP Can Be Used for Speculative Decoding

After MTP predicts E F G H, the target / verifier can verify the whole chain in parallel:

A B C D -> E accepted
             F accepted
             G accepted
             H rejected

The final result accepts E F G, discards H, and lets the target correct after G. So MTP is itself a ready-made speculative decoding drafter — this is where the common DeepSeek pattern of “Target Model with MTP modules attached, verified in parallel” comes from.

3. How Is EAGLE-3 Different from MTP

This is the most important layer. EAGLE-3 is a specifically trained Draft Head / Speculator, not a few MTP Heads bolted onto the original model:

Target Model -> hidden states -> Eagle3 Draft Head -> autoregressively generate t1->t2->t3->t4 -> Target Verify

EAGLE-3’s draft tokens have clear sequential dependency (the vLLM docs describe it as “autoregressively predict draft tokens using Llama-style draft layers”). MTP, by contrast, is native to the model, a multi-token head embedded during training. Both can serve as drafters, but their origin and draft topology differ.

4. What Is DFlash Then

DFlash goes in a completely different direction — one forward pass predicts an entire block in parallel:

Eagle3 (serial):   DFlash (parallel):
t1                Input |- t1
 |                      |- t2
t2                      |- t3
 |                      |- t4
t3                      |- t5
 |                      |- t6
t4

DFlash’s advantage is that it is very fast; the cost is insufficient dependency between tokens, so the acceptance rate drops the further back you go — this is acceptance decay / suffix decay (the DSpark paper states explicitly that the main weakness of a pure parallel drafter is insufficient in-block token dependency).

5. How Does DSpark Solve It

The body of DSpark is still DFlash’s parallel generation, but with a lightweight Markov Head on top so that the k-th token is aware of the previous token:

Anchor -> DFlash Parallel Backbone -> t1 t2 t3 ... tN
                                        |
                                   Markov Head
                                        |
                            adds local token-to-token dependency

That is, “parallel backbone + very light local sequential dependency,” which is where the paper title Semi-Autoregressive Generation comes from.

6. Why the Markov Head Helps

With pure parallelism, each t_k only depends on the context; after adding the Markov Head, t_k also depends on the previous token t_{k-1}. In the official implementation the Markov Head is a low-rank logit bias:

B = W1 @ W2      (default rank = 256)

It adds a bias to the current draft logits based on the previous token. The cost is almost zero, yet it patches the weakness of a parallel draft tail being incoherent and easily rejected.

7. Only Then Comes the Confidence Head

The Confidence Head is a lightweight head:

hidden state -> Linear -> sigmoid -> c_k

where:

c_k = P(token k is accepted | all previous tokens are accepted)

This is very important — it is not “how probable this token is by itself,” but “if everything before is correct, what is the conditional probability that this token is also correct.”

8. Is “High Confidence” the Same as “High Probability”

Yes, but it must be stated precisely: “high confidence” here means the acceptance probability predicted by the Confidence Head is high. For example c1=0.95, c2=0.90, c3=0.85, c4=0.40 corresponds to

P(t1 accepted | previous correct)
P(t2 accepted | t1 correct)
P(t3 accepted | t1,t2 correct)
P(t4 accepted | t1,t2,t3 correct)

and not to P(t1), P(t2), P(t3), P(t4), the tokens’ own generation probabilities. So the strict name is conditional acceptance probability.

9. Why It Is Called Prefix Survival Probability

What we really care about is “can this whole prefix survive.” Multiply the conditional probabilities:

Prefix 1: 0.95
Prefix 2: 0.95 x 0.90 = 0.855
Prefix 3: 0.95 x 0.90 x 0.85 = 0.727
Prefix 4: 0.95 x 0.90 x 0.85 x 0.40 = 0.291

The definition given in the DSpark paper is a_{r,j} = prod_{i=1}^{j} c_{r,i} — the prefix survival probability is the product of the conditional probabilities at each position.

Confidence Head predicts prefix survival probability The draft backbone produces gamma hidden states in one forward pass; each position goes through a Confidence Head to output a conditional acceptance probability c_k; the prefix survival probability a_j is the product of c_1 to c_j, monotonically decreasing; the scheduler greedily truncates the tail in descending order of a_j. Confidence Head to Prefix Survival Probability (core of DSpark verification scheduling) i.e. "Confidence Head predicts prefix survival probability" Anchor D Draft Backbone 1 forward, gamma positions h_1 h_2 h_3 h_4 h_5 Conf Head Conf Head Conf Head Conf Head Conf Head c_1 c_2 c_3 c_4 c_5

c_k = P(position k accepted | first k-1 all accepted) - conditional acceptance probability, supervised by the TV distance between draft and target distributions

a_1 a_2 a_3 a_4 a_5

prefix survival prob: a_j = c_1 c_2 … c_j (probability the first j are all accepted, monotonically decreasing in j) Scheduler: greedily add candidates in descending a_j, stop early when expected accepted tokens saturate, truncating the low-confidence tail

10. The Complete DSpark Structure

Putting it all together, the end-to-end flow is: the target first generates the anchor, then DSpark produces E F G H plus confidence in parallel, the scheduler truncates it to a prefix, and the target verifies in parallel.

DSpark end-to-end pipeline The Target generates the Anchor, the DFlash parallel backbone emits a block in one pass, the Markov Head adds local dependency, the Confidence Head outputs c_k, the scheduler decides the verification length by prefix survival probability, and control returns to the Target for parallel verification. DSpark end-to-end pipeline Target Model (generates Anchor) Anchor D DFlash Parallel Backbone 1 forward, block parallel t_1 t_2 t_3 t_4 t_N Markov Head prev token to logit bias (B=W1W2, r=256) Confidence Head to c_k (sigmoid) Hardware-aware Scheduler (truncate by a_j) back to Target, parallel verify accept / reject

11. How DSpark Relates to EAGLE-3 / DFlash / MTP

Put the four in one picture: EAGLE-3 = shallow autoregressive drafter; DFlash = parallel drafter; MTP = native multi-token head; DSpark = DFlash + Markov Head + Confidence Head + Hardware-aware Scheduler. So the relationship between DSpark and EAGLE-3 is not “DSpark is an improved EAGLE-3,” but “a sequential head and confidence scheduling stacked on top of a parallel draft (DFlash).” The official Speculators docs explicitly define DSpark as “extends DFlash with Markov head + confidence head.”

DSpark, EAGLE-3 tree, and MTP Head: three self-drafting paradigms compared Left: DSpark parallel block plus sequential head produces a flat block; middle: EAGLE-3 feature-level autoregression produces a candidate tree; right: MTP attaches multiple heads to the main model to predict offset tokens. All three are self-drafting with no separate small LM. DSpark / EAGLE-3 tree / MTP Head: three self-drafting paradigms compared DSpark EAGLE-3 (tree) MTP Head E F G H gamma flat tokens (block) root a b a candidate tree (tree attention verify) head1 head2 headD D offset heads, flat output - gamma positions in one forward - flat block (not a tree) - independently trained light drafter - shares embedding/LM head - includes confidence-head scheduling
<text x="250" y="222">- autoregressive per-position feature prediction</text>
<text x="250" y="244">- dynamic draft tree plus tree attention</text>
<text x="250" y="266">- multi-layer fusion plus Training-Time Test</text>
<text x="250" y="288">- up to 6.5x speedup</text>
<text x="250" y="310">- founder of the self-drafting school</text>

<text x="470" y="222">- D heads attached to the main model (pretrained)</text>
<text x="470" y="244">- each head predicts offset d+1</text>
<text x="470" y="266">- used as draft head at inference (zero cost)</text>
<text x="470" y="288">- reuses main model weights</text>
<text x="470" y="310">- replaced the V4 MTP-1 baseline</text>
All three are "self-drafting, no separate small LM." DSpark adds a sequential head plus confidence scheduling on a parallel draft (DFlash), improving accepted length by 26-31% over EAGLE-3 and 16-18% over DFlash, and replaces the V4 MTP-1 baseline. EAGLE-3 is a tree, MTP is offset multi-head, DSpark is a flat block: different draft topologies.

12. Does DSpark’s Draft “Generate Multiple Tokens at Once”

Yes. But it should be seen in two layers:

So DSpark is “parallel backbone + lightweight sequential dependency” — neither pure parallel nor pure autoregressive, which is exactly Semi-Autoregressive Generation.

13. The Easiest Comparison Table to Remember

MethodHow the draft is producedInter-token dependencyTokens per forwardMain issue / advantage
Plain decodeTargetstrong1 tokenslow
EAGLE-3autoregressive Draft Headstrongmultiple stepsdraft is serial
DFlashblock parallelweakmultiple tokenssuffix acceptance decay
DSparkDFlash + Markovmediummultiple tokensbalances speed and acceptance
MTPnative MTP Headdepends on the MTP structuremultiple tokensneeds native model support / training

One line to close: EAGLE-3 guesses one by one but guesses accurately; DFlash guesses a batch in one shot but the later ones drift; DSpark guesses a batch in one shot while using a Markov Head so later guesses refer to the previous token, then uses a Confidence Head to judge how far this chain is worth verifying. And the precise meaning of “high confidence = high probability” is: the probability that this position is accepted by the target conditioned on the prefix being correct; multiplying these conditional probabilities gives the prefix survival probability, and the scheduler decides the verification length from it. On DeepSeek-V4 production traffic the paper reports a per-user generation speed improvement of about 60%-85% over the production MTP-1 baseline.

This is Ep7 of the Speculative Decoding Notes series. Previously: DFlash deep dive, DSpark deep dive, DFlash vs DSpark, EAGLE-3 deep dive; back to the series index.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。