系列:Speculative Decoding Notes

Speculative Decoding: Lossless Multi-Token Generation

1. Motivation: the serial bottleneck of autoregressive generation

Today’s large language models (LLMs) are almost all autoregressive: to produce each new token, the model feeds the entire context — including that new token — back through a full forward pass. In other words, generation is strictly serial — token t+1 cannot start until token t finishes.

This implies two things:

Could one large-model forward pass emit several tokens at once? That is exactly what speculative decoding (SD) addresses.

2. Core idea: Draft + Verify, two stages

SD introduces a draft (small) model: much smaller and cheaper than the target large model, but trained for roughly the same task. Generation is no longer “the large model walks step by step”; it is a two-stage collaboration:

  1. Draft: the small model parallelizes a whole block of candidate tokens at once (say, 4 tokens).
  2. Verify: the candidate block is handed to the large model for a single parallel forward pass; the large model independently decides “accept / reject” for each candidate (via rejection sampling).

Accepted candidates are kept; at the first rejected position, the large model overrides with its own prediction and continues from there. The key point: no matter how long the draft, the large model ran only one forward pass — so if most candidates are accepted, we traded one forward pass for multiple tokens.

Intuition: let the cheap small model “guess” a sequence first; the expensive large model just “glances once” and stamps it. The more accurate the guess, the bigger the speedup.

Draft (small)
Target (large)
Click "Start" to watch the Draft → Verify loop

3. Why lossless? Rejection sampling

This is SD’s most appealing property: the output distribution matches the target large model token-by-token, with zero quality loss.

The secret is rejection sampling. During verification the large model does not simply “accept all or reject all”; it re-samples per token by probability:

So SD does not change what the model says, only how fast it says it. This is crucial for production: you can speed up with SD without worrying about answer quality dropping.

4. Where does the speedup come from? Acceptance length α

Let one large-model forward pass cost ≈ 1 unit, and drafting k tokens cost ≈ β·k (β≪1). If on average α candidates are accepted (acceptance length α), the total cost to produce α tokens ≈ 1 + β·k.

Speedup ≈ α / (1 + β·k).

So:

5. A minimal intuition example

Suppose we draft 4 tokens; the large model accepts the first 3 in one verification and rejects the 4th, overriding it with its own prediction:

In that step we compressed “4 passes” into “1” — that is where SD’s speedup comes from.

6. Summary and what’s next

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。