1. Why a dedicated look at EAGLE-3
Ep1–Ep5 covered the shape of speculative decoding (SD), the ceiling of autoregressive drafters, and the two new routes DFlash / DSpark. But one name keeps showing up — EAGLE — the originator of the whole “feature-level drafting” school. DFlash and DSpark both use it as the baseline (“about 2.5× faster than EAGLE-3”).
The EAGLE lineage has three generations:
- EAGLE-1 (ICML 2024): autoregresses at the feature level, borrowing the target model’s top-layer features to guess more accurately;
- EAGLE-2 (EMNLP 2024): dynamically schedules the draft tree by confidence, allocating compute more cleverly;
- EAGLE-3 (2025): drops the feature-prediction constraint, giving the draft model its first scaling law.
What makes EAGLE-3 remarkable is not “a bit faster” — it proved that a draft model can keep benefiting from more data, exactly like a large model. That is something the speculative-decoding field had never seen before.
2. Recap: how EAGLE guesses better
Vanilla SD uses a small LLM as the drafter; the small model’s distribution diverges from the large model’s, so acceptance rate α is low (~1.6–1.9×).
EAGLE-1’s insight: instead of a separate small model guessing tokens, reuse the target model’s intermediate features. It autoregresses at the feature level — concatenating the target’s top-layer feature with a “one-step-ahead token” and feeding it to a 1-layer draft Decoder that predicts the next feature, then uses the target’s LM head to turn it into a token. This reached ~3–4×.
EAGLE-2 found that the drafter’s confidence is highly calibrated with the actual acceptance rate, so it dynamically shapes the draft tree (fewer branches for easy positions, more for hard ones), pushing speedup to ~4.2×.
3. The feature-prediction constraint: why it is a scaling lock
EAGLE-1/2’s drafter is trained under two losses:
- feature loss (SmoothL1): make the draft output approximate the target model’s top-layer feature;
- token loss (CrossEntropy): get the token right.
When the authors tried scaling training data 8×, speedup barely moved. The culprit is the feature loss:
The drafter has only 1 Decoder layer and tiny capacity. The feature loss forces it to spend capacity on “geometric fitting” — placing the output vector close to the target’s feature in space — instead of on the real goal: predicting the token as accurately as possible.
In other words, feature prediction is an over-strong regularizer: it helps generalization when data is scarce, but becomes a lock when data is abundant. Worse, the feature constraint also locks the input to “top-layer features only” (since top-layer features and next-token logits are one-to-one).
4. EAGLE-3’s two changes
Two designs that work together — neither alone is enough.
4.1 Drop feature prediction, predict tokens directly via Training-Time Test
Core insight: feature prediction is a means, not an end — its only job was letting single-step training generalize to multi-step inference, at the cost of over-constraining the output. EAGLE-3 makes the drafter predict tokens directly, and simulates the multi-step generation process during training (wiring the target’s LM head and sampling into the drafter’s training loop). This is called Training-Time Test: the drafter practices, at training time, exactly the multi-step generation it will do at test time, eliminating the train–test distribution gap that plagued EAGLE.
4.2 Multi-layer feature fusion
With the feature constraint gone, the input is free too — instead of top-layer features only, EAGLE-3 fuses low / mid / high-level features from the target model, giving the drafter richer semantic context. This and “direct token prediction” are mutually prerequisite: without the feature constraint, you would not dare to change the input.
5. Three-generation comparison
| Dimension | EAGLE-1 | EAGLE-2 | EAGLE-3 |
|---|---|---|---|
| Draft style | feature-level AR | dynamic draft tree | direct token + multi-layer fusion |
| Feature constraint | yes | yes | no |
| Input | top-layer feature | top-layer feature | low/mid/high fusion |
| Scaling law | no | no | yes |
| Typical speedup | ~3–4× | ~4.2× | up to 6.5× |
6. Results
| Metric | Value |
|---|---|
| Max speedup | 6.5× (Vicuna 13B, HumanEval, T=0) |
| vs EAGLE-2 | ~1.4× lower latency |
| SGLang throughput (batch=64, H100) | +1.38× |
| Compatibility | fully compatible with EAGLE-2’s draft tree |
7. Takeaway
The EAGLE lineage in one line:
EAGLE-1 solved “how to borrow the large model’s info to guess better” → EAGLE-2 solved “how to spend the compute budget smarter” → EAGLE-3 solved “how to keep benefiting from more data.”
It sends a signal to the whole SD field: a drafter need not be a tiny toy — drop the wrong constraint, give it richer input and more capacity, and it can scale too. That is why DFlash / DSpark both benchmark against EAGLE-3: it defines the ceiling of the “feature-level drafting” route.
This post belongs to the “Speculative Decoding Notes” series, Ep6. See also: DFlash Deep Dive, DSpark Deep Dive, DFlash vs DSpark; back to the series index at Speculative Decoding Notes.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。