系列:Speculative Decoding Notes

EAGLE-3 Deep Dive: Drop the Feature Constraint, Let the Drafter Finally Scale with Data

1. Why a dedicated look at EAGLE-3

Ep1–Ep5 covered the shape of speculative decoding (SD), the ceiling of autoregressive drafters, and the two new routes DFlash / DSpark. But one name keeps showing up — EAGLE — the originator of the whole “feature-level drafting” school. DFlash and DSpark both use it as the baseline (“about 2.5× faster than EAGLE-3”).

The EAGLE lineage has three generations:

What makes EAGLE-3 remarkable is not “a bit faster” — it proved that a draft model can keep benefiting from more data, exactly like a large model. That is something the speculative-decoding field had never seen before.

2. Recap: how EAGLE guesses better

Vanilla SD uses a small LLM as the drafter; the small model’s distribution diverges from the large model’s, so acceptance rate α is low (~1.6–1.9×).

EAGLE-1’s insight: instead of a separate small model guessing tokens, reuse the target model’s intermediate features. It autoregresses at the feature level — concatenating the target’s top-layer feature with a “one-step-ahead token” and feeding it to a 1-layer draft Decoder that predicts the next feature, then uses the target’s LM head to turn it into a token. This reached ~3–4×.

EAGLE-2 found that the drafter’s confidence is highly calibrated with the actual acceptance rate, so it dynamically shapes the draft tree (fewer branches for easy positions, more for hard ones), pushing speedup to ~4.2×.

3. The feature-prediction constraint: why it is a scaling lock

EAGLE-1/2’s drafter is trained under two losses:

When the authors tried scaling training data 8×, speedup barely moved. The culprit is the feature loss:

The drafter has only 1 Decoder layer and tiny capacity. The feature loss forces it to spend capacity on “geometric fitting” — placing the output vector close to the target’s feature in space — instead of on the real goal: predicting the token as accurately as possible.

In other words, feature prediction is an over-strong regularizer: it helps generalization when data is scarce, but becomes a lock when data is abundant. Worse, the feature constraint also locks the input to “top-layer features only” (since top-layer features and next-token logits are one-to-one).

Figure 1: Draft training data volume vs speedup (LLaMA-Instruct 3.1 8B, MT-bench, schematic) 3.5 4.5 5.5 6.5 1× 2× 4× 8× Training data volume (× relative to ShareGPT) EAGLE / EAGLE-2 (flat) EAGLE-3 (rises w/ data) After dropping feature prediction, the drafter shows its first scaling law: more data, higher speedup, up to 6.5×.

4. EAGLE-3’s two changes

Two designs that work together — neither alone is enough.

4.1 Drop feature prediction, predict tokens directly via Training-Time Test

Core insight: feature prediction is a means, not an end — its only job was letting single-step training generalize to multi-step inference, at the cost of over-constraining the output. EAGLE-3 makes the drafter predict tokens directly, and simulates the multi-step generation process during training (wiring the target’s LM head and sampling into the drafter’s training loop). This is called Training-Time Test: the drafter practices, at training time, exactly the multi-step generation it will do at test time, eliminating the train–test distribution gap that plagued EAGLE.

Figure 2: EAGLE-2 (feature-constrained) vs EAGLE-3 (Training-Time Test) EAGLE-2 EAGLE-3 Input: feature f_t + ahead token t₍ₜ₊₁₎ Draft model (1-layer Decoder) Predict f₍ₜ₊₁₎ (SmoothL1 fits target feature) LM head → token Single-step train→multi-step; feature lock caps capacity & input Input: multi-layer fused feature + prev token Draft model (capacity scalable) Predict token directly (CrossEntropy loss) Training simulates multi-step generation: wire LM head + sampling into training loop Train = "rehearse" test-time multi-step, kills distribution gap

4.2 Multi-layer feature fusion

With the feature constraint gone, the input is free too — instead of top-layer features only, EAGLE-3 fuses low / mid / high-level features from the target model, giving the drafter richer semantic context. This and “direct token prediction” are mutually prerequisite: without the feature constraint, you would not dare to change the input.

5. Three-generation comparison

DimensionEAGLE-1EAGLE-2EAGLE-3
Draft stylefeature-level ARdynamic draft treedirect token + multi-layer fusion
Feature constraintyesyesno
Inputtop-layer featuretop-layer featurelow/mid/high fusion
Scaling lawnonoyes
Typical speedup~3–4×~4.2×up to 6.5×

6. Results

MetricValue
Max speedup6.5× (Vicuna 13B, HumanEval, T=0)
vs EAGLE-2~1.4× lower latency
SGLang throughput (batch=64, H100)+1.38×
Compatibilityfully compatible with EAGLE-2’s draft tree

7. Takeaway

The EAGLE lineage in one line:

EAGLE-1 solved “how to borrow the large model’s info to guess better” → EAGLE-2 solved “how to spend the compute budget smarter” → EAGLE-3 solved “how to keep benefiting from more data.”

It sends a signal to the whole SD field: a drafter need not be a tiny toy — drop the wrong constraint, give it richer input and more capacity, and it can scale too. That is why DFlash / DSpark both benchmark against EAGLE-3: it defines the ceiling of the “feature-level drafting” route.

This post belongs to the “Speculative Decoding Notes” series, Ep6. See also: DFlash Deep Dive, DSpark Deep Dive, DFlash vs DSpark; back to the series index at Speculative Decoding Notes.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。