系列:Inference Systems Infrastructure Notes

Inference Systems Infrastructure Notes (4): Knowledge Distillation — Passing the Dark Knowledge of a Large Model to a Small One

0. One-Line Positioning

Knowledge distillation (KD) = let a large model (the teacher) pass its “dark knowledge” to a small model (the student), so the student approaches the teacher’s performance with far fewer parameters.

The problem it solves: model capability is not the same as deployment cost. Large models are strong but expensive to serve; small models are cheap but, without enough training data, never learn that well. Distillation lets the small model “stand on the giant’s shoulders.”

1. Why Call It “Dark Knowledge”?

Take an image classification task where the teacher outputs:

cat:   0.70
dog:   0.20
table: 0.05
chair: 0.05

A hard label only tells the model “this is a cat.” The teacher’s soft label additionally implies “this looks like a cat but also somewhat like a dog, and not much like furniture.” Those inter-class similarity signals are the dark knowledge. A student that learns them generalizes better than one trained on one-hot labels alone.

2. Classic Hinton Distillation

Knowledge distillation: transferring dark knowledge Teacher large model emits soft labels Temperature T softens the distribution Student small model learns soft + hard Loss a * KL + (1-a) * CE p_T = softmax(z_T / T) p_S = softmax(z_S / T) L_distill = T^2 * KL(p_T || p_S) L_total = a * L_distill + (1-a) * CE(y, p_S_hard)

Figure: distillation softens the teacher's soft target with temperature T and uses it as the student's training objective.

3. What Temperature T Does

T = 1:      close to the raw distribution, little dark knowledge
T = 4 to 10: smoother distribution, inter-class relations clearer, dark knowledge rich
T -> inf:   the distribution approaches uniform and loses discriminative power

Empirical values: T = 4 to 10, with a distillation weight of alpha = 0.5 to 0.9.

4. Why Better Than Training the Small Model Directly?

Training methodInformation sourceGeneralization
Direct trainingHard label (one-hot)Weak: inter-class relations are lost
Knowledge distillationTeacher soft label + hard labelStrong: inherits the teacher’s dark knowledge and ranking

Typical gain: student parameters cut 10x to 50x, accuracy loss 1% to 3%.

5. Representative Work and Measured Gains

WorkTeacherStudentKey result
DeepSeek R1 distillationR1 671B MoEQwen2.5 / Llama 7B-70B7B reaches 92.8% on MATH-500 (up from 58.8%)
Meta Muse GlimmerMuse Spark 1.230B denseSWE-Bench Verified 76.0; 3.1x speedup on RTX 5090 with DFlash
DistilBERTBERT-base40% of the parameters60% faster, retains 97% of performance
MiniMax on-device distillationLarge MoESmall denseFor on-device deployment

6. Relationship to DeepSeek V4’s OPD

DeepSeek V4 uses OPD (On-Policy Distillation): the main model distills “domain expert” small models online, compressing specialized capability into one unified model.

How it differs from conventional offline distillation:

Offline:  teacher is frozen -> generates soft labels -> student learns
Online:   teacher and student train together; the student's input distribution
          keeps updating as the policy updates

OPD avoids distribution shift, which suits unified models spanning many tasks and domains.

7. Embodied Intelligence: On-Device VLA Deployment

Robots have limited onboard compute (a Jetson Orin is on the order of tens to just over a hundred TOPS), so a 7B+ VLA cannot run locally. Distillation is the key path:

Large cloud VLA (teacher)
    | distill
    v
Small on-device VLA (student)
    | deploy
    v
Real-time action inference on the robot

In NVIDIA GR00T’s “three-computer architecture,” DGX trains the large model and the distilled result goes to Jetson. Domestic VLA companies (Zhiyuan, Unitree, Fourier, and others) have to walk the same road.

8. Common Pitfalls

  1. T too small — close to a hard label, so the dark knowledge never transfers;
  2. T too large — the distribution flattens and the student learns no discrimination;
  3. Student and teacher architectures too different — the capability gap is too wide and distillation becomes inefficient;
  4. Wrong distillation data distribution — it must cover the teacher’s capability boundary, not be randomly sampled;
  5. Distilling only the final output — intermediate feature distillation transfers more information;
  6. Ignoring the hard label — using soft labels alone removes the strong ground-truth constraint.

9. Pseudocode

import torch.nn.functional as F

def distill_loss(teacher_logits, student_logits, labels, T=4.0, alpha=0.7):
    # soft target
    p_t = F.softmax(teacher_logits / T, dim=-1)
    p_s = F.log_softmax(student_logits / T, dim=-1)
    loss_kl = F.kl_div(p_s, p_t, reduction='batchmean') * (T * T)

    # hard label
    loss_ce = F.cross_entropy(student_logits, labels)

    return alpha * loss_kl + (1 - alpha) * loss_ce

10. Summary

ProblemDistillation’s answer
Large models are too expensiveDistill a small model, cost down 10x to 50x
Small models are not strong enoughInherit the teacher’s dark knowledge, generalization up
On-device deploymentLarge VLA to small VLA to Jetson / robot chip
Unifying many tasksOPD online distillation avoids distribution shift

One line to remember: distillation does not shrink a large model; it teaches the small model the things the large model “knows but never said out loud.”

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。