0. One-Line Positioning
Knowledge distillation (KD) = let a large model (the teacher) pass its “dark knowledge” to a small model (the student), so the student approaches the teacher’s performance with far fewer parameters.
The problem it solves: model capability is not the same as deployment cost. Large models are strong but expensive to serve; small models are cheap but, without enough training data, never learn that well. Distillation lets the small model “stand on the giant’s shoulders.”
1. Why Call It “Dark Knowledge”?
Take an image classification task where the teacher outputs:
cat: 0.70
dog: 0.20
table: 0.05
chair: 0.05
A hard label only tells the model “this is a cat.” The teacher’s soft label additionally implies “this looks like a cat but also somewhat like a dog, and not much like furniture.” Those inter-class similarity signals are the dark knowledge. A student that learns them generalizes better than one trained on one-hot labels alone.
2. Classic Hinton Distillation
Figure: distillation softens the teacher's soft target with temperature T and uses it as the student's training objective.
3. What Temperature T Does
T = 1: close to the raw distribution, little dark knowledge
T = 4 to 10: smoother distribution, inter-class relations clearer, dark knowledge rich
T -> inf: the distribution approaches uniform and loses discriminative power
Empirical values: T = 4 to 10, with a distillation weight of alpha = 0.5 to 0.9.
4. Why Better Than Training the Small Model Directly?
| Training method | Information source | Generalization |
|---|---|---|
| Direct training | Hard label (one-hot) | Weak: inter-class relations are lost |
| Knowledge distillation | Teacher soft label + hard label | Strong: inherits the teacher’s dark knowledge and ranking |
Typical gain: student parameters cut 10x to 50x, accuracy loss 1% to 3%.
5. Representative Work and Measured Gains
| Work | Teacher | Student | Key result |
|---|---|---|---|
| DeepSeek R1 distillation | R1 671B MoE | Qwen2.5 / Llama 7B-70B | 7B reaches 92.8% on MATH-500 (up from 58.8%) |
| Meta Muse Glimmer | Muse Spark 1.2 | 30B dense | SWE-Bench Verified 76.0; 3.1x speedup on RTX 5090 with DFlash |
| DistilBERT | BERT-base | 40% of the parameters | 60% faster, retains 97% of performance |
| MiniMax on-device distillation | Large MoE | Small dense | For on-device deployment |
6. Relationship to DeepSeek V4’s OPD
DeepSeek V4 uses OPD (On-Policy Distillation): the main model distills “domain expert” small models online, compressing specialized capability into one unified model.
How it differs from conventional offline distillation:
Offline: teacher is frozen -> generates soft labels -> student learns
Online: teacher and student train together; the student's input distribution
keeps updating as the policy updates
OPD avoids distribution shift, which suits unified models spanning many tasks and domains.
7. Embodied Intelligence: On-Device VLA Deployment
Robots have limited onboard compute (a Jetson Orin is on the order of tens to just over a hundred TOPS), so a 7B+ VLA cannot run locally. Distillation is the key path:
Large cloud VLA (teacher)
| distill
v
Small on-device VLA (student)
| deploy
v
Real-time action inference on the robot
In NVIDIA GR00T’s “three-computer architecture,” DGX trains the large model and the distilled result goes to Jetson. Domestic VLA companies (Zhiyuan, Unitree, Fourier, and others) have to walk the same road.
8. Common Pitfalls
- T too small — close to a hard label, so the dark knowledge never transfers;
- T too large — the distribution flattens and the student learns no discrimination;
- Student and teacher architectures too different — the capability gap is too wide and distillation becomes inefficient;
- Wrong distillation data distribution — it must cover the teacher’s capability boundary, not be randomly sampled;
- Distilling only the final output — intermediate feature distillation transfers more information;
- Ignoring the hard label — using soft labels alone removes the strong ground-truth constraint.
9. Pseudocode
import torch.nn.functional as F
def distill_loss(teacher_logits, student_logits, labels, T=4.0, alpha=0.7):
# soft target
p_t = F.softmax(teacher_logits / T, dim=-1)
p_s = F.log_softmax(student_logits / T, dim=-1)
loss_kl = F.kl_div(p_s, p_t, reduction='batchmean') * (T * T)
# hard label
loss_ce = F.cross_entropy(student_logits, labels)
return alpha * loss_kl + (1 - alpha) * loss_ce
10. Summary
| Problem | Distillation’s answer |
|---|---|
| Large models are too expensive | Distill a small model, cost down 10x to 50x |
| Small models are not strong enough | Inherit the teacher’s dark knowledge, generalization up |
| On-device deployment | Large VLA to small VLA to Jetson / robot chip |
| Unifying many tasks | OPD online distillation avoids distribution shift |
One line to remember: distillation does not shrink a large model; it teaches the small model the things the large model “knows but never said out loud.”
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。