Knowledge Distillation
A large model performs a task well but is too expensive for everyday use. A smaller model could try to reproduce its behavior. It needs examples or assessments showing more than just a list of correct labels, however.
Knowledge Distillation trains a student model using a teacher model's outputs or predictions. The student actually changes its parameters; simply sending a query to the teacher is not training.
If the teacher assigns A probability 0.60, B 0.35, and C 0.05, the label “A” omits the fact that B is a close alternative. The student can learn to reproduce the entire distribution. Another variant uses complete texts generated by the teacher.
The teacher need not be a single larger model, but classically serves as a source of behavior to imitate. Distillation can also transfer errors. KL Divergence helps measure agreement between predictions, while On-policy self-distillation is a more specific arrangement of data and teacher.
Mechanism and details
In Distilling the Knowledge in a Neural Network, §2, Hinton, Vinyals and Dean describe learning from soft probability distributions. Increased Temperature reveals more information about relationships between candidates. The same parameter is used on the teacher and student sides when comparing distributions. The authors combine this objective with learning from correct labels when available.
More information than a single answer
An original example: the teacher assigns illustrative tokens A, B, and C probabilities of 0.60, 0.35, and 0.05. The label “A” identifies only the winner. The full distribution also says that B was much more likely than C. The student can try to reproduce these proportions using Cross-entropy or appropriately directed KL Divergence. This is an additional training signal, not evidence that the teacher's labels are correct.
In text-generating models, comparing next-token distributions must be distinguished from training on entire generated sequences. Kim and Rush, Sequence-Level Knowledge Distillation, §2.2–3.2 describe both approaches for machine translation. In the sequence-level variant, the student trains on outputs found by the teacher's beam search. The text of such a response does not contain the full probabilities of all alternative tokens.
The result depends on the teacher, the data, and the student's capacity. It does not mean copying every capability or guaranteeing an improvement in every task. The separate note On-policy self-distillation describes a particular variant in TTPO; its conditions are not a general definition of Knowledge Distillation.