Back to archive
#ai#llm#training#papers#aigen

On-policy self-distillation

You want to improve a model but do not have a separate, stronger teacher. The same model can receive an additional hint, such as the answer in the task text. Can its better-informed version teach the version without that hint?

On-policy self-distillation uses this relationship. The student produces samples according to its current way of responding — hence “on-policy”. The teacher comes from the same model — hence “self” — and assesses its successive tokens, pieces of text. The student learns from these assessments rather than simply copying a ready document.

In the variant described in TTPO, §2.2 and §3.3, the teacher receives the answer as additional context. Next-token predictions are compared on text produced by the student. This is a particular form of Knowledge Distillation.

Additional context may improve the signal, but a wrong hint can distort it. The method does not prove the model created new truth without a source. What matters is where samples come from, what privilege the teacher has, and which of its predictions we actually use.

Connections

Voting lies. Disagreement usually does not