Knowledge Distillation

anchor

Train a smaller model to match a larger model's output distribution rather than the hard labels.

Findings