Math & training · Glossary term

What is Knowledge Distillation?

Training a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary target labels.

Why does Knowledge Distillation matter?

It can transfer useful behavior into a smaller or cheaper model when serving the teacher directly is impractical.

Knowledge Distillation in practice

Define teacher outputs, temperature, student loss, and a held-out eval set, then compare the student with both the teacher and a label-only baseline.

What is the common confusion about Knowledge Distillation?

Distillation transfers behavior on the training distribution; it does not copy every capability, fact, or safety property of the teacher.

Learn Knowledge Distillation in the course

No lesson links to this term yet. Search the course catalog for it.

  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.