Math & training · Glossary term
What is Knowledge Distillation?
Training a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary target labels.
Why does Knowledge Distillation matter?
It can transfer useful behavior into a smaller or cheaper model when serving the teacher directly is impractical.
Knowledge Distillation in practice
Define teacher outputs, temperature, student loss, and a held-out eval set, then compare the student with both the teacher and a label-only baseline.
What is the common confusion about Knowledge Distillation?
Distillation transfers behavior on the training distribution; it does not copy every capability, fact, or safety property of the teacher.
Learn Knowledge Distillation in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
- QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.