Math & training · Glossary term
What is Weight Decay?
An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.
“Regularization that shrinks weights during optimization.”
Why does Weight Decay matter?
It can improve generalization, but the useful coefficient and excluded parameter groups depend on model, optimizer, schedule, and data.
What is the common confusion about Weight Decay?
Decoupled weight decay is equivalent to an L2 loss penalty for some simple optimizers, but not generally for adaptive optimizers such as Adam.
Learn Weight Decay in the course
Lessons that name Weight Decay in a title or section
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- Regularization
Your model gets 99% on training data and 60% on test data. It memorized instead of learning. Regularization is the tax you impose on complexity to force generalization.
Covered in Phase 03: Deep Learning Core.
Related terms
- AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
- OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.