Math & training · Glossary term
What is AdamW?
An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.
“Adam with weight decay fixed.”
What is the common confusion about AdamW?
Decoupled weight decay does not make AdamW universally optimal. Model, data, and training scale still determine the best optimizer and schedule.
Learn AdamW in the course
Lessons that name AdamW in a title or section
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- Training Loop and Evaluation
A loop that does not measure is a loop that lies. This lesson builds the training loop that drives the GPT model: AdamW with weight decay split, a warmup plus cosine learning rate schedule, a…
Covered in Phase 03: Deep Learning Core and Phase 19: Capstone Projects.
Related terms
- Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
- Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.