Math & training · Glossary term

What is AdamW?

An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.

What people say

“Adam with weight decay fixed.”

What is the common confusion about AdamW?

Decoupled weight decay does not make AdamW universally optimal. Model, data, and training scale still determine the best optimizer and schedule.

Learn AdamW in the course

Lessons that name AdamW in a title or section

  • Optimizers

    Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.

    Phase 03: Deep Learning Core

  • Training Loop and Evaluation

    A loop that does not measure is a loop that lies. This lesson builds the training loop that drives the GPT model: AdamW with weight decay split, a warmup plus cosine learning rate schedule, a…

    Phase 19: Capstone Projects

Covered in Phase 03: Deep Learning Core and Phase 19: Capstone Projects.

  • Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
  • Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.