Math & training · Glossary term

What is Weight Decay?

An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.

What people say

“Regularization that shrinks weights during optimization.”

Why does Weight Decay matter?

It can improve generalization, but the useful coefficient and excluded parameter groups depend on model, optimizer, schedule, and data.

What is the common confusion about Weight Decay?

Decoupled weight decay is equivalent to an L2 loss penalty for some simple optimizers, but not generally for adaptive optimizers such as Adam.

Learn Weight Decay in the course

Lessons that name Weight Decay in a title or section

  • Optimizers

    Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.

    Phase 03: Deep Learning Core

  • Regularization

    Your model gets 99% on training data and 60% on test data. It memorized instead of learning. Regularization is the tax you impose on complexity to force generalization.

    Phase 03: Deep Learning Core

Covered in Phase 03: Deep Learning Core.

  • AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
  • OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.