Math & training · Glossary term

What is Adam (Optimizer)?

Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias correction, and adapts the update scale per parameter. It is a useful baseline, but it still needs a suitable learning rate and schedule.

What people say

“The optimizer you use without thinking about it.”

What is the common confusion about Adam (Optimizer)?

Adam is a strong baseline, not a universal best optimizer.

Learn Adam (Optimizer) in the course

Lessons that name Adam (Optimizer) in a title or section

  • Optimization

    Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.

    Phase 01: Math Foundations

  • Optimizers

    Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.

    Phase 03: Deep Learning Core

  • Build Your Own Mini Framework

    You have built neurons, layers, networks, backprop, activations, loss functions, optimizers, regularization, initialization, and LR schedules. All as separate pieces.

    Phase 03: Deep Learning Core

Covered in Phase 01: Math Foundations and Phase 03: Deep Learning Core.

  • AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.