Math & training · Glossary term

What is Optimizer?

An algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and other optimizers change the update using history or adaptive scaling. Each choice has different memory, stability, and tuning behavior.

What people say

“The algorithm that updates weights.”

What is the common confusion about Optimizer?

The optimizer consumes gradients; backpropagation computes them.

Learn Optimizer in the course

Lessons that name Optimizer in a title or section

  • Optimizers

    Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.

    Phase 03: Deep Learning Core

  • ZeRO Optimizer State Sharding

    Adam stores two moment estimates per parameter, both in float32. A 7B-parameter model carries 56 GB of optimiser state. ZeRO stage 1 shards that across N ranks; each rank owns 1/N of the optimiser.

    Phase 19: Capstone Projects

  • Build Your Own Mini Framework

    You have built neurons, layers, networks, backprop, activations, loss functions, optimizers, regularization, initialization, and LR schedules. All as separate pieces.

    Phase 03: Deep Learning Core

  • Introduction to PyTorch

    You built the engine from pistons and crankshafts. Now learn the one everyone actually drives. Build and train neural networks using PyTorch's nn.Module, nn.Sequential, and autograd.

    Phase 03: Deep Learning Core

  • Introduction to JAX

    PyTorch mutates tensors. TensorFlow builds graphs. JAX compiles pure functions. That last one changes how you think about deep learning.

    Phase 03: Deep Learning Core

Covered in Phase 03: Deep Learning Core and Phase 19: Capstone Projects.

  • Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
  • AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
  • Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
  • Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
  • Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
  • Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
  • WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
  • Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.