Math & training · Glossary term
What is Optimizer?
An algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and other optimizers change the update using history or adaptive scaling. Each choice has different memory, stability, and tuning behavior.
“The algorithm that updates weights.”
What is the common confusion about Optimizer?
The optimizer consumes gradients; backpropagation computes them.
Learn Optimizer in the course
Lessons that name Optimizer in a title or section
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- ZeRO Optimizer State Sharding
Adam stores two moment estimates per parameter, both in float32. A 7B-parameter model carries 56 GB of optimiser state. ZeRO stage 1 shards that across N ranks; each rank owns 1/N of the optimiser.
- Build Your Own Mini Framework
You have built neurons, layers, networks, backprop, activations, loss functions, optimizers, regularization, initialization, and LR schedules. All as separate pieces.
- Introduction to PyTorch
You built the engine from pistons and crankshafts. Now learn the one everyone actually drives. Build and train neural networks using PyTorch's nn.Module, nn.Sequential, and autograd.
- Introduction to JAX
PyTorch mutates tensors. TensorFlow builds graphs. JAX compiles pure functions. That last one changes how you think about deep learning.
Covered in Phase 03: Deep Learning Core and Phase 19: Capstone Projects.
Related terms
- Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
- AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
- Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
- WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
- Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.