Math & training · Glossary term
What is Gradient Accumulation?
Summing or averaging gradients from several microbatches before performing one optimizer update.
Why does Gradient Accumulation matter?
It lets you approximate a larger effective batch when one device cannot hold all examples and activations at once.
Gradient Accumulation in practice
Scale the loss consistently, call the optimizer only after the chosen number of microbatches, and measure whether normalization or distributed synchronization changes behavior.
What is the common confusion about Gradient Accumulation?
Gradient accumulation reduces per-step activation memory, but it does not reproduce every property of processing the full batch simultaneously.
Learn Gradient Accumulation in the course
Lessons that name Gradient Accumulation in a title or section
- Gradient Accumulation
Train at an effective batch you cannot afford, one micro-batch at a time. Scale the loss, hold the optimizer step, and let the gradients pile up.
Covered in Phase 19: Capstone Projects.
Related terms
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.