Math & training · Glossary term

What is Gradient Accumulation?

Summing or averaging gradients from several microbatches before performing one optimizer update.

Why does Gradient Accumulation matter?

It lets you approximate a larger effective batch when one device cannot hold all examples and activations at once.

Gradient Accumulation in practice

Scale the loss consistently, call the optimizer only after the chosen number of microbatches, and measure whether normalization or distributed synchronization changes behavior.

What is the common confusion about Gradient Accumulation?

Gradient accumulation reduces per-step activation memory, but it does not reproduce every property of processing the full batch simultaneously.

Learn Gradient Accumulation in the course

Lessons that name Gradient Accumulation in a title or section

  • Gradient Accumulation

    Train at an effective batch you cannot afford, one micro-batch at a time. Scale the loss, hold the optimizer step, and let the gradients pile up.

    Phase 19: Capstone Projects

Covered in Phase 19: Capstone Projects.

  • Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
  • Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.