Math & training · Glossary term

What is Batch Size?

The number of examples whose losses contribute to one gradient estimate before an optimizer update. Larger batches can improve hardware utilization and reduce gradient noise, but they require more memory and may need different learning-rate or scheduling choices.

What people say

“How many examples are processed at once.”

What is the common confusion about Batch Size?

There is no universal batch-size range or rule that says every batch increase should produce the same learning-rate increase.

Learn Batch Size in the course

No lesson links to this term yet. Search the course catalog for it.

  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • EpochOne traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the…
  • Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
  • HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
  • Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
  • Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
  • WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.