Math & training · Glossary term

What is Warmup?

An initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.

Why does Warmup matter?

Early gradients and optimizer statistics can be unstable, especially in large-batch or transformer training, so abrupt full-size updates may damage optimization.

Warmup in practice

Define warmup in steps or processed tokens, log the realized curve, and tune it with the batch, optimizer, and total training budget held visible.

What is the common confusion about Warmup?

Warmup is not required for every model and does not make an otherwise unsuitable learning rate safe.

Learn Warmup in the course

Lessons that name Warmup in a title or section

  • Learning Rate Schedules and Warmup

    The learning rate is the single most important hyperparameter. Not the architecture. Not the dataset size. Not the activation function. The learning rate. If you tune nothing else, tune this.

    Phase 03: Deep Learning Core

  • Cosine LR with Linear Warmup

    The learning-rate schedule is the second most important decision after the loss function. AdamW with a cosine decay and a linear warmup is the modern default for language-model training because it…

    Phase 19: Capstone Projects

  • Training Loop and Evaluation

    A loop that does not measure is a loop that lies. This lesson builds the training loop that drives the GPT model: AdamW with weight decay split, a warmup plus cosine learning rate schedule, a…

    Phase 19: Capstone Projects

Covered in Phase 03: Deep Learning Core and Phase 19: Capstone Projects.

  • Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
  • AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.