Math & training · Glossary term
What is Learning Rate Schedule?
A policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
Why does Learning Rate Schedule matter?
Different training stages can benefit from different update scales, so one constant rate may be unstable early or wasteful late.
Learning Rate Schedule in practice
Version the schedule with the optimizer configuration, log the actual rate at every step, and compare schedules under the same token or update budget.
What is the common confusion about Learning Rate Schedule?
A scheduler controls the learning rate over time; it does not decide when an optimizer step occurs or guarantee convergence.
Learn Learning Rate Schedule in the course
Lessons that name Learning Rate Schedule in a title or section
- Learning Rate Schedules and Warmup
The learning rate is the single most important hyperparameter. Not the architecture. Not the dataset size. Not the activation function. The learning rate. If you tune nothing else, tune this.
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
Covered in Phase 01: Math Foundations and Phase 03: Deep Learning Core.
Related terms
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- EpochOne traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the…
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.