Math & training · Glossary term
What is Learning Rate?
A scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that are too small can make useful progress impractically slow.
“How large each optimization step is.”
What is the common confusion about Learning Rate?
The effective update also depends on the optimizer, schedule, gradient scale, batch, and parameter history.
Learn Learning Rate in the course
Lessons that name Learning Rate in a title or section
- Learning Rate Schedules and Warmup
The learning rate is the single most important hyperparameter. Not the architecture. Not the dataset size. Not the activation function. The learning rate. If you tune nothing else, tune this.
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
- Hyperparameter Tuning
Hyperparameters are the knobs you turn before training starts. Turning them well is the difference between a mediocre model and a great one.
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- Introduction to PyTorch
You built the engine from pistons and crankshafts. Now learn the one everyone actually drives. Build and train neural networks using PyTorch's nn.Module, nn.Sequential, and autograd.
- Debugging Neural Networks
Your network compiled. It ran. It produced a number. The number is wrong and nothing crashed. Welcome to the hardest kind of debugging -- the kind where there is no error message.
- Transfer Learning & Fine-Tuning
Somebody else spent a million GPU hours teaching a network what edges, textures, and object parts look like. You should borrow those features before training your own.
Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core and Phase 04: Computer Vision.
Related terms
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
- Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
- WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.