Math & training · Glossary term
What is Gradient Clipping?
Limiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
Why does Gradient Clipping matter?
It can prevent an unusually large gradient from destabilizing a training step and producing non-finite values.
Gradient Clipping in practice
Log unclipped norms, clip after unscaling mixed-precision gradients, and investigate repeated clipping instead of treating it as a substitute for diagnosing instability.
What is the common confusion about Gradient Clipping?
Clipping controls update magnitude; it does not repair invalid data, a broken loss, or a consistently unsuitable learning rate.
Learn Gradient Clipping in the course
Lessons that name Gradient Clipping in a title or section
- Gradient Clipping and Mixed Precision
The optimizer and schedule from the previous lesson assume gradients are sane. They usually are not. A single bad batch can spike the gradient norm by three orders of magnitude.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
Covered in Phase 01: Math Foundations and Phase 19: Capstone Projects.
Related terms
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.