Math & training · Glossary term
What is Gradient?
A vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent) to minimize the loss.
“The slope of the loss.”
What is the common confusion about Gradient?
Optimizers can transform, average, clip, or adapt gradients instead of taking a plain negative-gradient step.
Learn Gradient in the course
Lessons that name Gradient in a title or section
- Policy Gradient — REINFORCE from Scratch
Stop estimating value. Parameterize the policy directly, compute the gradient of expected return, step uphill. Williams (1992) wrote it in one theorem.
- Gradient Checkpointing and Activation Recomputation
Backprop keeps every intermediate activation. At 70B parameters and 128K context that is 3 TB of activations per rank. Checkpointing trades FLOPs for memory: recompute instead of save.
- Gradient Clipping and Mixed Precision
The optimizer and schedule from the previous lesson assume gradients are sane. They usually are not. A single bad batch can spike the gradient norm by three orders of magnitude.
- Gradient Accumulation
Train at an effective batch you cannot afford, one micro-batch at a time. Scale the loss, hold the optimizer step, and let the gradients pile up.
- Calculus for Machine Learning
Derivatives tell you which way is downhill. That is all a neural network needs to learn. Language: Python Compute numerical and analytical derivatives for common ML functions (x^2, sigmoid,…
- Chain Rule & Automatic Differentiation
The chain rule is the engine behind every neural network that learns. Language: Python Build a minimal autograd engine (Value class) that records operations and computes gradients via reverse-mode…
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 05: NLP: Foundations to Advanced, Phase 08: Generative AI, Phase 09: Reinforcement Learning, Phase 10: LLMs from Scratch, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- Activation FunctionA function applied after a linear or affine layer that introduces nonlinearity. Without it, composing layers with weights and biases…
- AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- ReLURectified Linear Unit, defined as `f(x) = max(0, x)`. It is inexpensive and has a non-saturating positive branch, though zero gradients on…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.