Math & training · Glossary term
What is Backpropagation?
An efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph. It computes gradients; an optimizer uses those gradients to update parameters.
“How neural networks learn.”
What is the common confusion about Backpropagation?
Backpropagation calculates gradients. It does not choose the update rule or learning rate.
Why is it called Backpropagation?
Derivative information moves backward from the loss toward earlier operations.
Learn Backpropagation in the course
Start with
- Backpropagation from Scratch
Backpropagation is the algorithm that makes learning possible. Without it, neural networks are just expensive random number generators.
Taught in Phase 03: Deep Learning Core.
Related terms
- AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- Activation CheckpointingA training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
- Activation FunctionA function applied after a linear or affine layer that introduces nonlinearity. Without it, composing layers with weights and biases…
- Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.