Math & training · Glossary term
What is Activation Checkpointing?
A training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
Why does Activation Checkpointing matter?
It lets you train larger models or sequences within a fixed memory budget by trading additional computation for lower activation storage.
Activation Checkpointing in practice
Checkpoint the memory-heavy transformer blocks, measure the extra step time, and keep recovery checkpoints separate from activation-recomputation settings.
What is the common confusion about Activation Checkpointing?
Activation checkpointing is not a durable training checkpoint. It helps one forward and backward pass fit in memory but cannot resume a crashed run.
Learn Activation Checkpointing in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
- BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
Sources
More terms in Math & training
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.