Math & training · Glossary term
What is Mixed Precision?
A numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher precision for values that need more range or stability.
“Using lower-precision arithmetic for speed and memory savings.”
What is the common confusion about Mixed Precision?
Speed, memory, and accuracy effects depend on hardware, data type, scaling method, kernels, and model. They are not a fixed multiplier.
Learn Mixed Precision in the course
Lessons that name Mixed Precision in a title or section
- Gradient Clipping and Mixed Precision
The optimizer and schedule from the previous lesson assume gradients are sane. They usually are not. A single bad batch can spike the gradient norm by three orders of magnitude.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
- Scaling: Distributed Training, FSDP, DeepSpeed
Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.
Covered in Phase 01: Math Foundations, Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.
Related terms
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- CUDANVIDIA's platform and programming model for general-purpose computation on compatible GPUs.
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
- Activation CheckpointingA training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
- FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
- Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
- Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.