Math & training · Glossary term

What is Mixed Precision?

A numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher precision for values that need more range or stability.

What people say

“Using lower-precision arithmetic for speed and memory savings.”

What is the common confusion about Mixed Precision?

Speed, memory, and accuracy effects depend on hardware, data type, scaling method, kernels, and model. They are not a fixed multiplier.

Learn Mixed Precision in the course

Lessons that name Mixed Precision in a title or section

  • Gradient Clipping and Mixed Precision

    The optimizer and schedule from the previous lesson assume gradients are sane. They usually are not. A single bad batch can spike the gradient norm by three orders of magnitude.

    Phase 19: Capstone Projects

  • Numerical Stability

    Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…

    Phase 01: Math Foundations

  • Scaling: Distributed Training, FSDP, DeepSpeed

    Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.

    Phase 10: LLMs from Scratch

Covered in Phase 01: Math Foundations, Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.

  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • CUDANVIDIA's platform and programming model for general-purpose computation on compatible GPUs.
  • NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • Activation CheckpointingA training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
  • FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
  • Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
  • Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
  • NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.