Models & inference · Glossary term

What is Quantization?

Representing weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost. Methods differ in calibration, granularity, data type, and whether conversion happens before, during, or after training.

What people say

“Storing or computing model values with fewer bits.”

What is the common confusion about Quantization?

Moving from one nominal bit width to another does not guarantee the same end-to-end memory or speed ratio because metadata, kernels, caches, and hardware support also matter.

Learn Quantization in the course

Lessons that name Quantization in a title or section

Covered in Phase 06: Speech & Audio, Phase 10: LLMs from Scratch, Phase 11: LLM Engineering and Phase 17: Infrastructure & Production.

  • QLoRAA parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training…
  • Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • Knowledge DistillationTraining a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.