Math & training · Glossary term

What is QLoRA?

A parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training LoRA adapters with higher-precision computation where needed.

What people say

“LoRA with a quantized base model.”

Why does QLoRA matter?

It can reduce the memory needed to adapt large models, but savings and quality depend on model, rank, optimizer, sequence length, hardware, and implementation.

What is the common confusion about QLoRA?

QLoRA does not guarantee a particular memory footprint or a fixed quality gap from full fine-tuning.

Learn QLoRA in the course

Start with

  • Fine-Tuning with LoRA & QLoRA

    Full fine-tuning a 7B model requires 56GB of VRAM. You don't have that. Neither do most companies. LoRA lets you fine-tune the same model in 6GB by training less than 1% of the parameters.

    Phase 11: LLM Engineering

Lessons that name QLoRA in a title or section

Taught in Phase 11: LLM Engineering.

Also covered in Phase 04: Computer Vision.

  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.