Math & training · Glossary term
What is LoRA (Low-Rank Adaptation)?
A method that keeps base weights frozen and learns low-rank update matrices for selected layers. It reduces the number of trainable parameters and can lower training memory relative to full-parameter fine-tuning.
“Parameter-efficient fine-tuning.”
What is the common confusion about LoRA (Low-Rank Adaptation)?
Actual memory and speed savings depend on rank, target modules, optimizer state, activation memory, quantization, and implementation.
Learn LoRA (Low-Rank Adaptation) in the course
Start with
- Fine-Tuning with LoRA & QLoRA
Full fine-tuning a 7B model requires 56GB of VRAM. You don't have that. Neither do most companies. LoRA lets you fine-tune the same model in 6GB by training less than 1% of the parameters.
Lessons that name LoRA (Low-Rank Adaptation) in a title or section
- ControlNet, LoRA & Conditioning
Text alone is a clumsy control signal. ControlNet lets you clone a pretrained diffusion model and steer it with a depth map, pose skeleton, scribble, or edge image.
- Stable Diffusion — Architecture & Fine-Tuning
Stable Diffusion is a DDPM that runs in the latent space of a pretrained VAE, conditioned on text via cross-attention, sampled with a fast deterministic ODE solver, and steered by classifier-free…
- Vision-Language Models — The ViT-MLP-LLM Pattern
A vision encoder converts an image into tokens. An MLP projector maps those tokens into the LLM's embedding space. A language model does the rest.
- Whisper — Architecture & Fine-Tuning
Whisper is a 30-second-window transformer encoder-decoder, trained on 680k hours of multilingual weakly-supervised audio-text pairs. One architecture, multiple tasks, robust across 99 languages.
- Differential Privacy for LLMs
DP-SGD remains the standard — noise-injected gradient updates provide formal (epsilon, delta) guarantees. Overhead in compute, memory, and utility is substantial; parameter-efficient DP fine-tuning…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 04: Computer Vision, Phase 06: Speech & Audio, Phase 08: Generative AI and Phase 18: Ethics, Safety & Alignment.
Related terms
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- QLoRAA parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training…
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.