Math & training · Glossary term
What is SFT (Supervised Fine-Tuning)?
Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training distribution.
“Training on example inputs and desired outputs.”
What is the common confusion about SFT (Supervised Fine-Tuning)?
SFT can adapt many behaviors beyond chat, and example quality determines what behavior is reinforced.
Learn SFT (Supervised Fine-Tuning) in the course
Lessons that name SFT (Supervised Fine-Tuning) in a title or section
- Instruction Tuning (SFT)
A base model predicts the next token. That's it. It doesn't follow instructions, answer questions, or refuse harmful requests. SFT is the bridge between a token predictor and a useful assistant.
- Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve)
An 8B model trained on your own data, DPO-aligned on your own preferences, quantized, speculative-decoded, and served at measurable $/1M tokens.
- Capstone Lesson 39: Instruction Tuning by Supervised Fine-Tuning
A pretrained base model can extend a sequence but cannot follow an instruction. Supervised fine-tuning is the smallest change that fixes this: feed the model paired examples of an instruction and a…
- Instruction-Following as Alignment Signal
Every later critique of RLHF argues against this pipeline. Before you study how optimization pressure distorts a proxy, you have to see the proxy.
Covered in Phase 10: LLMs from Scratch, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
- RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
- Instruction FollowingA model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
- Transfer LearningStarting from representations or parameters learned on one data distribution or objective and adapting them for another.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.