Math & training · Glossary term

What is SFT (Supervised Fine-Tuning)?

Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training distribution.

What people say

“Training on example inputs and desired outputs.”

What is the common confusion about SFT (Supervised Fine-Tuning)?

SFT can adapt many behaviors beyond chat, and example quality determines what behavior is reinforced.

Learn SFT (Supervised Fine-Tuning) in the course

Lessons that name SFT (Supervised Fine-Tuning) in a title or section

Covered in Phase 10: LLMs from Scratch, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
  • RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
  • Instruction FollowingA model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
  • Transfer LearningStarting from representations or parameters learned on one data distribution or objective and adapting them for another.

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.