Math & training · Glossary term

What is DPO (Direct Preference Optimization)?

A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy. It avoids running an explicit reward model and reinforcement-learning loop during this stage.

What people say

“Preference training without a separate reward-model stage.”

What is the common confusion about DPO (Direct Preference Optimization)?

DPO still depends on the quality and coverage of preference data and does not eliminate evaluation or alignment risk.

Learn DPO (Direct Preference Optimization) in the course

Start with

  • DPO: Direct Preference Optimization

    RLHF works. It also requires training three models (SFT, reward model, policy), managing PPO's instability, and tuning a KL penalty. DPO asks: what if you could skip all of that?

    Phase 10: LLMs from Scratch

Lessons that name DPO (Direct Preference Optimization) in a title or section

Taught in Phase 10: LLMs from Scratch.

Also covered in Phase 15: Autonomous Systems, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
  • SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.