Math & training · Glossary term

What is RLHF (Reinforcement Learning from Human Feedback)?

A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.

What people say

“Training a model from human preferences.”

What is the common confusion about RLHF (Reinforcement Learning from Human Feedback)?

RLHF optimizes a proxy learned from collected feedback. It does not guarantee broad alignment with every user or situation.

Learn RLHF (Reinforcement Learning from Human Feedback) in the course

Start with

  • RLHF: Reward Model + PPO

    SFT teaches the model to follow instructions. But it doesn't teach the model which response is BETTER. Two grammatically correct, factually accurate answers can differ enormously in helpfulness.

    Phase 10: LLMs from Scratch

Lessons that name RLHF (Reinforcement Learning from Human Feedback) in a title or section

  • Reward Modeling & RLHF

    Humans cannot write a reward function for "good assistant response," but they can compare two responses and pick the better one.

    Phase 09: Reinforcement Learning

  • Sycophancy as RLHF Amplification

    Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…

    Phase 18: Ethics, Safety & Alignment

  • DPO: Direct Preference Optimization

    RLHF works. It also requires training three models (SFT, reward model, policy), managing PPO's instability, and tuning a KL penalty. DPO asks: what if you could skip all of that?

    Phase 10: LLMs from Scratch

  • Constitutional AI and RLAIF

    Bai et al. (arXiv:2212.08073, 2022) asked: what if we replaced the human labeler with an AI that reads a list of principles?

    Phase 18: Ethics, Safety & Alignment

Taught in Phase 10: LLMs from Scratch.

Also covered in Phase 09: Reinforcement Learning and Phase 18: Ethics, Safety & Alignment.

  • DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
  • SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…

Sources

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.