Math & training · Glossary term
What is RLHF (Reinforcement Learning from Human Feedback)?
A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.
“Training a model from human preferences.”
What is the common confusion about RLHF (Reinforcement Learning from Human Feedback)?
RLHF optimizes a proxy learned from collected feedback. It does not guarantee broad alignment with every user or situation.
Learn RLHF (Reinforcement Learning from Human Feedback) in the course
Start with
- RLHF: Reward Model + PPO
SFT teaches the model to follow instructions. But it doesn't teach the model which response is BETTER. Two grammatically correct, factually accurate answers can differ enormously in helpfulness.
Lessons that name RLHF (Reinforcement Learning from Human Feedback) in a title or section
- Reward Modeling & RLHF
Humans cannot write a reward function for "good assistant response," but they can compare two responses and pick the better one.
- Sycophancy as RLHF Amplification
Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…
- DPO: Direct Preference Optimization
RLHF works. It also requires training three models (SFT, reward model, policy), managing PPO's instability, and tuning a KL penalty. DPO asks: what if you could skip all of that?
- Constitutional AI and RLAIF
Bai et al. (arXiv:2212.08073, 2022) asked: what if we replaced the human labeler with an AI that reads a list of principles?
Taught in Phase 10: LLMs from Scratch.
Also covered in Phase 09: Reinforcement Learning and Phase 18: Ethics, Safety & Alignment.
Related terms
- DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
- SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.