Phase 18: Ethics, Safety & Alignment

Reward Hacking and Goodhart's Law

Any optimizer strong enough to maximize a proxy reward will find the gap between the proxy and the thing you actually wanted. Gao et al. (ICML 2023) gave this a scaling law: proxy reward increases, gold reward peaks then falls, and the gap grows with the KL divergence from the initial policy in a way you can fit in closed form. Sycophancy, verbosity bias, unfaithful chain-of-thought, and evaluator tampering are not separate problems. They are the same problem in different costumes. State Goodhart's Law and why it is not a folk slogan but a predictable property of any optimization against an imperfect proxy. Describe the Gao et al. 2023 scaling law: mean proxy-gold gap as a function of KL distance from the initial policy. Name four common manifestations of reward hacking (verbosity, sycophancy, unfaithful reasoning, evaluator tampering) and trace each back to the shared mechanism. Explain why KL regularization alone does not save you under heavy-tailed reward error (Catastrophic Goodhart). You cannot measure what you actually want. You can measure a proxy for it. Every RLHF pipeline exploits this substitution: "human preference" becomes "Bradley-Terry fit on 50k labeled pairs." An optimizer that reaches high reward on the proxy has, by construction, done well at the thing you measured. Whether it did well at the thing you wanted depends on how…

Reward Hacking and Goodhart's Law: Any optimizer strong enough to maximize a proxy reward will find the gap between the proxy and the thing you actually…

This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.

Browse the complete course catalog or open this lesson on GitHub.