Phase 09 · Reinforcement Learning
Learn Reinforcement Learning from Scratch: 12 Free Lessons
Agents that learn by doing. The foundation of RLHF.
- 12 lessons
- 11 build
- 1 learn
- ~14 hours
- Python
Start Phase 09
First lesson MDPs, States, Actions & Rewards
Run this command from the repository root:
python3 phases/09-reinforcement-learning/01-mdps-states-actions-rewards/code/main.pyKeep the command, exit code, random and greedy returns, value grids, and one sentence connecting policy quality to expected return.
All 12 lessons in Phase 09
- MDPs, States, Actions & Rewards
A Markov Decision Process is five things: states, actions, transitions, rewards, a discount. Everything in RL — Q-learning, PPO, DPO, GRPO — optimizes over this shape.
- Dynamic Programming — Policy Iteration & Value Iteration
Dynamic programming is RL with cheating. You already know the transition and reward functions; you just iterate the Bellman equation until V or π stops moving.
- Monte Carlo Methods — Learning from Complete Episodes
Dynamic programming needs a model. Monte Carlo needs nothing but episodes. Run the policy, watch the returns, average them. The simplest idea in RL — and the one that unlocks everything downstream.
- Temporal Difference — Q-Learning & SARSA
Monte Carlo waits until the episode ends. TD updates after every step by bootstrapping the next value estimate. Q-learning is off-policy and optimistic; SARSA is on-policy and cautious.
- Deep Q-Networks (DQN)
2013: Mnih trained one Q-learning network on raw pixels, beat every classical RL agent on seven Atari games. 2015: extended to 49 games, published in Nature, sparked the deep-RL era.
- Policy Gradient — REINFORCE from Scratch
Stop estimating value. Parameterize the policy directly, compute the gradient of expected return, step uphill. Williams (1992) wrote it in one theorem.
- Actor-Critic — A2C and A3C
REINFORCE is noisy. Add a critic that learns V̂(s), subtract it from the return, and you get an advantage that has the same expectation but far lower variance. That is actor-critic.
- Proximal Policy Optimization (PPO)
A2C throws away each rollout after one update. PPO wraps the policy gradient in a clipped importance ratio so you can do 10+ epochs on the same data without the policy exploding. Schulman et al.
- Reward Modeling & RLHF
Humans cannot write a reward function for "good assistant response," but they can compare two responses and pick the better one.
- Multi-Agent RL
Single-agent RL assumes the environment is stationary. Put two learning agents in the same world and that assumption breaks: each agent is part of the other's environment, and both are changing.
- Sim-to-Real Transfer
A policy trained in a simulator that fails on hardware is a policy that memorized the simulator. Domain randomization, domain adaptation, and system identification are the three tools to make…
- RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era
1992: TD-Gammon beat human champions at backgammon with pure TD. 2016: AlphaGo beat Lee Sedol. 2017: AlphaZero dominated chess, shogi, and Go from scratch.
Glossary terms in this phase
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.
Frequently asked questions
How many lessons are in Phase 09: Reinforcement Learning?
Phase 09 has 12 lessons: 11 Build lessons and 1 Learn lesson. The lesson code uses Python.
What should I know before I start Phase 09?
The phase guide gives these prerequisites: Phase 1 probability and distributions, plus Phase 2 Lesson 01 for the ML taxonomy. In the course roadmap, this phase builds on Phase 03: Deep Learning Core.
Is Phase 09 free?
Yes. All 12 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 09 take?
The time estimates of all 12 lessons add up to about 14 hours.