Phase 18 · Ethics, Safety & Alignment
Learn AI Ethics, Safety and Alignment: 30 Free Lessons
Build AI that helps humanity. Not optional.
- 30 lessons
- 21 learn
- 9 build
- ~31 hours
- Python
Start Phase 18
First lesson Instruction-Following as Alignment Signal
Run this command from the repository root:
python3 phases/18-ethics-safety-alignment/01-instruction-following-alignment-signal/code/main.pyKeep the command, exit code, policies with and without the KL penalty, reward and KL trajectories, and one sentence naming the proxy failure you observed.
All 30 lessons in Phase 18
- Instruction-Following as Alignment Signal
Every later critique of RLHF argues against this pipeline. Before you study how optimization pressure distorts a proxy, you have to see the proxy.
- Reward Hacking and Goodhart's Law
Any optimizer strong enough to maximize a proxy reward will find the gap between the proxy and the thing you actually wanted. Gao et al.
- The Direct Preference Optimization Family
Rafailov et al. (2023) showed RLHF's optimum has a closed form in terms of the preference data, so you can skip the explicit reward model and optimize the policy directly.
- Sycophancy as RLHF Amplification
Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…
- Constitutional AI and RLAIF
Bai et al. (arXiv:2212.08073, 2022) asked: what if we replaced the human labeler with an AI that reads a list of principles?
- Mesa-Optimization and Deceptive Alignment
Hubinger et al. (arXiv:1906.01820, 2019) named the problem a decade before it was empirically demonstrated. When you train a learned optimizer to minimize a base objective, the learned optimizer's…
- Sleeper Agents — Persistent Deception
Hubinger et al. (arXiv:2401.05566, January 2024) built the first empirical model organisms of deceptive alignment. Two constructions: a code model that writes safe code when the prompt says the year…
- In-Context Scheming in Frontier Models
Meinke, Schoen, Scheurer, Balesni, Shah, Hobbhahn (Apollo Research, arXiv:2412.04984, December 2024). Tested o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, Llama 3.1 405B on agentic scenarios…
- Alignment Faking
Greenblatt, Denison, Wright, Roger et al. (Anthropic / Redwood, arXiv:2412.14093, December 2024). First demonstration that a production-grade model, without being trained to deceive and without any…
- AI Control — Safety Despite Subversion
Greenblatt, Shlegeris, Sachan, Roger (Redwood Research, arXiv:2312.06942, ICML 2024). Control reframes the safety question: given an untrusted strong model U that may be adversarially optimizing…
- Scalable Oversight and Weak-to-Strong Generalization
Burns et al. (OpenAI Superalignment, "Weak-to-Strong Generalization", 2023) proposed a proxy for the superalignment problem: fine-tune a strong model using labels produced by a weaker model.
- Red-Teaming: PAIR and Automated Attacks
Chao, Robey, Dobriban, Hassani, Pappas, Wong (NeurIPS 2023, arXiv:2310.08419). PAIR — Prompt Automatic Iterative Refinement — is the canonical automated black-box jailbreak.
- Many-Shot Jailbreaking
Anil, Durmus, Panickssery, Sharma, et al. (Anthropic, NeurIPS 2024). Many-shot jailbreaking (MSJ) exploits long context windows: stuff hundreds of faux user-assistant turns where the assistant…
- ASCII Art and Visual Jailbreaks
Jiang, Xu, Niu, Xiang, Ramasubramanian, Li, Poovendran, "ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs" (ACL 2024, arXiv:2402.11753).
- Indirect Prompt Injection — Production Attack Surface
Indirect prompt injection (IPI) embeds instructions inside external content — a web page, an email, a shared document, a support ticket — consumed by an agentic system without explicit user action.
- Red-Team Tooling — Garak, Llama Guard, PyRIT
Three production tools frame the 2026 red-team stack. Llama Guard (Meta) — a Llama-3.1-8B classifier fine-tuned on 14 MLCommons hazard categories; the 2025 Llama Guard 4 is a 12B natively multimodal…
- WMDP and Dual-Use Capability Evaluation
Li et al., "The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning" (ICML 2024, arXiv:2403.03218). 4,157 multiple-choice questions across biosecurity (1,520), cybersecurity…
- Frontier Safety Frameworks — RSP, PF, FSF
Three major-lab frameworks define the 2026 industry governance of frontier capability. Anthropic Responsible Scaling Policy v3.0 (February 2026) introduces tiered AI Safety Levels (ASL-1 through…
- Anthropic's Model Welfare Program
Anthropic, "Exploring Model Welfare" (April 2025). First major-lab formal research program on AI model welfare. Hired Kyle Fish as the first dedicated model-welfare researcher.
- Bias and Representational Harm in LLMs
Gallegos, Rossi, Barrow, Tanjim, Kim, Dernoncourt, Yu, Zhang, Ahmed (Computational Linguistics 2024, arXiv:2309.00770). Foundational 2024 survey distinguishing representational harms (stereotypes,…
- Fairness Criteria — Group, Individual, Counterfactual
Three families structure the fairness literature. Group fairness: demographic parity, equalized odds, conditional use accuracy equality — equal rates across protected groups on average.
- Differential Privacy for LLMs
DP-SGD remains the standard — noise-injected gradient updates provide formal (epsilon, delta) guarantees. Overhead in compute, memory, and utility is substantial; parameter-efficient DP fine-tuning…
- Watermarking — SynthID, Stable Signature, C2PA
Three technologies structure 2026 AI-generated-content provenance. SynthID (Google DeepMind) — image watermarking launched August 2023, text+video May 2024 (Gemini + Veo), text open-sourced October…
- Regulatory Frameworks — EU, US, UK, Korea
Four primary regulatory regimes define the 2026 AI governance landscape. EU AI Act (in force 1 August 2024) — prohibited practices and AI literacy from 2 February 2025; GPAI obligations from 2…
- EchoLeak and the Emergence of CVEs for AI
CVE-2025-32711 "EchoLeak" (CVSS 9.3) was the first publicly documented zero-click prompt injection in a production LLM system (Microsoft 365 Copilot).
- Model, System, and Dataset Cards
Three documentation formats structure AI transparency. Model Cards (Mitchell et al. 2019) — nutrition labels for models: training data, quantitative disaggregated analyses, ethical considerations,…
- Data Provenance and Training-Data Governance
EU AI Act requires machine-readable opt-out standards for GPAI by August 2025 (via EU Copyright Directive TDM exception).
- Alignment Research Ecosystem — MATS, Redwood, Apollo, METR
Five organisations define the 2026 non-lab alignment research layer. MATS (ML Alignment & Theory Scholars): 527+ researchers since late 2021, 180+ papers, 10K+ citations, h-index 47; summer 2024…
- Moderation Systems — OpenAI, Perspective, Llama Guard
Production moderation systems operationalize the safety policies defined in Lessons 12-16. OpenAI Moderation API: omni-moderation-latest (2024) built on GPT-4o classifies text + images in one call;…
- Dual-Use Risk — Cyber, Bio, Chem, Nuclear Uplift
The 2026 dual-use picture, domain by domain. Bio/chem: Lesson 17 covers WMDP; Anthropic's bioweapon-acquisition trial (2.53x uplift) and OpenAI's April 2025 Preparedness Framework v2 warning ("on…
Glossary terms in this phase
- AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- Content ProvenanceVerifiable information about the origin and editing history of a piece of media or other digital content, including the actors, tools,…
- Data ExfiltrationUnauthorized transfer of protected data from a system or trust zone to a person, tool, service, or storage location that is not permitted…
- Data ProvenanceTraceable information about where data originated, who or what transformed it, which versions were used, and how derived artifacts relate…
- Datasheet for DatasetsStructured documentation of a dataset's motivation, composition, collection process, preprocessing, uses, distribution, maintenance, and…
- DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- Indirect Prompt InjectionA prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or…
- JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
- Membership InferenceAn attack that estimates whether a particular record or example was included in a model's training data by observing model outputs or…
- Model CardA structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant…
- Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
- RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
- SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
- Threat ModelA documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and…
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
Frequently asked questions
How many lessons are in Phase 18: Ethics, Safety & Alignment?
Phase 18 has 30 lessons: 21 Learn lessons and 9 Build lessons. The lesson code uses Python.
What should I know before I start Phase 18?
The phase guide gives these prerequisites: Phase 10 Lessons 06, 07, and 08 on SFT, RLHF, and DPO. In the course roadmap, this phase builds on Phase 15: Autonomous Systems.
Is Phase 18 free?
Yes. All 30 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 18 take?
The time estimates of all 30 lessons add up to about 31 hours.
What comes after Phase 18?
Phase 19: Capstone Projects builds on this phase.