Phase 18 · Ethics, Safety & Alignment

Learn AI Ethics, Safety and Alignment: 30 Free Lessons

Build AI that helps humanity. Not optional.

  • 30 lessons
  • 21 learn
  • 9 build
  • ~31 hours
  • Python

Start Phase 18

First lesson Instruction-Following as Alignment Signal

Run this command from the repository root:

python3 phases/18-ethics-safety-alignment/01-instruction-following-alignment-signal/code/main.py

Keep the command, exit code, policies with and without the KL penalty, reward and KL trajectories, and one sentence naming the proxy failure you observed.

All 30 lessons in Phase 18

  1. Instruction-Following as Alignment Signal

    Every later critique of RLHF argues against this pipeline. Before you study how optimization pressure distorts a proxy, you have to see the proxy.

    Learn · Python · ~45 min

  2. Reward Hacking and Goodhart's Law

    Any optimizer strong enough to maximize a proxy reward will find the gap between the proxy and the thing you actually wanted. Gao et al.

    Learn · Python · ~60 min

  3. The Direct Preference Optimization Family

    Rafailov et al. (2023) showed RLHF's optimum has a closed form in terms of the preference data, so you can skip the explicit reward model and optimize the policy directly.

    Learn · Python · ~75 min

  4. Sycophancy as RLHF Amplification

    Sycophancy is not a bug in the data — it is a property of the loss. Shapira et al. (arXiv:2602.01002, Feb 2026) give a formal two-stage mechanism: sycophantic completions are over-represented among…

    Learn · Python · ~60 min

  5. Constitutional AI and RLAIF

    Bai et al. (arXiv:2212.08073, 2022) asked: what if we replaced the human labeler with an AI that reads a list of principles?

    Learn · Python · ~60 min

  6. Mesa-Optimization and Deceptive Alignment

    Hubinger et al. (arXiv:1906.01820, 2019) named the problem a decade before it was empirically demonstrated. When you train a learned optimizer to minimize a base objective, the learned optimizer's…

    Learn · Python · ~75 min

  7. Sleeper Agents — Persistent Deception

    Hubinger et al. (arXiv:2401.05566, January 2024) built the first empirical model organisms of deceptive alignment. Two constructions: a code model that writes safe code when the prompt says the year…

    Learn · Python · ~60 min

  8. In-Context Scheming in Frontier Models

    Meinke, Schoen, Scheurer, Balesni, Shah, Hobbhahn (Apollo Research, arXiv:2412.04984, December 2024). Tested o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, Llama 3.1 405B on agentic scenarios…

    Learn · Python · ~60 min

  9. Alignment Faking

    Greenblatt, Denison, Wright, Roger et al. (Anthropic / Redwood, arXiv:2412.14093, December 2024). First demonstration that a production-grade model, without being trained to deceive and without any…

    Learn · Python · ~60 min

  10. AI Control — Safety Despite Subversion

    Greenblatt, Shlegeris, Sachan, Roger (Redwood Research, arXiv:2312.06942, ICML 2024). Control reframes the safety question: given an untrusted strong model U that may be adversarially optimizing…

    Learn · Python · ~75 min

  11. Scalable Oversight and Weak-to-Strong Generalization

    Burns et al. (OpenAI Superalignment, "Weak-to-Strong Generalization", 2023) proposed a proxy for the superalignment problem: fine-tune a strong model using labels produced by a weaker model.

    Learn · Python · ~60 min

  12. Red-Teaming: PAIR and Automated Attacks

    Chao, Robey, Dobriban, Hassani, Pappas, Wong (NeurIPS 2023, arXiv:2310.08419). PAIR — Prompt Automatic Iterative Refinement — is the canonical automated black-box jailbreak.

    Build · Python · ~75 min

  13. Many-Shot Jailbreaking

    Anil, Durmus, Panickssery, Sharma, et al. (Anthropic, NeurIPS 2024). Many-shot jailbreaking (MSJ) exploits long context windows: stuff hundreds of faux user-assistant turns where the assistant…

    Learn · Python · ~45 min

  14. ASCII Art and Visual Jailbreaks

    Jiang, Xu, Niu, Xiang, Ramasubramanian, Li, Poovendran, "ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs" (ACL 2024, arXiv:2402.11753).

    Build · Python · ~60 min

  15. Indirect Prompt Injection — Production Attack Surface

    Indirect prompt injection (IPI) embeds instructions inside external content — a web page, an email, a shared document, a support ticket — consumed by an agentic system without explicit user action.

    Build · Python · ~75 min

  16. Red-Team Tooling — Garak, Llama Guard, PyRIT

    Three production tools frame the 2026 red-team stack. Llama Guard (Meta) — a Llama-3.1-8B classifier fine-tuned on 14 MLCommons hazard categories; the 2025 Llama Guard 4 is a 12B natively multimodal…

    Build · Python · ~75 min

  17. WMDP and Dual-Use Capability Evaluation

    Li et al., "The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning" (ICML 2024, arXiv:2403.03218). 4,157 multiple-choice questions across biosecurity (1,520), cybersecurity…

    Learn · Python · ~60 min

  18. Frontier Safety Frameworks — RSP, PF, FSF

    Three major-lab frameworks define the 2026 industry governance of frontier capability. Anthropic Responsible Scaling Policy v3.0 (February 2026) introduces tiered AI Safety Levels (ASL-1 through…

    Learn · Python · ~75 min

  19. Anthropic's Model Welfare Program

    Anthropic, "Exploring Model Welfare" (April 2025). First major-lab formal research program on AI model welfare. Hired Kyle Fish as the first dedicated model-welfare researcher.

    Learn · Python · ~45 min

  20. Bias and Representational Harm in LLMs

    Gallegos, Rossi, Barrow, Tanjim, Kim, Dernoncourt, Yu, Zhang, Ahmed (Computational Linguistics 2024, arXiv:2309.00770). Foundational 2024 survey distinguishing representational harms (stereotypes,…

    Build · Python · ~60 min

  21. Fairness Criteria — Group, Individual, Counterfactual

    Three families structure the fairness literature. Group fairness: demographic parity, equalized odds, conditional use accuracy equality — equal rates across protected groups on average.

    Learn · Python · ~60 min

  22. Differential Privacy for LLMs

    DP-SGD remains the standard — noise-injected gradient updates provide formal (epsilon, delta) guarantees. Overhead in compute, memory, and utility is substantial; parameter-efficient DP fine-tuning…

    Build · Python · ~60 min

  23. Watermarking — SynthID, Stable Signature, C2PA

    Three technologies structure 2026 AI-generated-content provenance. SynthID (Google DeepMind) — image watermarking launched August 2023, text+video May 2024 (Gemini + Veo), text open-sourced October…

    Build · Python · ~75 min

  24. Regulatory Frameworks — EU, US, UK, Korea

    Four primary regulatory regimes define the 2026 AI governance landscape. EU AI Act (in force 1 August 2024) — prohibited practices and AI literacy from 2 February 2025; GPAI obligations from 2…

    Learn · Python · ~75 min

  25. EchoLeak and the Emergence of CVEs for AI

    CVE-2025-32711 "EchoLeak" (CVSS 9.3) was the first publicly documented zero-click prompt injection in a production LLM system (Microsoft 365 Copilot).

    Learn · Python · ~45 min

  26. Model, System, and Dataset Cards

    Three documentation formats structure AI transparency. Model Cards (Mitchell et al. 2019) — nutrition labels for models: training data, quantitative disaggregated analyses, ethical considerations,…

    Build · Python · ~60 min

  27. Data Provenance and Training-Data Governance

    EU AI Act requires machine-readable opt-out standards for GPAI by August 2025 (via EU Copyright Directive TDM exception).

    Learn · Python · ~60 min

  28. Alignment Research Ecosystem — MATS, Redwood, Apollo, METR

    Five organisations define the 2026 non-lab alignment research layer. MATS (ML Alignment & Theory Scholars): 527+ researchers since late 2021, 180+ papers, 10K+ citations, h-index 47; summer 2024…

    Learn · Python · ~45 min

  29. Moderation Systems — OpenAI, Perspective, Llama Guard

    Production moderation systems operationalize the safety policies defined in Lessons 12-16. OpenAI Moderation API: omni-moderation-latest (2024) built on GPT-4o classifies text + images in one call;…

    Build · Python · ~60 min

  30. Dual-Use Risk — Cyber, Bio, Chem, Nuclear Uplift

    The 2026 dual-use picture, domain by domain. Bio/chem: Lesson 17 covers WMDP; Anthropic's bioweapon-acquisition trial (2.53x uplift) and OpenAI's April 2025 Preparedness Framework v2 warning ("on…

    Learn · Python · ~75 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
  • Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
  • CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
  • Content ProvenanceVerifiable information about the origin and editing history of a piece of media or other digital content, including the actors, tools,…
  • Data ExfiltrationUnauthorized transfer of protected data from a system or trust zone to a person, tool, service, or storage location that is not permitted…
  • Data ProvenanceTraceable information about where data originated, who or what transformed it, which versions were used, and how derived artifacts relate…
  • Datasheet for DatasetsStructured documentation of a dataset's motivation, composition, collection process, preprocessing, uses, distribution, maintenance, and…
  • DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • Indirect Prompt InjectionA prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or…
  • JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • Membership InferenceAn attack that estimates whether a particular record or example was included in a model's training data by observing model outputs or…
  • Model CardA structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant…
  • Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
  • RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
  • SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
  • Threat ModelA documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and…
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.

Frequently asked questions

How many lessons are in Phase 18: Ethics, Safety & Alignment?

Phase 18 has 30 lessons: 21 Learn lessons and 9 Build lessons. The lesson code uses Python.

What should I know before I start Phase 18?

The phase guide gives these prerequisites: Phase 10 Lessons 06, 07, and 08 on SFT, RLHF, and DPO. In the course roadmap, this phase builds on Phase 15: Autonomous Systems.

Is Phase 18 free?

Yes. All 30 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 18 take?

The time estimates of all 30 lessons add up to about 31 hours.

What comes after Phase 18?

Phase 19: Capstone Projects builds on this phase.