Evaluation & safety · Glossary term

What is Alignment?

The effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected and adversarial situations.

What people say

“Making AI safe.”

Why does Alignment matter?

A system can optimize the stated metric while violating the user's real intent, so alignment requires evaluation, oversight, and system controls as well as model training.

Learn Alignment in the course

Lessons that name Alignment in a title or section

  • Automated Alignment Research (Anthropic AAR)

    Anthropic ran parallel teams of Claude Opus 4.6 Autonomous Alignment Researchers in independent sandboxes, coordinating via a shared forum whose logs live outside any sandbox (so agents cannot…

    Phase 15: Autonomous Systems

  • Recursive Self-Improvement — Capability vs Alignment

    Recursive self-improvement (RSI) is no longer speculation. The ICLR 2026 RSI Workshop in Rio (April 23-27) framed it as an engineering problem with concrete tooling.

    Phase 15: Autonomous Systems

  • Instruction-Following as Alignment Signal

    Every later critique of RLHF argues against this pipeline. Before you study how optimization pressure distorts a proxy, you have to see the proxy.

    Phase 18: Ethics, Safety & Alignment

  • Mesa-Optimization and Deceptive Alignment

    Hubinger et al. (arXiv:1906.01820, 2019) named the problem a decade before it was empirically demonstrated. When you train a learned optimizer to minimize a base objective, the learned optimizer's…

    Phase 18: Ethics, Safety & Alignment

  • Alignment Faking

    Greenblatt, Denison, Wright, Roger et al. (Anthropic / Redwood, arXiv:2412.14093, December 2024). First demonstration that a production-grade model, without being trained to deceive and without any…

    Phase 18: Ethics, Safety & Alignment

  • Alignment Research Ecosystem — MATS, Redwood, Apollo, METR

    Five organisations define the 2026 non-lab alignment research layer. MATS (ML Alignment & Theory Scholars): 527+ researchers since late 2021, 180+ papers, 10K+ citations, h-index 47; summer 2024…

    Phase 18: Ethics, Safety & Alignment

  • Projection Layer for Modality Alignment

    A vision encoder produces image tokens. A text decoder consumes text tokens. The two live in different vector spaces. A small two-layer MLP projects image tokens into the text embedding space, and a…

    Phase 19: Capstone Projects

  • Vision-Language Models — The ViT-MLP-LLM Pattern

    A vision encoder converts an image into tokens. An MLP projector maps those tokens into the LLM's embedding space. A language model does the rest.

    Phase 04: Computer Vision

Covered in Phase 04: Computer Vision, Phase 12: Multimodal AI, Phase 15: Autonomous Systems, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
  • DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
  • Instruction FollowingA model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
  • Model CardA structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant…
  • RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.