Phase 14 · Agent Engineering

Build AI Agents from Scratch: 54 Free Lessons

The core of modern AI engineering. Build agents from first principles.

  • 54 lessons
  • 48 build
  • 6 learn
  • ~57 hours
  • Python

Start Phase 14

First lesson The Agent Loop: Observe, Think, Act

Run this command from the repository root:

python3 phases/14-agent-engineering/01-the-agent-loop/code/main.py

Keep the command, exit code, thought-action-observation trace, final answer, turn count, and the exact stop condition that ended the run.

All 54 lessons in Phase 14

  1. The Agent Loop: Observe, Think, Act

    Every agent in 2026 is a variant of the ReAct loop from 2022 — Claude Code, Cursor, Devin, Operator included. Reasoning tokens interleave with tool calls and observations until a stop condition fires.

    Build · Python · ~60 min

  2. ReWOO and Plan-and-Execute: Decoupled Planning

    ReAct interleaves thought and action in one stream. ReWOO separates them: one big plan up front, then execute. 5x fewer tokens, +4% accuracy on HotpotQA, and you can distill the planner into a 7B…

    Build · Python · ~60 min

  3. Reflexion: Verbal Reinforcement Learning

    Gradient-based RL needs thousands of trials and a GPU cluster to fix a failure mode. Reflexion (Shinn et al., NeurIPS 2023) does it in natural language: after each failed trial, the agent writes a…

    Build · Python · ~60 min

  4. Tree of Thoughts and LATS: Deliberate Search

    A single chain-of-thought trajectory has no room to backtrack. ToT (Yao et al., 2023) turns reasoning into a tree with self-evaluation on each node.

    Build · Python · ~75 min

  5. Self-Refine and CRITIC: Iterative Output Improvement

    Self-Refine (Madaan et al., 2023) uses one LLM in three roles — generate, feedback, refine — in a loop. Average gain: +20 absolute on 7 tasks.

    Build · Python · ~60 min

  6. Tool Use and Function Calling

    Toolformer (Schick et al., 2023) started self-supervised tool annotation. Berkeley Function Calling Leaderboard V4 (Patil et al., 2025) sets the 2026 bar: 40% agentic, 30% multi-turn, 10% live, 10%…

    Build · Python · ~60 min

  7. Agent Memory — Virtual Context and Memory Paging

    Context windows are finite. Conversations, documents, and tool traces are not. The fix is OS virtual memory restated — main context is RAM, external store is disk, the agent pages between them.

    Build · Python · ~75 min

  8. Memory Blocks and Sleep-Time Compute

    Discrete functional memory blocks the model can edit directly, and a sleep-time agent that consolidates memory asynchronously while the primary agent is idle.

    Build · Python · ~75 min

  9. Hybrid Memory: Vector + Graph + KV

    Hybrid memory runs three stores in parallel — vector for semantic similarity, KV for fast fact lookup, graph for entity-relationship reasoning — with a scoring layer that fuses them on retrieval.

    Build · Python · ~75 min

  10. Skill Libraries and Lifelong Learning (Voyager)

    Voyager (Wang et al., TMLR 2024) treats executable code as a skill. Skills are named, retrievable, composable, and refined by environment feedback.

    Build · Python · ~75 min

  11. Planning with HTN and Evolutionary Search

    Symbolic planning handles the cases where the plan is provably correct. Evolutionary code search handles the cases where the fitness function is machine-checkable.

    Build · Python · ~75 min

  12. Anthropic's Workflow Patterns: Simple Over Complex

    Schluntz and Zhang (Anthropic, Dec 2024) distinguish workflows (predefined paths) from agents (dynamic tool-use). Five workflow patterns cover most cases. Start with direct API calls.

    Build · Python · ~60 min

  13. Stateful Graph Orchestration — Durable Execution and Checkpoints

    Agent is a state machine; nodes are functions; edges are transitions; state is checkpointed after each node. Resume from any failure at the last successful checkpoint.

    Build · Python · ~75 min

  14. The Actor Model for Agents — Async Messages and Typed Runtimes

    Agents as actors: async message exchange, event-driven handlers, fault isolation, natural concurrency. AutoGen v0.4 (Microsoft Research, Jan 2025) redesigned agent orchestration around this model;…

    Build · Python · ~75 min

  15. Role-Based Agent Teams — Roles, Tasks, Processes

    Four primitives: Agent, Task, Crew, Process. Two top-level shapes: Crews (autonomous, role-based collaboration) and Flows (event-driven, deterministic).

    Build · Python · ~75 min

  16. OpenAI Agents SDK: Handoffs, Guardrails, Tracing

    OpenAI Agents SDK is the lightweight multi-agent framework built on the Responses API. Five primitives: Agent, Handoff, Guardrail, Session, Tracing. Handoffs are tools named transferto .

    Build · Python · ~75 min

  17. The Harness as a Library — Subagents and Session Store

    A harness you can import: built-in tools, subagents for context isolation, hooks, W3C trace propagation, session persistence.

    Build · Python · ~75 min

  18. Production Agent Runtimes — Fast Instantiation and Typed Workflows

    A production agent runtime optimizes what prototyping frameworks ignore: instantiation cost, typed workflow surfaces, and a serving-ready backend.

    Learn · Python · ~45 min

  19. Benchmarks: SWE-bench, GAIA, AgentBench

    Three benchmarks anchor agent evaluation in 2026. SWE-bench tests code patching. GAIA tests generalist tool use. AgentBench tests multi-environment reasoning.

    Learn · Python · ~60 min

  20. Benchmarks: WebArena and OSWorld

    WebArena tests web-agent capability across four self-hosted apps. OSWorld tests desktop-agent capability across Ubuntu, Windows, macOS.

    Learn · Python · ~60 min

  21. Computer Use: Claude, OpenAI CUA, Gemini

    Three production computer-use models in 2026. All three are vision-based. All three treat screenshots, DOM text, and tool outputs as untrusted input. Only direct user instructions count as permission.

    Build · Python · ~60 min

  22. Voice Agents: Pipecat and LiveKit

    Voice agents are a first-class production category in 2026. Pipecat gives you a Python frame-based pipeline (VAD → STT → LLM → TTS → transport). LiveKit Agents bridges AI models to users over WebRTC.

    Build · Python · ~60 min

  23. OpenTelemetry GenAI Semantic Conventions

    OpenTelemetry's GenAI SIG (launched April 2024) defines the standard schema for agent telemetry. Span names, attributes, and content-capture rules converge across vendors so agent traces mean the…

    Build · Python · ~60 min

  24. Agent Observability: Langfuse, Phoenix, Opik

    Three open-source agent observability platforms dominate 2026. Langfuse (MIT) — 6M+ installs/month, tracing + prompt management + evals + session replay.

    Learn · Python · ~45 min

  25. Multi-Agent Debate and Collaboration

    Du et al. (ICML 2024, "Society of Minds") run N model instances that independently propose answers, then iteratively critique each other over R rounds to converge.

    Build · Python · ~60 min

  26. Failure Modes: Why Agents Break

    MASFT (Berkeley, 2025) catalogs 14 multi-agent failure modes in 3 categories. Microsoft's Taxonomy documents how existing AI failures amplify in agentic settings.

    Build · Python · ~60 min

  27. Prompt Injection and the PVE Defense

    Greshake et al. (AISec 2023) established indirect prompt injection as the defining agent security problem. Attacker plants instructions in data the agent retrieves; on ingest, those instructions…

    Build · Python · ~75 min

  28. Orchestration Patterns: Supervisor, Swarm, Hierarchical

    Four orchestration patterns recur across 2026 frameworks: supervisor-worker, swarm / peer-to-peer, hierarchical, debate.

    Build · Python · ~60 min

  29. Production Runtimes: Queue, Event, Cron

    Production agents run on six runtime shapes: request-response, streaming, durable execution, queue-based background, event-driven, and scheduled. Pick the shape before you pick the framework.

    Learn · Python · ~60 min

  30. Eval-Driven Agent Development

    Anthropic's guidance: "start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when needed." Evaluation is not the last step.

    Build · Python · ~60 min

  31. Agent Workbench Engineering: Why Capable Models Still Fail

    A capable model is not enough. Reliable agents need a workbench: instructions, state, scope, feedback, verification, review, and handoff.

    Learn · Python · ~45 min

  32. The Minimal Agent Workbench

    The smallest useful workbench is three files: a root instructions router, a state file, and a task board. Everything else is layered on top. If a repo cannot carry these three, no model will save it.

    Build · Python · ~45 min

  33. Agent Instructions as Executable Constraints

    Instructions written as prose are wishes. Instructions written as constraints are tests. The workbench turns each rule into something an agent can check at runtime and a reviewer can verify after…

    Build · Python · ~50 min

  34. Repo Memory and Durable State

    Chat history is volatile. The repo is durable. The workbench stores agent state in versioned files so the next session, the next agent, and the next reviewer all read from the same source of truth.

    Build · Python · ~60 min

  35. Initialization Scripts for Agents

    Every session that starts cold pays a tax. The agent reads the same files, retries the same probes, and rediscovers the same paths. An init script pays the tax once and writes the answers into state.

    Build · Python · ~45 min

  36. Scope Contracts and Task Boundaries

    The model does not know where the work ends. A scope contract is a per-task file that says where the work begins, where it ends, and how to roll back if it spills.

    Build · Python · ~50 min

  37. Runtime Feedback Loops

    Agents that do not see real command output guess. A feedback runner captures stdout, stderr, exit code, and timing into a structured record the next turn can read.

    Build · Python · ~50 min

  38. Verification Gates

    The agent does not get to mark its own work as done. A verification gate reads the scope contract, the feedback log, the rule report, and the diff, and answers a single question: is this task…

    Build · Python · ~55 min

  39. Reviewer Agent: Separate Builder from Marker

    The agent that wrote the code cannot grade it. A reviewer is a second loop with a different system prompt, a different goal, and read-only access to everything the builder produced.

    Build · Python · ~55 min

  40. Multi-Session Handoff

    The session is going to end. The work is not. The handoff packet is the artifact that turns "the agent worked for an hour" into "the next session is productive in the first minute." Build it on…

    Build · Python · ~50 min

  41. The Workbench on a Real Repo

    Eleven lessons of surfaces are worth nothing if they do not survive contact with a real codebase. This lesson runs the same task twice on a small sample app: prompt-only versus workbench-guided.

    Build · Python · ~60 min

  42. Capstone: Ship a Reusable Agent Workbench Pack

    The mini-track ends with a pack you drop into any repo. Eleven lessons of surfaces compressed into a directory you can cp -r and have an agent working reliably the next morning.

    Build · Python · ~75 min

  43. Frame the Task Before the Agent Writes Code

    A coding agent can implement a clear task quickly. It can also implement an unclear task quickly. The speed is the same. The cost is not. Turn a request into a bounded task frame before editing.

    Build · Python · ~60 min

  44. Build an Evidence-Backed Execution Plan

    A plan is not a prettier to-do list. It is a dependency graph in which every change has a reason and every terminal node has proof. Convert a task frame into work items with evidence and proof.

    Build · Python · ~65 min

  45. Delegate Agent Work with Isolation and Merge Contracts

    Parallel agents save wall time only when the work is independent. Otherwise they convert one clear task into a coordination problem with a faster failure rate.

    Build · Python · ~70 min

  46. Turn Every Agent Correction into a System Improvement

    A correction that lives only in chat fixes one run. A correction promoted into a test, boundary, example, or tool improves every later run. Convert agent corrections into durable controls.

    Build · Python · ~65 min

  47. Define the Outcome Before You Choose the Output

    Fast implementation increases the penalty for choosing the wrong problem. Shape the outcome first so speed points in the right direction. Write an outcome frame without naming a solution.

    Build · Python · ~60 min

  48. Discover the Workflow People Actually Perform

    Requirements are not waiting in a meeting to be collected. They are scattered across actions, workarounds, records, and disagreements. Model the current workflow as ordered actions with evidence.

    Build · Python · ~70 min

  49. Map Assumptions and Resolve the Riskiest One First

    A roadmap hides uncertainty inside features. An assumption map exposes what must be true before those features deserve to exist. Convert proposed work into explicit assumptions.

    Build · Python · ~65 min

  50. Choose the Smallest Slice That Can Change the Decision

    Small is useful only when it proves something important. A tiny build that cannot change the next decision is merely incomplete. Define a slice by the assumptions it proves.

    Build · Python · ~65 min

  51. Write Specifications That Preserve Judgment

    A useful specification fixes invariants and evidence while leaving reversible implementation choices open. It is a decision boundary, not a screenplay.

    Build · Python · ~75 min

  52. Design Success Metrics Before the Result Exists

    Measurement should answer a decision, not decorate a dashboard. Start with the goal, derive questions, then choose the smallest metrics that answer them.

    Build · Python · ~70 min

  53. Choose Prototype, Pilot, or Production Deliberately

    These are different learning environments, not levels of polish. Choose the stage that answers the current unknown with the least unnecessary consequence.

    Build · Python · ~70 min

  54. Build a Feedback Ratchet with Ownership and Retirement

    Shipping closes one build loop and opens the learning loop. Evidence must change the system or it becomes telemetry nobody owns.

    Build · Python · ~75 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • Agent HarnessThe runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the…
  • Agent MemoryInformation stored outside the model and selected for use in later agent steps, such as prior decisions, user preferences, task episodes,…
  • Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
  • Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • Durable ExecutionRunning a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • Function CallingA provider or application interface through which a model emits a structured request naming a tool and its arguments.
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • HallucinationGenerated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth.
  • HandoffA structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints,…
  • Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • LLM-as-a-JudgeUsing a language model to score, compare, classify, or critique another system's output against a rubric.
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • OrchestrationThe control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
  • PatchA reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
  • PlanningConstructing, selecting, or revising a sequence of actions and dependencies intended to move from the current state to a goal.
  • Progressive DisclosureSupplying a person or model with the minimum useful context first, then revealing deeper detail when the task or evidence requires it.
  • Prompt EngineeringDesigning model-facing instructions, examples, constraints, and output requirements to improve behavior on a defined task.
  • Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
  • ReActAn agent pattern that interleaves task reasoning, a concrete action, and an observation returned by the environment before deciding the…
  • Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
  • Repository MapA compact, maintained description of a repository's important directories, ownership boundaries, entry points, build commands, tests,…
  • Reviewer AgentAn agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
  • RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
  • SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
  • Scope ContractA concrete agreement that defines a task's goal, allowed and forbidden surfaces, expected artifacts, verification requirements, and…
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • SwarmA loosely coordinated multi-agent pattern in which local agent decisions and message exchange produce system-level behavior.
  • System PromptA provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that…
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • Tool ContractThe complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts,…
  • TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
  • Verification GateA control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
  • WorktreeIn Git, a working directory attached to a repository and branch or commit, with shared object storage but its own checked-out files and…

Frequently asked questions

How many lessons are in Phase 14: Agent Engineering?

Phase 14 has 54 lessons: 48 Build lessons and 6 Learn lessons. The lesson code uses Python.

What should I know before I start Phase 14?

The phase guide gives these prerequisites: Phase 11 LLM Engineering and Phase 13 Tools and Protocols. In the course roadmap, this phase builds on Phase 13: Tools & Protocols.

Is Phase 14 free?

Yes. All 54 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 14 take?

The time estimates of all 54 lessons add up to about 57 hours.

What comes after Phase 14?

Phase 15: Autonomous Systems and Phase 17: Infrastructure & Production build on this phase.