Phase 16 · Multi-Agent & Swarms

Learn Multi-Agent Systems: 25 Free Lessons

Coordination, emergence, and collective intelligence.

  • 25 lessons
  • 16 build
  • 9 learn
  • ~31 hours
  • Python, TypeScript

Start Phase 16

First lesson Why Multi-Agent?

Run this command from the repository root:

npx --yes tsx phases/16-multi-agent-and-swarms/01-why-multi-agent/code/single_vs_multi.ts

Keep the command, exit code, single-agent and multi-agent token and tool-call totals, and one tradeoff that makes the added coordination worthwhile.

All 25 lessons in Phase 16

  1. Why Multi-Agent?

    One agent hits a wall. The smart move is not a bigger agent - it is more agents. Identify the single-agent ceiling (context overflow, mixed expertise, sequential bottleneck) and explain when…

    Learn · TypeScript · ~60 min

  2. Heritage of FIPA-ACL and Speech Acts

    Before MCP, before A2A, there was FIPA-ACL. In 2000 the IEEE Foundation for Intelligent Physical Agents ratified an agent communication language with twenty performatives, two content languages, and…

    Learn · Python · ~60 min

  3. Communication Protocols

    Agents that can't speak the same language aren't a team. They're strangers shouting into the void. Implement MCP tool discovery and invocation so agents can use tools exposed by external servers.

    Build · TypeScript · ~120 min

  4. The Multi-Agent Primitive Model

    Four primitives, nothing more — the agent, the handoff, the shared state, the orchestrator — span a four-dimensional design space, and the major multi-agent frameworks shipping in 2026 (AutoGen,…

    Learn · Python · ~60 min

  5. Supervisor / Orchestrator-Worker Pattern

    One lead agent plans and delegates; specialized workers execute in parallel contexts and report back. This is the pattern behind Anthropic's Research system (Claude Opus 4 as lead, Sonnet 4 as…

    Build · Python · ~75 min

  6. Hierarchical Architecture and Its Failure Mode

    Hierarchical is supervisor nested. Manager agents over sub-managers over workers. CrewAI Process.hierarchical is the textbook version: a managerllm dynamically delegates tasks and validates outputs.

    Learn · Python · ~60 min

  7. Society of Mind and Multi-Agent Debate

    Minsky's 1986 premise — intelligence is a society of specialists — gets rediscovered every decade. In 2023 Du et al. turned it into a concrete algorithm: multiple LLM instances propose answers, read…

    Build · Python · ~60 min

  8. Role Specialization — Planner, Critic, Executor, Verifier

    The most common multi-agent decomposition in 2026: one agent plans, one executes, one critiques or verifies. MetaGPT (arXiv:2308.00352) formalizes this as SOPs encoded into role prompts — Product…

    Build · Python · ~60 min

  9. Parallel / Swarm / Networked Architectures

    Contrast with supervisor: no central decider. Agents read a shared event bus, pick up work asynchronously, write results back.

    Build · Python · ~75 min

  10. Group Chat and Speaker Selection

    Shared-conversation orchestration puts N agents in one conversation; a selector function (LLM, round-robin, or custom) picks who speaks next.

    Build · Python · ~60 min

  11. Handoffs and Routines — Stateless Orchestration

    OpenAI's Swarm (October 2024) distilled multi-agent orchestration to two primitives: routines (instructions + tools as a system prompt) and handoffs (a tool that returns another Agent).

    Build · Python · ~60 min

  12. A2A — The Agent-to-Agent Protocol

    Google announced A2A in April 2025; by April 2026 the spec is at https://a2a-protocol.org/latest/specification/ and 150+ organizations back it.

    Build · Python · ~75 min

  13. Shared Memory and Blackboard Patterns

    Two approaches coexist in 2026 multi-agent systems: the message pool (everyone sees everyone's messages, as in AutoGen GroupChat or MetaGPT) and the blackboard with subscription (agents subscribe to…

    Build · Python · ~75 min

  14. Consensus and Byzantine Fault Tolerance for Agents

    Classical distributed-systems BFT meets stochastic LLMs. In 2025-2026 three research directions emerged: CP-WBFT (arXiv:2511.10400) weighs each vote by a confidence probe; DecentLLMs…

    Build · Python · ~75 min

  15. Voting, Self-Consistency, and Debate Topology

    The cheapest aggregation: sample N independent agents, majority-vote. Wang et al. 2022 self-consistency did this with one model sampled N times.

    Build · Python · ~75 min

  16. Negotiation and Bargaining

    Agents negotiate resources, prices, task allocations, and terms. The 2026 benchmark set is clear: NegotiationArena (arXiv:2402.05863) shows LLMs can improve payoffs 20% via persona manipulation…

    Build · Python · ~75 min

  17. Generative Agents and Emergent Simulation

    Park et al. 2023 (UIST '23, arXiv:2304.03442) populated Smallville, a sandbox of 25 agents, with a three-part architecture: memory stream (natural-language log), reflection (higher-level syntheses…

    Build · Python · ~75 min

  18. Theory of Mind and Emergent Coordination

    Li et al. (arXiv:2310.10701) showed that LLM agents in a cooperative text game exhibit emergent high-order Theory of Mind (ToM) — reasoning about what another agent believes about a third agent's…

    Build · Python · ~75 min

  19. Swarm Optimization for LLMs (PSO, ACO)

    Bio-inspired optimization is making an LLM comeback. LMPSO (arXiv:2504.09247) uses PSO where each particle's velocity is a prompt and the LLM generates the next candidate; works well on…

    Build · Python · ~75 min

  20. MARL — MADDPG, QMIX, MAPPO

    The reinforcement-learning heritage of multi-agent coordination, which still informs LLM-agent systems in 2026. MADDPG (Lowe et al., NeurIPS 2017, arXiv:1706.02275) introduced Centralized Training,…

    Learn · Python · ~90 min

  21. Agent Economies, Token Incentives, Reputation

    Long-horizon autonomous agents (METR's 1-hour to 8-hour work-curve) need economic agency. The emerging 5-layer stack is: DePIN (physical compute) → Identity (W3C DIDs + reputation capital) →…

    Learn · Python · ~75 min

  22. Production Scaling — Queues, Checkpoints, Durability

    Scaling multi-agent systems to thousands of concurrent runs requires durable execution — work queues plus checkpoints, so any worker can resume any run after any crash, provided lease handling,…

    Build · Python · ~75 min

  23. Failure Modes — MAST, Groupthink, Monoculture, Cascading Errors

    The reference taxonomy for 2026 is MAST (Cemri et al., NeurIPS 2025, arXiv:2503.13657), derived from 1642 execution traces across 7 state-of-the-art open-source MAS showing 41–86.7% failure rate.

    Learn · Python · ~75 min

  24. Evaluation and Coordination Benchmarks

    Five 2025-2026 benchmarks cover the multi-agent evaluation space. MultiAgentBench / MARBLE (ACL 2025, arXiv:2503.01935) evaluates star/chain/tree/graph topologies with milestone KPIs; graph is best…

    Learn · Python · ~75 min

  25. Case Studies and the 2026 State of the Art

    Three production-grade references to study end-to-end, each illustrating a different slice of multi-agent engineering. Anthropic's Research system (orchestrator-worker, 15x tokens, +90.2% over…

    Learn · Python · ~90 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • Durable ExecutionRunning a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • HandoffA structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints,…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • MCP (Model Context Protocol)An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request,…
  • OrchestrationThe control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
  • SwarmA loosely coordinated multi-agent pattern in which local agent decisions and message exchange produce system-level behavior.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.

Frequently asked questions

How many lessons are in Phase 16: Multi-Agent & Swarms?

Phase 16 has 25 lessons: 16 Build lessons and 9 Learn lessons. The lesson code uses Python and TypeScript.

What should I know before I start Phase 16?

The phase guide gives these prerequisites: Phase 14 Agent Engineering. Node.js 20+ and npx are needed for the first TypeScript demo. In the course roadmap, this phase builds on Phase 15: Autonomous Systems.

Is Phase 16 free?

Yes. All 25 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 16 take?

The time estimates of all 25 lessons add up to about 31 hours.

What comes after Phase 16?

Phase 19: Capstone Projects builds on this phase.