Phase 13 · Tools & Protocols
Learn Tool Calling and MCP: 31 Free Lessons
The interfaces between AI and the real world.
- 31 lessons
- 24 build
- 7 learn
- ~43 hours
- Python, TypeScript
This phase moves from function calls and tool schemas into interoperable protocols, Agent Skills, security, and production governance. Numeric order is useful for browsing.
Start Phase 13
First full-phase lesson The Tool Interface — Why Agents Need Structured I/O
Run this command from the repository root:
python3 phases/13-tools-and-protocols/01-the-tool-interface/code/main.pyKeep the command, exit code, describe-decide-execute-observe trace, rejected input evidence, and one sentence explaining the turn limit.
All 31 lessons in Phase 13
- The Tool Interface — Why Agents Need Structured I/O
A language model produces tokens. A program takes actions. The gap between those two is the tool interface: a contract that lets the model request an action and the host execute it.
- Function Calling Deep Dive — OpenAI, Anthropic, Gemini
The three frontier providers converged on the same tool-call loop in 2024 and then diverged on everything else. OpenAI uses tools and toolcalls. Anthropic uses tooluse and toolresult blocks.
- Parallel Tool Calls and Streaming with Tools
Three independent weather lookups serialized is three round trips. Run them in parallel and total time collapses to the slowest single call.
- Structured Output — JSON Schema, Pydantic, Zod, Constrained Decoding
"Ask the model nicely to return JSON" fails 5 to 15 percent of the time, even on frontier models. Structured outputs close that gap with constrained decoding: the model is literally prevented from…
- Tool Schema Design — Naming, Descriptions, Parameter Constraints
A correct tool fails silently when the model cannot tell when to use it. Naming, descriptions, and parameter shapes drive 10 to 20 percentage-point swings in tool-selection accuracy on benchmarks…
- MCP Fundamentals: Stateless Requests and JSON-RPC
Modern MCP has no handshake and no protocol session. Each request must carry enough metadata to be understood, authorized, routed, and retried on its own.
- Building an MCP Server: Stateless Python and TypeScript
A modern MCP server does not remember a handshake. It validates the metadata on every request, runs one handler, and returns one typed result. Implement mandatory server/discover for MCP 2026-07-28.
- Building an MCP Client: Discovery, Routing, and Dual-Era Fallback
A modern MCP client repeats its contract on every request. Its hardest compatibility decision is knowing when an old server is truly old and when a modern server is reporting a correctable error.
- MCP Transports: stdio and Stateless Streamable HTTP
Transport carries MCP messages. It does not supply missing protocol state. In 2026-07-28, local stdio and remote Streamable HTTP both carry self-describing requests.
- MCP Resources and Prompts: Addressable Context for Stateless Servers
Tools perform operations. Resources expose addressable content. Prompts package user-selected message templates. A good MCP server keeps those contracts separate and predictable.
- MCP Model Input: Sampling Migration and Stateless MRTR
MCP 2026-07-28 deprecates Sampling for new designs and removes the server-to-client request channel. If an existing workflow still needs the client's model, the server returns an inputrequired…
- Explicit Scope and Stateless Elicitation
Roots are deprecated in MCP 2026-07-28 and were never a security sandbox. Put scope in visible tool arguments or resource URIs, authorize it on the server, and use MRTR when a tool genuinely needs…
- MCP Tasks Extension: Durable Work on a Stateless Core
Stateless MCP does not mean every operation must finish in one request. The official Tasks extension gives long-running work an explicit durable handle.
- MCP Apps on the Stateless Protocol
An interactive result is still an MCP tool and resource exchange. The 2026-07-28 core makes that exchange self-contained, while the Apps extension adds the sandboxed browser surface.
- MCP Security: Poisoned Metadata, Routing, and MRTR State
Stateless does not mean trustless. It means every request exposes the evidence a server and gateway need to validate the call independently.
- MCP Authorization: CIMD, Issuer Binding, PKCE, and Step-Up
A remote MCP request is stateless, but its authorization is not anonymous. Bind every credential to the issuer that created it and every token to the resource that receives it.
- Stateless MCP Gateways and Registry Admission
A gateway should make every route explicit. The 2026-07-28 protocol gives it method, name, version, capability, identity, cache, and trace boundaries without a transport session.
- MCP Auth in Production: Issuer-Bound Enrollment and Tokens
Lesson 16 built the OAuth 2.1 state machine. This lesson hardens its production boundaries for MCP 2026-07-28: Client ID Metadata Documents first, deprecated dynamic registration only for…
- A2A — Agent-to-Agent Protocol
MCP is agent-to-tool. A2A (Agent2Agent) is agent-to-agent — an open protocol for letting opaque agents built on different frameworks collaborate.
- OpenTelemetry GenAI — Tracing Tool Calls End-to-End
An agent calls five tools, three MCP servers, and two sub-agents. You need one trace across all of it. The OpenTelemetry GenAI semantic conventions (stable attributes in v1.37 and up) are the 2026…
- LLM Routing Layer — LiteLLM, OpenRouter, Portkey
Provider lock-in is expensive. Different tool-calling workloads suit different models. Routing gateways give one API surface, retries, failover, cost tracking, and guardrails.
- Agent Skills: Portable Contract and Runtime Boundary
A skill is not a long prompt with a better filename. It is a discoverable package of instructions, resources, and executable helpers that enters an agent's context through a runtime contract.
- Capstone: Stateless Tool Ecosystem
A production agent system is a set of boundaries, not a pile of features. This capstone separates a readable in-process simulation from the protocol clients, authorization server, sandbox, and…
- Skill Discovery and Progressive Disclosure
A skill becomes useful before its body is loaded. Its name and description earn a place in the catalog; its deeper files earn context only when the task reaches them.
- Skill Invocation and Routing
Invocation is an authority decision followed by a relevance decision. A good description helps the model choose; a good policy decides whether that choice is allowed.
- Skill Permissions, Sandboxes, and Trust
A skill can suggest an action. Only the host can authorize it, only an isolation boundary can contain it, and only verification can tell you whether it worked.
- Skill Evals, Packaging, and Portability
A skill is finished when its package survives linting, routes on the right requests, improves a measured task, stays inside policy, and degrades honestly on another host.
- MCP Tool Contracts and Content
A tool is safe to automate only when discovery, arguments, results, pagination, and transport metadata agree on one contract. Define tool inputs and outputs with JSON Schema 2020-12.
- MCP Reliability, Cancellation, and Flow Control
A request ID correlates a message. It does not make a side effect safe, stop a worker, or protect a stream from a slow consumer.
- MCP Registry Supply Chain: Admission, Drift, and Rollback
A registry entry tells you what a publisher declared. Production admission proves what you fetched, what you observed, what you approved, and what you can safely restore.
- MCP Conformance Engineering: Versioning, Evidence, and Operations
A server is not conformant because its happy path worked through one SDK. Conformance lives at the wire, at version boundaries, through intermediaries, and during rollback.
Glossary terms in this phase
- AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
- Agent SkillA discoverable directory of procedural instructions whose entry point is `SKILL.md`, with optional references, scripts, and assets that a…
- Circuit BreakerA reliability control that temporarily stops calls to a dependency after failures cross a threshold, then probes whether the dependency…
- Coding AgentAn agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- Function CallingA provider or application interface through which a model emits a structured request naming a tool and its arguments.
- GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- MCP (Model Context Protocol)An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request,…
- Multi Round-Trip Request (MRTR)An MCP request pattern in which an operation returns `resultType: input_required` with one or more `inputRequests`, then the client…
- OrchestrationThe control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- Progressive DisclosureSupplying a person or model with the minimum useful context first, then revealing deeper detail when the task or evidence requires it.
- Rate LimitA policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.
- Repository InstructionsVersion-controlled guidance that tells coding agents how a repository is organized, which commands and conventions apply, what boundaries…
- RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
- SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
- Skill BundleThe complete installable skill directory, including `SKILL.md` and every reference, script, asset, fixture, or companion file required by…
- Skill CatalogThe compact model-visible inventory of eligible skills, usually containing routing metadata such as name, description, and an internal…
- Skill DiscoveryA runtime pipeline that searches configured roots, identifies candidate skill directories, validates their package contract, attaches…
- Skill InvocationThe runtime-mediated process in which an eligible human, model, application, or other skill selects a skill and causes its instructions to…
- Stateless MCPThe MCP 2026-07-28 request model in which every request carries the protocol version and client capabilities in `params._meta`, while…
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- Structured OutputModel output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Tool ContractThe complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts,…
- TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
- Trust BoundaryAn interface where data, instructions, identity, or authority crosses between components or principals that operate under different trust…
Frequently asked questions
How many lessons are in Phase 13: Tools & Protocols?
Phase 13 has 31 lessons: 24 Build lessons and 7 Learn lessons. The lesson code uses Python and TypeScript.
What should I know before I start Phase 13?
The phase guide gives these prerequisites: Phase 11 LLM completion APIs. In the course roadmap, this phase builds on Phase 11: LLM Engineering.
Is Phase 13 free?
Yes. All 31 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 13 take?
The time estimates of all 31 lessons add up to about 43 hours.
What comes after Phase 13?
Phase 14: Agent Engineering builds on this phase.