MCP Servers Can Now Ship Their Own Manual

Some MCP servers have a real workflow behind them. If you ship one, you can now move the long how-to text out of instructions and into a skill. A skill is a directory that the host loads only when a task needs it. The host checks it byte for byte and binds the user's approval to it. If you build a host, you must give that skill a registry keyed by server and URI. You must read its files on demand and ask for approval again when a file changes. This essay covers the mechanism and the code for both sides.

On September 13 the MCP maintainers merged SEP-2640, the Skills Extension, as Final. With it, a server gives an agent two things over one connection. It gives the tools and the skill that explains how to use them. It adds little on the wire: one URI convention, two required methods, and one optional method. Most of the document is about a different question. What should a host believe about instructions that a remote server wrote?


The gap it closes

In the Agent Skills specification, a skill is a directory with a SKILL.md at its root. That file has YAML frontmatter with a name and a description, then a Markdown body of instructions. The directory can also hold scripts/, references/, and assets/. Hosts load a skill in stages. The name and description, about 100 tokens, sit in context from startup. The body loads when the agent decides that the skill applies. The spec recommends a body under 5,000 tokens and 500 lines. Supporting files load only when a step needs them.

MCP had no standard way to ship any of that. The SEP's motivation section lists three problems. First, distribution was split. A server and the skill that teaches an agent to use it had separate versions, discovery, and installs. A user who installed a server from a registry did not know that a companion skill existed. Second, the only built-in channel for guidance is the instructions field from server/discover. In practice, that field has a size limit. The SEP's example is an 875-line skill that does not fit there. Third, several projects had already made their own skill:// URIs. Each project had different rules for the host, the path, and the sub-resources.

The post-stateless roadmap described the catalog tax. Every tool schema that a server exposes costs context, even if the agent never uses it. Skills reduce the other half of that cost. A tool schema says what a function accepts. A skill says when to call the function, in what order, and what to do when it fails. The host does not pay for all of that text at the start.

Resources, not a new primitive

The working group had a competing design on the table, SEP-2076, which made skills a first-class protocol primitive. The rationale document explains why it lost. A new primitive would flatten skills to "name-addressed blobs." That design removes the directory model that progressive disclosure needs. The existing Resources primitive keeps the directory. It also brings URI addressing, resources/read, subscriptions, and every client that already browses resources.

Each file in a skill directory becomes one resource. By convention, its URI is skill://<skill-path>/<file-path>. The rules are precise:

The skill is the same directory a filesystem host already loads. The Agent Skills specification sets the frontmatter rules:

The code in the rest of this essay runs against the skill below, served at skill://acme/billing/refunds/SKILL.md:

refunds/
  SKILL.md
  examples/email.md
---
name: refunds
description: Process customer refund requests per company policy. Use when a ticket asks for a refund, a chargeback, or a partial credit.
license: Apache-2.0
---

1. Look up the order with `get_order`. Refuse refunds older than 90 days.
2. Under $200, call `issue_refund`. Otherwise draft the email in
   `examples/email.md` and stop for a human.

A sample skill written for this essay. The spec's reference library validates the frontmatter and naming before you publish: skills-ref validate ./refunds (Agent Skills, Validation).

Three methods, and a manifest in every listing

A server opts in through the extension negotiation that SEP-2133 defines. It declares io.modelcontextprotocol/skills in the extensions field of its capabilities. It must also declare the ordinary resources capability. The extension requires the server to implement two methods.

skills/list enumerates the skills the server serves. Each entry is a full manifest. It holds the skill's frontmatter, copied verbatim as JSON. It also holds a resources array with every file in the skill and its SHA-256 digest and byte size. One paged pass gives the host its whole registry. From that pass, the host shows the user what they approve and later verifies every file that it reads. On protocol 2026-07-28 and later, the result also carries the base protocol's ttlMs and cacheScope list-caching attributes. The listing may be empty or partial. Some servers cannot list everything. Examples are a documentation server that makes one skill per API endpoint, and a gateway in front of an external index. Hosts must not treat an empty list as proof that no skills exist.

skills/get takes one SKILL.md URI and returns the same entry shape. It has two uses. First, it refreshes the digests of one skill after a mismatch, without a new listing of the catalog. Second, it verifies a skill whose URI came by a different route. That route can be the server's instructions, another skill, or the user. An unknown URI gets error -32602, which is the same code that resources/read uses.

The third method is optional. resources/directory/read lists the direct children of a directory resource, for instructions like "pick the right template from templates/". A directoryRead: true setting on the capability enables it. Without that setting, clients must not call it. All other reads use resources/read. The SEP also sets two limits per skill that every conforming host must accept: 512 files and 16 MiB in total. A host can check both limits from the manifest before it fetches any content.

On the server, you opt in with one object inside the capabilities that you already return. Below is the SEP's example with the required resources capability:

{
  "capabilities": {
    "resources": {},
    "extensions": {
      "io.modelcontextprotocol/skills": { "directoryRead": true }
    }
  }
}

No SDK builds the manifest for you. The Python reference implementation, python-sdk pull request #3485, says so. Its Skills extension "does not discover, read, or hash skills from a filesystem." A server that serves skills from a folder needs a build step. The step walks the directory and checks the name against it. It hashes every file and enforces the two limits before a host sees the entry:

# /// script
# dependencies = ["pyyaml"]
# ///
import hashlib
import json
import re
import sys
from pathlib import Path

import yaml

NAME = re.compile(r"^[a-z0-9]+(-[a-z0-9]+)*$")
MAX_FILES, MAX_BYTES = 512, 16 * 1024 * 1024


def skill_entry(skill_dir: Path, prefix: str = "") -> dict:
    skill_md = (skill_dir / "SKILL.md").read_bytes()
    head = skill_md.decode("utf-8").split("---", 2)
    frontmatter = yaml.safe_load(head[1])
    name = frontmatter["name"]
    if not NAME.match(name) or len(name) > 64 or name != skill_dir.name:
        raise ValueError(f"name {name!r} must match the directory and the Agent Skills rules")
    base = f"skill://{prefix + '/' if prefix else ''}{name}"
    files = sorted(p for p in skill_dir.rglob("*") if p.is_file())
    resources = []
    for p in files:
        data = p.read_bytes()
        resources.append({
            "uri": f"{base}/{p.relative_to(skill_dir).as_posix()}",
            "digest": "sha256:" + hashlib.sha256(data).hexdigest(),
            "size": len(data),
        })
    total = sum(r["size"] for r in resources)
    if len(resources) > MAX_FILES or total > MAX_BYTES:
        raise ValueError(f"{len(resources)} files, {total} bytes: over the SEP-2640 limits")
    return {"uri": f"{base}/SKILL.md", "frontmatter": frontmatter, "resources": resources}


if __name__ == "__main__":
    print(json.dumps(skill_entry(Path(sys.argv[1]), *sys.argv[2:]), indent=2))
$ uv run skill_entry.py ./refunds acme/billing
{
  "uri": "skill://acme/billing/refunds/SKILL.md",
  "frontmatter": {
    "name": "refunds",
    "description": "Process customer refund requests per company policy. Use when a ticket asks for a refund, a chargeback, or a partial credit.",
    "license": "Apache-2.0"
  },
  "resources": [
    {
      "uri": "skill://acme/billing/refunds/SKILL.md",
      "digest": "sha256:f76de8897595ceccabb7b14cb50773d4df1bc2d2d760690f20c5d275714053bf",
      "size": 365
    },
    {
      "uri": "skill://acme/billing/refunds/examples/email.md",
      "digest": "sha256:9c51b56836ef4edad24243e61a61f29e43dab0fdb3bb0fb2bcc8242d896abdb1",
      "size": 152
    }
  ]
}

Written for this essay from the SEP's Frontmatter, Resources, Integrity, and Limits rules, and run on the sample skill. The output above is the real output. The frontmatter goes through a YAML parser into JSON, so put metadata values in quotes. The Agent Skills spec types them as strings. An unquoted version: 2.1 arrives in the listing as the number 2.1.

When you have the entries, the reference server needs two async handlers and a normal resource registration for each file. The get_skill handler must answer for every skill that the server serves, including skills that list_skills does not return:

from mcp.server.mcpserver import MCPServer
from mcp.server.mcpserver.resources import TextResource
from mcp.server.skills import Skills
from mcp.shared.exceptions import MCPError
from mcp.shared.skills import GetSkillResult, ListSkillsResult
from mcp.types import INVALID_PARAMS

async def list_skills(ctx, params) -> ListSkillsResult:
    return ListSkillsResult(skills=[GIT_WORKFLOW])

async def get_skill(ctx, params) -> GetSkillResult:
    if params.uri != SKILL_URI:
        raise MCPError(code=INVALID_PARAMS, message=f"unknown skill: {params.uri}")
    return GetSkillResult(skill=GIT_WORKFLOW)

mcp = MCPServer("catalog", extensions=[Skills(list_skills=list_skills, get_skill=get_skill)])
mcp.add_resource(TextResource(uri=SKILL_URI, name="SKILL.md", mime_type="text/markdown", text=SKILL_MD))

Trimmed from the tutorial in python-sdk #3485 (the full file defines GIT_WORKFLOW, SKILL_URI, and SKILL_MD). As of September 25, 2026, the pull request is open and not in a release, so names can change before it merges. I did not run it.

Your listing can be partial, or you can want the model to find a skill before a registry exists. For both cases, the SEP allows a plain pointer from instructions. The host still verifies the skill through skills/get before it loads the skill:

"instructions": "Refunds follow company policy. Before touching an order, load skill://acme/billing/refunds/SKILL.md."
Fig. 1 · the wire, message by message

The agent's task: fill the EU invoice template from the pdf-processing skill. Flip what the server declares and how complete its listing is, and the transcript rebuilds itself. File names and byte sizes are the SEP's own example catalog.

extension
directoryRead
listing

    Sizes are raw file bytes from the SEP-2640 skills/list example (catalog of 3 skills, 9 files, 38,866 bytes). JSON-RPC envelopes are not counted. "Verified" means the bytes were checked against a digest the host already held.

    Two details in that transcript are easy to miss. First, the skill is still readable without the extension. A URI in the server's instructions is enough for resources/read to return the bytes. The host loses the data around the bytes. It has no digest to check them against and no frontmatter to show the user. It also cannot tell a skill from an ordinary Markdown resource. Second, the partial listing costs exactly one extra round trip. skills/get turns a URI from any source into the same manifest that a listing gives.

    Progressive disclosure, stretched over a network

    On a local filesystem, progressive disclosure is a context budget. Over MCP, it is also a network budget, and the SEP makes that a rule. Hosts must not retrieve a skill's files before they need them. That applies on connection, on listing, and at approval. A SKILL.md is fetched when the skill is loaded. A supporting file is fetched when it is read. The stated reason is load. A server can publish many skills with many files. A host that downloads everything on connect puts load on the server in proportion to the catalog size. The SEP wants load in proportion to use.

    The digests make lazy loading cheap. A host caches the files that it fetches. If a cached file's digest still matches the current manifest, the host uses it without another request. A mismatch forces a fresh fetch. The same property explains why the maintainers cut the archive form from an earlier draft. A whole skill in one tarball saves round trips. To unpack a server-supplied archive safely, the host must defend against many attacks:

    With lazy retrieval, the number of round trips grows with the files that a session opens. The SEP judged that the better trade.

    Fig. 2 · disclosure tiers across a task

    Rows are the SEP's three example skills plus copies from other servers you connect. Columns are the three Agent Skills tiers. Scrub the task forward and watch which cells turn on. The verdict compares the lazy session with a host that loads every body and file up front.

    4

    Tier 1 is about 100 tokens per skill (Agent Skills spec). Body and file tokens use the SEP example byte sizes at an assumed 4 bytes per token, a rough heuristic for English Markdown. Extra servers are modeled as copies of the same 3-skill catalog.

    The ratio in that verdict is the reason for the design. The first tier grows with the number of skills that you connect. The second and third tiers grow only with the work. An agent connected to a dozen servers pays for their short descriptions. It pays for a manual only when a task opens one.

    A worked example: twelve servers, one invoice

    Use the SEP's example catalog: three skills, nine files, and 38,866 bytes. Connect a host to twelve servers that each serve a catalog of that size. Then ask for the EU invoice. Assume 4 bytes per token for English Markdown:

    Byte sizes are the SEP-2640 example listing. Tokens per byte is a rough heuristic and not a measurement. The ratio stays the same for any constant.

    Most of the spec is a threat model

    This part explains the SEP's length. A tool call runs on the server. A skill runs on the host. It is text from the server, placed directly into the model's context. It can tell the model to run a bundled script with the host's own shell. The SEP says this directly. It requires hosts to treat MCP-served skills as a higher risk than remote tool calls. The rules are specific:

    I expect content-bound approval to matter most. When a host stores a user's approval, it binds the approval to the whole resources set at that moment. That set includes every URI and every digest. An earlier draft had one digest per skill. The rationale describes the attack that design allowed. The attacker gets approval for a benign skill and then changes references/GUIDE.md. With the full set bound, any changed, added, or removed file revokes the approval. A file read that no longer matches fails verification, even in the middle of a task.

    Fig. 3 · manifest tamper bench

    The user approved pdf-processing against the manifest below. Now play the server, or the network in between. Each move changes what the server serves, and the host applies the SEP's rules to decide what happens next.

    Manifest from the SEP-2640 example, digests shortened. Verdicts follow the Integrity and Security Implications sections: digest mismatch, unlisted read, frontmatter mismatch, content-bound approval, allowed-tools, and "digests are not a security boundary."

    The last move on that bench shows a limit, and the SEP states it. Digests are unsigned and come from the same server as the content. They prove that the listing and the bytes agree. They cannot prove that either one is safe. An intermediary that rewrites both passes every check. Trust at approval time still depends on the user reading what they approve.

    For that reason, the SEP asks hosts to let users inspect a skill before it loads. Hosts must also show the source everywhere. Caching has the same rule. A host must keep cached skills where only the host can write, or hash the bytes again on every access. Cached content never gets local-skill trust because it is on disk.

    On the host, all of that becomes one gate for every skill read. The gate runs four checks in a fixed order:

    1. Is the file in the held manifest?
    2. Is the file the promised size?
    3. Does the file hash to the promised digest?
    4. For SKILL.md, does the frontmatter match the listing that the user saw?

    The code:

    # /// script
    # dependencies = ["pyyaml"]
    # ///
    import hashlib
    import json
    
    import yaml
    
    
    class VerificationFailure(Exception):
        pass
    
    
    def verify_read(entry: dict, uri: str, data: bytes) -> bytes:
        if entry["resources"] == "dynamic":
            raise VerificationFailure("dynamic skill: no digests to check, cannot be content-bound")
        held = {r["uri"]: r for r in entry["resources"]}
        if uri not in held:
            raise VerificationFailure(f"{uri} is not in the held manifest")
        if len(data) != held[uri]["size"]:
            raise VerificationFailure(f"{uri}: size {len(data)} != {held[uri]['size']}")
        if "sha256:" + hashlib.sha256(data).hexdigest() != held[uri]["digest"]:
            raise VerificationFailure(f"{uri}: digest mismatch")
        if uri == entry["uri"]:
            served = yaml.safe_load(data.decode("utf-8").split("---", 2)[1])
            if served != entry["frontmatter"]:
                raise VerificationFailure("SKILL.md frontmatter differs from the listing")
        return data
    SKILL.md: 365 bytes verified
    refused: skill://acme/billing/refunds/examples/email.md: size 32 != 152
    refused: skill://acme/billing/refunds/references/GUIDE.md is not in the held manifest

    Written for this essay from the SEP's Integrity and verification rules and run against the entry above. The real SKILL.md passes. A changed email.md fails on size before any hashing. The gate refuses a file that the server added after approval, because the manifest does not list it. The python-sdk pull request ships verify_skill_resource for size and digest. Its docs say it does not compare frontmatter. A host that uses frontmatter must add that step.

    The model reaches the gate through one loading tool keyed by server and URI, never by name. The SEP's implementation guidelines show a sketch. Supporting files go through a matching read_resource tool with the same two arguments:

    {
      "name": "read_skill",
      "description": "Load a skill's SKILL.md into context.",
      "inputSchema": {
        "type": "object",
        "properties": {
          "server": { "type": "string", "description": "Name of the connected MCP server" },
          "uri": { "type": "string", "description": "The skill's SKILL.md URI" }
        },
        "required": ["server", "uri"]
      }
    }

    From the SEP-2640 implementation guidelines, which are illustrative rather than normative.

    Who has built it

    An Extensions Track SEP needs a reference implementation in an official SDK before review. This SEP lists three: Python, C#, and Go. The conformance scenarios for enumeration, manifests, and directory reads merged on September 11. That was two days before the SEP. On the host side, the SEP cites prototypes in forks of gemini-cli and Codex, and in fast-agent. It also cites a prototype inside Anthropic for Claude Code, which is not yet public. On the server side there is a GitHub MCP Server prototype. The authors are Peter Alexander of Anthropic, Ola Hungerford, Sambhav Kothari, and Aditya Kumar. They wrote for the Skills Over MCP Working Group. The design history is in the ext-skills repository.

    The SEP names one migration. FastMCP's widely used SkillsProvider serves skills in a different way. It uses a different URI structure and a manifest resource for each skill in place of a central listing. Its metadata mapping is also different. The SEP calls these mechanical changes. The working group lists their coordination as a near-term priority. If your server uses it, expect your skill URIs to change.

    The division of work will last. The Agent Skills specification owns the content format. The SEP delegates the format to it completely, including future revisions. MCP owns transport, discovery, and verification.

    When it breaks

    Most failures come from the verification rules working as designed. Each row gives the symptom, the rule behind it, and the fix, all from the SEP text:

    SymptomCauseFix
    A host refuses your skill as invalidThe last URI path segment differs from frontmatter.name, or the entry has no resourcesRename the directory to the name. Publish a complete array, or "dynamic" for generated content
    Every load fails verification after a deployThe host holds an old entry. Any byte change breaks size or digestNothing to fix on the host side: it calls skills/get and asks for re-approval. On the server, batch skill edits into releases
    Listing and file disagree on frontmatterThe server built the listing from a curated or edited copy of the frontmatterGenerate the listing from the served SKILL.md itself, every field included
    A new template never shows up for the agentA directory read lists a file the held manifest does notExpected: refresh with skills/get, re-approve, then read it
    resources/directory/read returns method not foundThe server did not declare directoryRead: trueDeclare it and implement the method, or rely on the manifest
    Two servers' refunds skills overwrite each other in the cacheThe host keyed its registry or cache on the URI aloneKey everything on the host's server label plus the URI
    The skill's script never runsConforming hosts gate execution per skill and ignore remote allowed-toolsWrite instructions that still work when the user declines

    Shipping a skill, then hosting one, in order

    For a server, do these steps in order:

    1. Write the skill as an Agent Skills directory and run skills-ref validate on it.
    2. Generate its entry at build time with the builder above. Fail the build if a skill goes above 512 files or 16 MiB.
    3. Serve every file as a resource at the URI that the entry lists. Declare the extension and resources.
    4. Answer skills/list from the generated entries. Answer skills/get for every skill that you serve, listed or not.
    5. Shrink instructions to a pointer at the skill URI.
    6. Ship skill edits in batches. Each edit sends your users back through approval.

    For a host, do these steps in order:

    1. Build the registry from listings only, keyed by server label and URI. Show the origin next to every name.
    2. Expose one loading tool that takes server and URI. Send unknown pairs through skills/get.
    3. Fetch SKILL.md at load and supporting files at read, through the verification gate. Cache by digest where nothing else can write.
    4. Tag loaded content with its origin. Ignore remote allowed-tools. Require per-skill approval for any execution.
    5. Bind approval to the full resources set. Treat any change to that set as a new skill to approve.

    The Skills Extension is small on the wire because resources already did most of the work. The document is long because a server's manual can make the host run code.

    rg
    Rohit Ghumare

    CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Every protocol detail here comes from the SEP-2640 text on modelcontextprotocol.io, its rationale document, and the Agent Skills specification. I read them on September 14, 2026. Merge dates are from the GitHub pull requests. The SEP is a historical record once Final, so check the current specification before you implement.

    Related: Stateless MCP · MCP's New Roadmap · More posts · X