Docker Sandboxes and Docker Agent 101
Isolated microVMs for coding agents, and the agent runtime that runs inside them, command by command
- The Two Products on One Page
- sbx run
- sbx policy, sbx secret, sbx mcp
- Kits and sbxenv.yaml
- docker-agent run and the Agent File
- docker-agent serve and share
- Operating It
How to read this manual
Every command output in this manual comes from one recorded run of sbx and docker-agent on one Mac, and every rule traces to a ranked source.
This manual explains Docker Sandboxes, the sbx CLI at v0.47.0, and Docker Agent, the docker-agent CLI at v1.149.0. It follows one task from start to end: run an agent you do not trust on files you do.
Who this manual is for
You run coding agents such as Claude Code or Codex on real repositories, and you know Docker, git, and YAML. You have not used Docker Sandboxes or Docker Agent.
After the last part, you can:
- Run a coding agent inside a sandbox, give it only the files it needs, and get its commits back without giving it your credentials.
- Read one line of
sbx policy logand say which rule allowed or blocked the connection, and what to change. - Write an agent file that runs a team of agents over local tools and a local model, record it, and replay it from a cassette.
- Serve the same agent file over the HTTP API, MCP, ACP, and A2A, and share it as a signed OCI artifact.
- Map any command from the 2025 Docker Sandboxes plugin or the cagent material to its current form.
The sandbox lesson explains when a task needs a process, a container, or a microVM boundary. This manual starts where that lesson stops, with two products that build the microVM boundary.
How it was made
The manual pins sbx v0.47.0 at commit 0411f50e, released 2026-10-05, and docker-agent v1.149.0 at tag commit bf4169c, released 2026-10-07. The facts were verified on 2026-10-08. sbx is a proprietary binary with no public source, so its help text ranks first.
Facts come from seven sources, ranked by authority. When two of them disagree, the higher one wins and the text says so.
| Rank | Source | What the manual takes from it | Cited as |
|---|---|---|---|
| 1 | the help text of sbx 0.47.0 and docker-agent 1.149.0, as the binaries print it | every command, flag, default, and path | help-sbx or help-agent and a command |
| 2 | agent-schema.json at v1.149.0 | every key, type, and default of the agent file | schema and a definition |
| 3 | the Sandbox Kit Spec v3 at tag v3.0.0-m.8 | the kit descriptor and each capability | kitspec or kitcap and a section or a capability |
| 4 | the Docker documentation, fetched 2026-10-08 | rules and limits that the help text does not state | docs-sbx, docs-agent, docs-dmr, docs-mcp, or docs-desktop and a page |
| 5 | the release notes of sbx and the docker-agent changelog | when a behaviour appeared, changed, or went away | rel-sbx or rel-agent and a version |
| 6 | Docker blog posts and talks | history and positioning, never a rule | blog or talk and a date |
| 7 | the capture kit in capture/ | every command output, file, and record shown | the capture file name |
Sources 1 to 5 are vendored under research/sources/, so the audit checks each quote word for word. Blog posts and talks are not vendored, and their quotes are not checked.
The sources disagree in many places. The largest are the kit generations (conflict C13) and the agent names and template image of each launch path (C9 and C61). Others are host secrets that sbx never injects (C7) and the two agent binaries on one Mac (C48). The sources reference lists every conflict with the ruling this manual prints.
The capture kit
The capture kit drives the real sbx, docker-agent, and docker binaries and writes what they print to capture/out/. It masks the values that change on every run, uses the Python standard library only, and has three recorded tiers:
| Tier | What it needs | What it records |
|---|---|---|
| A | sbx, docker-agent, and the docker client installed, with no daemon, network, model, or key | versions, root help, docker-agent doctor, toolsets, models list, a dry run, and the legacy Desktop commands (files 00, 01, 16, 29) |
| B | sbx signed in, the sandboxd daemon, image pulls from Docker Hub, and HTTPS to example.com and mcp.deepwiki.com | the sandbox lifecycle, ports, cp, templates, --clone, policy, secrets, MCP, kits, env plan, and skills (files 02 to 15) |
| M | Docker Desktop 4.94.0 with Docker Model Runner on TCP port 12434 and ai/qwen3:4b pulled, plus containers, Compose, buildx, and a local registry | agent runs, teams, permissions, sessions, eval, five servers, share, Model Runner, Compose, run --sandbox, and the v3 kit build (files 17 to 28, 98) |
The kit was recorded on one Mac with macOS 26.2 on Apple silicon. It ran sbx v0.47.0, docker-agent v1.149.0, Docker Desktop 4.94.0, and Docker Model Runner with ai/qwen3:4b. The documented default model, ai/qwen3:latest, failed twice to pull on that Mac with a digest mismatch (capture/out/25-model-pull-latest.txt). Every agent file therefore names the 4B tag of the same repository.
This kit cannot run offline in CI. sbx starts real microVMs from images on Docker Hub after a Docker sign-in, and the agent runs need a model. Where the tools are absent, python3 capture/run.py --check prints skipped with the reason and exits 0. Tier C needs Cloud Sandboxes, a GitHub token, or a change to ~/.ssh/config. It is not recorded, and the sections that need it cite the docs.
Most runs of tier M replay a recorded cassette with --fake, so a check compares like with like. The eval, three of the servers, one run with an explicit base_url, and run --sandbox call the model live with temperature: 0. The kit README lists what the recorded run showed that the docs did not say.
Conventions
docker-agent names the Homebrew binary, v1.149.0, in every command. docker agent names the Docker Desktop plugin, v1.144.0 on the capture Mac, and appears only where a program prints it.
$CAPTURE stands for the capture directory and $HOME for the home directory. Values that change on every run appear as tokens such as <uuid>, <hash>, <ts>, and <port>. The command lines of tier M leave out --config-dir, --data-dir, and --cache-dir, which point at capture/work/.
Citations name a place: help-sbx sbx run a command in the help text, schema AgentConfig a definition in the schema, and kitspec §3.4 a section of the kit spec. A docs citation such as docs-sbx Network access policies names a page or heading, and rel-sbx v0.43.0 names a release. A listing copies lines from a capture file unchanged, and a … marks a cut that the note explains.
Figures animate on the web, where a Replay button runs the steps again, and print and reduced motion show the final frame.
A first run
You need sbx v0.47.0, a Docker sign-in with sbx login, and a global network policy, which you set once. The capture set it to the balanced preset:
$ sbx policy init balanced
Global network policy initialized to "balanced".
[exit 0]Then make a scratch directory and open a shell sandbox on it:
mkdir -p ~/sbx-scratch
sbx run shell ~/sbx-scratch
sbx run shell gives you a Bash login shell in a new sandbox, named shell-sbx-scratch after the default <agent>-<workdir> help-sbx sbx run. Inside, pwd prints the same path as on your host. Leave the shell, then list the sandbox and remove it:
sbx ls
sbx rm --force shell-sbx-scratch
The capture made its own shell sandbox with sbx create, which prints this summary:
$ sbx create shell $CAPTURE/fixtures/repo --name m101-demo
sandbox m101-demo
agent shell
workspace $CAPTURE/fixtures/repo (rw)
image docker/sandbox-templates:shell-docker
cpu 10
memory 32 GiB
Pulling image
<layers>
✓ Image ready
✓ Created sandbox m101-demo
…For the agent side, turn on Docker Model Runner with host TCP on port 12434 and pull the model with docker model pull ai/qwen3:4b. Save this agent file as dmr.yaml:
version: "16"
providers:
runner:
provider: dmr
base_url: http://localhost:12434/engines/llama.cpp/v1
models:
qwen:
provider: runner
model: ai/qwen3:4b
temperature: 0
agents:
root:
model: qwen
description: Answers with one word.
instruction: Reply with exactly one word.Run docker-agent doctor first. It finds the runner and the model:
…
Docker Model Runner
Status: reachable, 1 model(s) pulled:
- docker.io/ai/qwen3:4b
Model auto-selection
auto -> dmr/docker.io/ai/qwen3:4b
…Then run the agent without the terminal interface:
$ docker-agent run --exec --last fixtures/agents/dmr.yaml 'Say hello.'
Hello
[exit 0]Open capture/out/27-sandbox-run.txt next. It is the same kind of agent run inside a sandbox, and the end-to-end section reads it line by line.
How the parts are ordered
Part 1 puts both products on one page. Parts 2 to 4 take sbx apart: the sandbox and its daemon, the isolation layers, and kits with environment files. Parts 5 and 6 take docker-agent apart: the agent file and its run, then the served and shared agent. Part 7 covers operation, and the Reference holds the lookup tables:
| Part | What it covers |
|---|---|
| 1 · The Two Products on One Page | sbx runs any agent inside a microVM it owns, docker-agent runs an agent file, and docker-agent run --sandbox joins them through a kit and the proxy. |
| 2 · sbx run | A sandbox has a name, a workspace, a template, and a lifecycle that seven verbs and one daemon control. |
| 3 · sbx policy, sbx secret, sbx mcp | Five layers isolate the agent, and each layer has its own commands: the hypervisor, the workspace, the policy and its proxy, the secret store, and the MCP gateway. |
| 4 · Kits and sbxenv.yaml | A kit declares what a sandbox contains and can reach, and an environment file declares the sandbox and its secrets for approval before anything runs. |
| 5 · docker-agent run and the Agent File | An agent file is agents, models, toolsets, and the rules between them, and docker-agent run drives the loop, records it, and replays it. |
| 6 · docker-agent serve and share | One agent file answers over REST and SSE, an OpenAI-compatible endpoint, MCP, ACP, and A2A, travels as a signed OCI artifact, and runs on a local model. |
| 7 · Operating It | The same sandbox runs in the cloud, an organization narrows it with policies and reads its audit log, old commands map to new ones, and the claims are checked. |
| R · Reference | Every sbx and docker-agent command, settings keys and paths, agent file keys, toolsets, providers, sources and conflicts, a glossary, and the figure index. |
Colour in figures
Each hue keeps one meaning in every figure:
| blue | sbx, sandboxd, and the sandbox microVM |
| violet | docker-agent, agent files, and the agent loop |
| green | the workspace: files, mounts, clones, kits, and templates |
| amber | lifecycle: sandbox states, sessions, and tool call phases |
| teal | stored records: session.db, policy rules, the secret store, the audit log |
| indigo | network traffic: the proxy, served endpoints, and streams |
| plum | models and model calls: Docker Model Runner and hosted providers |
| olive | tools and effects outside the agent: MCP servers, shell commands, HTTP APIs |
| rose | failure: blocked connections, denied tool calls, errors |
| grey | the host, the operator, Docker Hub, and anything outside both products |
Figure 0.1 teaches the eight arrow styles with exchanges from the capture.
The model call in row 8 passes through the proxy, which figure 1.3 draws as a separate hop.
Sources:help-sbx sbx run (research/sources/help-sbx.md); manual.json; research/sources/README.md; research/plan.md; research/conflicts-register.md rows C7, C9, C13, C48, C61; capture/README.md, capture/run.py; capture/fixtures/agents/dmr.yaml; capture/cassettes/18-files.yaml.gz; capture/out/README.md, 00-versions.txt, 03-policy-init.txt, 04-auto-stop.txt, 04-create.txt, 04-workspace.txt, 09-allow.txt, 09-allowed.txt, 09-blocked.txt, 25-doctor.txt, 25-model-ls.txt, 25-model-pull-latest.txt, 25-run-dmr.txt, 27-inside-run.txt, 27-sandbox-run.txt, 29-docker-agent-plugin.txt
The Two Products on One Page
sbx runs any agent inside a microVM it owns, docker-agent runs an agent file, and docker-agent run --sandbox joins them through a kit and the proxy.
sbx and docker-agent
Each binary owns one thing: sbx owns the microVM, its proxy, and its secret store, and docker-agent owns the loop between a model, tools, and sub-agents.
You want a coding agent to work on a repository, but not to read your SSH keys or call any host it likes. Docker gives you two command-line programs for that job, sbx and docker-agent. A Mac with Docker Desktop also has a third name, docker agent, and this section draws the line between all three.
When you finish this section, you can name the owner of each part of a sandboxed run and pick a launch path. You can also tell two agent binaries on one machine apart.
What sbx owns
Sandbox: a microVM that sbx creates for one agent. "Every sandbox runs inside a lightweight microVM with its own Linux kernel" docs-sbx Isolation layers, and it has a private Docker Engine. Inside m101-demo, the capture reads Linux 7.0.14 on Ubuntu 26.04.1 and Docker Engine 29.8.1 (capture/out/04-guest.txt). The workspace appears at the same absolute path, and "containers started by the agent never appear in your host's docker ps" docs-sbx Develop and test locally.
sandboxd: the host daemon behind every local sbx command, reached at $HOME/Library/Application Support/com.docker.sandboxes/sandboxes/sandboxd/sandboxd.sock (capture/out/02-daemon-status.txt). "You don't need Docker Desktop or Docker Engine to use sbx" docs-sbx Install Docker Sandboxes. You do need a Docker sign-in with sbx login, and sbx diagnose reports it as Authentication — authenticated (capture/out/02-diagnose.txt). Conflict C15 records why both claims hold.
The daemon also runs the proxy, and one of its two log categories is proxy help-sbx sbx daemon log-level set. "All outbound TCP traffic from the sandbox routes through a proxy on your host" docs-sbx Architecture. On macOS, sbx secret set keeps each key in the system Keychain docs-sbx Where secrets are stored. The proxy adds the key to a request after the request leaves the VM, so the VM sees only the placeholder proxy-managed (capture/out/04-env.txt). The proxy and secrets take each one apart, and figure 1.1 shows who owns what.
$ sbx --help
Docker Sandboxes creates isolated sandbox environments for AI agents, powered by Docker.
…
Sandbox Commands:
…
create Create a sandbox for an agent
exec Execute a command inside a sandbox
…
run Run an agent in a sandbox
…
Management Commands:
daemon Manage sandboxd daemon
diagnose Diagnose common issues with your sbx installation
mcp Manage MCP servers
policy Manage sandbox policies
…
secret Manage stored secrets
…The agents sbx runs
sbx run accepts eleven agent names: claude, codex, copilot, cursor, devin, docker-agent, droid, gemini, kiro, opencode, and shell help-sbx sbx run. sbx create has eight of them as subcommands and leaves out copilot, droid, and kiro. Those three moved to public kits, and since v0.43.0 they "can be launched by name again" rel-sbx v0.43.0. sbx create docker-agent also answers to the alias cagent help-sbx sbx create docker-agent. shell gives you "a Bash login shell inside a sandbox with no pre-installed agent binary" docs-sbx Shell. Conflict C9 lists the shorter agent lists in older posts.
What docker-agent owns
Agent file: a YAML file that names agents, their models, their toolsets, and the rules between them. Docker Agent is "an open-source framework for building teams of specialized AI agents" docs-agent Docker Agent. Its run command drives the loop: a model call, the tool calls the model asks for, and the sub-agents it delegates to. The agent loop lesson explains that loop, and Part 6 serves and shares the same file.
The docker-agent binary isolates nothing by itself. On the host, its shell toolset runs commands "in the user's environment" (capture/out/16-toolsets.txt).
$ docker-agent --help
Docker AI Agent Runner.
…
Core Commands:
getting-started Learn docker agent with a hands-on interactive tour
run Run an agent
setup Interactively set up a model (built-in provider, local, custom endpoint, or Claude Code)
share Share agents
…
Advanced Commands:
…
eval Run evaluations for an agent
…
sandbox Manage docker-agent sandbox settings
serve Start an agent as a server
sessions Inspect recorded sessions
…Two launch paths
sbx run docker-agent ~/my-project starts from sbx. The sandbox uses docker/sandbox-templates:docker-agent and runs docker-agent run --yolo when you pass no arguments docs-sbx Docker Agent. Only project-level configuration in the workspace reaches it.
docker-agent run --sandbox agent.yaml starts from your agent file. Here docker-agent "orchestrates the installed sbx CLI" docs-agent Sandbox Mode, and --template defaults to docker/docker-agent-sbx-templates:latest help-agent docker-agent run. Two paths use two images, as conflict C61 rules. The capture kit recorded only the second path, and the next section follows it.
Two agent binaries on one machine
Docker Desktop 4.94.0 bundles "Docker Agent v1.144.0" docs-desktop 4.94.0 as the CLI plugin docker agent. Homebrew installs v1.149.0 as docker-agent, which the docs say "can be used as a standalone binary" docs-agent Docker Agent.
$ docker agent version
docker agent version v1.144.0
Commit: 3873760f47ecf22f72f65bd776c056a337d6f35c
[exit 0]
$ docker-agent version
docker-agent version v1.149.0
Commit: Homebrew
[exit 0]The hints that v1.149.0 prints still say docker agent sandbox allow <host> (capture/out/16-sandbox-list.txt), and on this Mac that command runs v1.144.0. This manual writes docker-agent in every command and pins v1.149.0, as conflict C48 rules. A third build runs inside the sandbox template, as the next section shows.
The same plugin list holds sandbox v0.13.0, and docker sandbox prints only its removal notice (capture/out/29-docker-sandbox.txt). The migration section maps its commands. The list also holds ai v1.31.0, the plugin behind docker ai, which the docs call Gordon, "Docker's built-in AI assistant" docs-agent Docker Agent. Gordon is outside this manual.
Sources:help-sbx sbx run, sbx create, sbx create docker-agent, sbx daemon log-level set (research/sources/help-sbx.md); help-agent docker-agent run (research/sources/help-docker-agent.md); docs-sbx Isolation layers, Develop and test locally, Install Docker Sandboxes, Architecture, Where secrets are stored, Shell, Docker Agent (research/sources/docs-sandboxes.md); docs-agent Docker Agent, Sandbox Mode (research/sources/docs-docker-agent.md); docs-desktop 4.94.0 (research/sources/docs-desktop-release-notes.md); rel-sbx v0.43.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C9, C15, C48, C61, C72; capture/out/01-help-sbx.txt, 01-help-docker-agent.txt, 02-daemon-status.txt, 02-diagnose.txt, 04-env.txt, 04-guest.txt, 16-sandbox-list.txt, 16-toolsets.txt, 27-sandbox-run.txt, 27-inside.txt, 29-docker-agent-plugin.txt, 29-docker-plugins.txt, 29-docker-sandbox.txt
docker-agent run --sandbox, one run end to end
One recorded docker-agent run --sandbox shows the kit, the sandbox, and the allowlist, and one sbx exec into the same VM shows the model call through the proxy.
You have an agent file with the filesystem and shell toolsets, and its model is ai/qwen3:4b on Docker Model Runner on your Mac. On the host, shell runs with your permissions. --sandbox moves the tools into a microVM and leaves the model on the host.
This section follows the recorded run in capture/out/27-*, step by step, and figure 1.3 draws it. When you finish this section, you can read the launch summary of a sandboxed run and explain the two failures this run met.
The agent file
The agent file differs from the host version only in its provider. Your host's 127.0.0.1 is "not reachable from inside the sandbox" docs-sbx Accessing host services from a sandbox, so the provider points at host.docker.internal:
version: "16"
providers:
host-runner:
provider: dmr
base_url: http://host.docker.internal:12434/engines/llama.cpp/v1
models:
local:
provider: host-runner
model: ai/qwen3:4b
temperature: 0
…Step 1: allow the model host
Persistent allowlist: the hosts that docker-agent sandbox allow stores, which "are added to the sandbox proxy's allow rules on every subsequent --sandbox run" help-agent docker-agent sandbox allow. The help names it the fix for a Blocked by network policy 403. The proxy translates host.docker.internal to localhost, so the capture allows localhost:12434 before the run:
$ docker-agent sandbox allow localhost:12434 (--config-dir, --data-dir, --cache-dir under fixtures/repo-sandbox/.m101, inside the workspace)
…
Added 1 host(s) to the persistent sandbox allowlist:
+ localhost:12434
[exit 0]
…Step 2: keep the state directories in the workspace
The capture kit keeps docker-agent out of ~/.cagent with --data-dir, --cache-dir, and --config-dir. With --sandbox, a data directory outside the workspace is refused, and the help gives only the default, ~/.cagent help-agent docker-agent run. The capture kit therefore puts all three under .m101/ in the workspace:
…
Error: --data-dir must be inside the sandbox's writable workspace $CAPTURE/fixtures/repo-sandbox
[exit 1]Step 3: the launch summary
…
Models gateway: none configured
Models catalog: allowlisting models.dev in the sandbox proxy
User sandbox allowlist: allowlisting 1 host(s) from `docker agent sandbox allow`:
- localhost:12434
sandbox docker-agent-<hash>
agent docker-agent
workspace $CAPTURE/fixtures/repo-sandbox (rw)
$CAPTURE/fixtures/agents (ro)
$CAPTURE/fixtures/repo-sandbox/.m101/cache/sandbox-kits/<hash> (ro)
$CAPTURE/fixtures/repo-sandbox/.m101/cfg (ro)
image docker/docker-agent-sbx-templates:latest
cpu 10
memory 32 GiB
…
✓ Created sandbox docker-agent-<hash>The run set no --models-gateway. It allows models.dev, because "without it the first catalog lookup fails with a 403 Blocked by network policy error" docs-agent Auto-Kit. The next line repeats the host from step 1.
Auto-kit: a directory that docker-agent stages on the host, "bind-mounted read-only into the VM at the same path" docs-agent Auto-Kit. Here it is .m101/cache/sandbox-kits/<hash>, and its manifest.json holds only agent_ref and built_at (capture/out/27-kit-cache.txt). The workspace is read-write, and the agent file's directory, the kit, and the config directory are read-only.
Step 4: an ordinary sandbox
During the run, sbx ls lists docker-agent-<hash> with the agent docker-agent, the status running, and the same four workspaces, three of them :ro (capture/out/27-sbx-ls.txt). The sandbox is an ordinary one, and sbx lists and removes it like any other.
Step 5: the session store fails
Inside the VM, the run stopped before any model call:
…
Error: creating session store: migration failed even after database reset: failed to create migrations table: attempt to write a readonly database (1032)
Error:
[exit 1]The data directory sits on the virtiofs workspace mount, which mount lists as rw. A Python sqlite3 write on the same mount fails the same way (capture/out/27-sqlite-probe.txt), so on this host SQLite cannot write through that mount. The run printed no safety mode, so conflict C62 stays open. Docker Agent "exits but does not stop or remove the sandbox VM" docs-agent Sandbox Mode, and the VM stayed running.
Step 6: the model call from inside
docker-agent version main
Commit: 154b78f2d517dae1397866dfb7586c4a302a3151
…
ANTHROPIC_API_KEY=proxy-managed
…
HTTP_PROXY=http://gateway.docker.internal:3128
…
NO_PROXY=localhost,127.0.0.1,::1,gateway.docker.internal
OPENAI_API_KEY=proxy-managed
…
Persistent sandbox allowlist is empty.The template runs docker-agent version main, not v1.149.0, although the docs build :latest from "The most recent v* release" docs-agent Sandbox Mode. The provider keys read proxy-managed. The in-VM allowlist is empty, because the allowed hosts arrive as proxy rules from the host.
The capture then starts the same agent file inside the VM by hand, with its state in /tmp/m101. This run shows the model path from the VM, not a working --sandbox launch:
$ sbx exec docker-agent-<hash> sh -c 'docker-agent --config-dir /tmp/m101 --data-dir /tmp/m101 --cache-dir /tmp/m101 run --exec --last …
…
The working directory contains README.md with 1 line.
[exit 0]The request to host.docker.internal:12434 leaves through HTTP_PROXY, because that host is not in NO_PROXY. "The sandbox proxy translates host.docker.internal to localhost before forwarding the request" docs-sbx Accessing host services from a sandbox, and the rule from step 1 admits it. The answer is correct: the capture kit copies the same README into every workspace, and it holds the one line hello (capture/out/04-workspace.txt).
Step 7: remove the sandbox
$ sbx rm --force docker-agent-<hash>
Deleting sandbox docker-agent-<hash>...
Sandbox 'docker-agent-<hash>' removed
[exit 0]A later run from the same workspace reuses the VM, and creates a new one "only when the mount set has changed" docs-agent Sandbox Mode.
Sources:help-agent docker-agent run, docker-agent sandbox allow (research/sources/help-docker-agent.md); docs-agent Sandbox Mode, Auto-Kit (research/sources/docs-docker-agent.md); docs-sbx Accessing host services from a sandbox (research/sources/docs-sandboxes.md); research/conflicts-register.md rows C61, C62; capture/README.md; capture/fixtures/agents/files-sandbox.yaml; capture/out/27-sandbox-allow.txt, 27-data-dir-outside.txt, 27-sandbox-run.txt, 27-kit-cache.txt, 27-sbx-ls.txt, 27-sqlite-probe.txt, 27-inside.txt, 27-inside-run.txt, 27-sbx-rm.txt, 04-workspace.txt; capture/cassettes/18-files.yaml.gz
sbx run
A sandbox has a name, a workspace, a template, and a lifecycle that seven verbs and one daemon control.
sbx run, sbx create, sbx stop, and sbx rm
sbx run creates a sandbox when none exists and attaches to it, sbx create only creates, and stop, rm, and prune end a sandbox in three different ways.
You have a repository with two commits and a README, and you want a shell agent to work on it and on nothing else. The capture does that with sbx create shell $CAPTURE/fixtures/repo --name m101-demo, where $CAPTURE stands for the capture directory, and the rest of this part reuses that sandbox. When you finish this section, you can name the verb that creates, attaches, stops, removes, or prunes a sandbox, and what each one keeps.
The agent and the workspace
Agent: the first positional argument of sbx run and sbx create, "a built-in agent name or a sandbox kit reference" help-sbx sbx run. The built-in names are claude, codex, copilot, cursor, devin, docker-agent, droid, gemini, kiro, opencode, and shell. A kit reference is a directory, a ZIP file, a git repository, or an OCI reference, and relative paths "must be explicit paths such as ./my-kit or ../my-kit.zip" help-sbx sbx run. So sbx run my-kit looks for an agent called my-kit, and sbx run ./my-kit reads a kit, as the kit descriptor explains.
Workspace: the directory after the agent, mounted inside the microVM at the same absolute path. Inside m101-demo, pwd prints $CAPTURE/fixtures/repo, and mount lists that path as a virtiofs mount. sbx run mounts the current directory when you give no path, and sbx create without a path makes a sandbox with no mount, where "the agent then works in the container's own filesystem instead of on your files" help-sbx sbx create. Extra paths follow the first one, and :ro mounts one of them read-only. A read-only argument can name a single file, "which holds that one path out of reach inside a workspace the sandbox can otherwise write" help-sbx sbx run.
$ sbx create shell $CAPTURE/fixtures/repo --name m101-demo
sandbox m101-demo
agent shell
workspace $CAPTURE/fixtures/repo (rw)
image docker/sandbox-templates:shell-docker
cpu 10
memory 32 GiB
Pulling image
<layers>
✓ Image ready
✓ Created sandbox m101-demo
To connect to this sandbox, run:
sbx run --name m101-demo
[exit 0]$ sbx exec m101-demo sh -c 'pwd; echo; ls -la; echo; cat README.md; echo; git log --oneline'
$CAPTURE/fixtures/repo
…
c7dfd56 second
086df68 first
[exit 0]Names and re-attach
Name: the key of a sandbox in every later command, by default <agent>-<workdir>, the agent name and the workspace directory name. The rule is "at least two characters, starting with a letter or number, containing only letters, numbers, hyphens and periods (periods are rejected with --cloud); 'default' is reserved" help-sbx sbx create. Since v0.43.0 the daemon also rejects "names longer than 63 characters or ending in a hyphen or period" rel-sbx v0.43.0. The 2025 plugin allowed _ and + in a name, and conflict C25 records the change.
Two sandboxes can share one workspace. The capture creates m101-demo-2 on the same fixtures/repo, and sbx ls --json lists both as running. A blog post of 2026-03-11 said that Docker enforces one sandbox per workspace blog 2026-03-11. A talk of 2026-08-13 runs claude and codex on one workspace in two terminals talk 2026-08-13, and S27 rules that the names decide.
--name on sbx run re-attaches to an existing sandbox, and then "the agent positional is optional when the named sandbox already exists and is read from its spec" help-sbx sbx run. Give the agent too, and sbx run checks it against the stored one. sbx create prints that re-attach command at the end of its output.
{
"sandboxes": [
{
"name": "m101-demo",
…
"status": "running",
…
"workspaces": [
"$CAPTURE/fixtures/repo"
],
…
},
{
"name": "m101-demo-2",
…
"status": "running",
…
"workspaces": [
"$CAPTURE/fixtures/repo"
],
…
}
]
}What is fixed at creation
Some flags act on every attach, and most act once, when the sandbox is created. A re-attach with a different -p or --skills changes nothing, and the help says so flag by flag.
| Flag | When it applies | Help page |
|---|---|---|
-e KEY=VALUE, --env-file FILE | the agent session on every attach, and stored in the sandbox when this run creates it | sbx run |
--cpus N, -m SIZE | at creation, memory defaults to half of host memory, between 512 MiB and 32 GiB | sbx create |
-p PORT, --deny-network HOST | at creation, and -p is ignored when re-attaching | sbx run |
--skills MODE, --static-mcp NAMES, --profile NAME | at creation only | sbx run |
-t IMAGE, --pull POLICY | at creation, and --pull is always, missing, or never, default always | sbx run |
-d | prints the sandbox id and exits without a session | sbx run |
--rm | removes the sandbox after the agent session exits, new in v0.47.0 | sbx run |
The recording Mac gave m101-demo cpu 10 and memory 32 GiB, the upper bound of the memory clamp. --rm "cannot be combined with --detached" rel-sbx v0.47.0, because a detached run has no session to end.
stop, rm, and prune
sbx stop keeps the sandbox and ends its microVM. The capture prints state preserved with the restart command, and the docs list what persists: "installed packages, Docker images, configuration changes, command history, and mountless workspace files all persist across stops and restarts" docs-sbx usage. The daemon also stops a sandbox on its own. After the capture closed its last session and waited, sbx ls showed m101-demo as stopped, and daemon.log gives the reason.
$ grep auto-stop daemon.log | grep m101-demo | tail -2
{"time":"<ts>","level":"INFO","msg":"auto-stop grace period expired, stopping runtime","version":"v0.47.0 0411f50ee4700fe7bd37e6e7e3aced563e850ca9","runtime":"m101-demo"}
{"time":"<ts>","level":"INFO","msg":"auto-stopped runtime after last session disconnected","version":"v0.47.0 0411f50ee4700fe7bd37e6e7e3aced563e850ca9","runtime":"m101-demo"}sbx run -d against an existing sandbox "keeps it running after sessions disconnect, until you stop or remove it" docs-sbx release-notes, and a kit can declare com.docker.sandbox/long-running@1 for the same effect, as the capabilities show.
sbx rm is the opposite of stop. For a local sandbox it "stops them, removes their containers, cleans up any Git worktrees, deletes sandbox state, and deletes secrets scoped to each removed sandbox" help-sbx sbx rm. --force skips the prompt and removes a sandbox with an open SSH connection, and --all removes every local sandbox.
sbx prune removes stopped sandboxes only, because "a running sandbox is never removed" help-sbx sbx prune. --filter until=168h keeps anything stopped within the last week, and a sandbox whose stop time the daemon cannot report is left alone. --dry-run --json shows both sets before you commit.
{
"would_remove": [
{
"name": "m101-demo-2",
"agent": "shell",
"stopped_at": "<ts>",
"workspaces": [
"$CAPTURE/fixtures/repo"
]
}
],
"skipped_unknown_stop": []
}$ sbx prune --force
Deleting sandbox m101-demo...
Sandbox 'm101-demo' removed
Deleting sandbox m101-demo-2...
Sandbox 'm101-demo-2' removed
[exit 0]
$ sbx ls --json
{
"sandboxes": []
}
[exit 0]Both rm and prune delete the secrets scoped to the sandbox, which sbx secret explains. Figure 2.1 puts the verbs on the edges of one state machine.
Sources:help-sbx sbx run, sbx create, sbx stop, sbx rm, sbx prune (research/sources/help-sbx.md); rel-sbx v0.43.0 and v0.47.0 (research/sources/sbx-releases.md); docs-sbx usage and release-notes (research/sources/docs-sandboxes.md); conflicts C25, S27 (research/conflicts-register.md); capture/out/04-create.txt, 04-workspace.txt, 04-guest.txt, 04-ls.json, 04-second-sandbox.txt, 04-second-sandbox.json, 04-stop.txt, 04-auto-stop.txt, 04-prune-dry-run.json, 04-prune.txt
sbx exec, sbx cp, and sbx ports
Three verbs reach into a sandbox: exec runs a command and starts a stopped sandbox, cp moves files across the boundary, and ports binds tcp4 unless told otherwise.
A server inside m101-demo listens on port 8080. From the host you want to open it, copy a file in, and read a result out. docker exec, docker cp, and docker run -p do those jobs for a container, and sbx has the same three verbs with a changed rule in each. When you finish this section, you can run a command as any user, copy a file either way, and publish a port that localhost reaches.
sbx exec
exec: sbx exec [flags] SANDBOX COMMAND [ARG...] runs one command inside the sandbox, and "If the sandbox is stopped, it is started first" help-sbx sbx exec. The flags follow docker exec, "except detached exec (-d/--detach) is not supported" help-sbx sbx exec, so -d appears in the help only to say so. The 2025 plugin ran docker sandbox exec -d detached, which conflict C31 records. With --cloud, the flags -d, --user, and --privileged "are rejected rather than silently ignored" help-sbx sbx exec.
The command runs as the agent user, uid 1000, with the sudo and docker groups, in the workspace directory. -u root runs it as root, which the help shows for apt-get update and the capture uses to create /opt/marker-from-demo. -w sets another working directory, -e KEY=VALUE and --env-file FILE set variables for that one command, and -it opens a shell.
Inside, a sandbox knows its own name. Since v0.39.0 sandboxes "expose their own identity as SANDBOX_NAME and SANDBOX_ID environment variables" rel-sbx v0.39.0, and SANDBOX_VM_ID stays as a deprecated copy of the name. WORKSPACE_DIR names the mount, and HOME is /home/agent.
$ sbx exec m101-demo sh -c 'env | sort'
…
HOME=/home/agent
…
PWD=$CAPTURE/fixtures/repo
…
SANDBOX_ID=<uuid>
SANDBOX_NAME=m101-demo
SANDBOX_VM_ID=m101-demo
…
WORKSPACE_DIR=$CAPTURE/fixtures/repo
…
[exit 0]sbx cp
cp: sbx cp SRC DST, where "Either SRC or DST must be a sandbox path, written as SANDBOX:PATH" help-sbx sbx cp and the other side is a host path. "Copying between two sandboxes is not supported" help-sbx sbx cp. The capture's sbx cp m101-demo:/tmp/out.txt m101-demo-2:/tmp/x exits 1 with that message, so route such a copy through the host. A directory copy places the directory itself at the destination. When the destination is an existing directory, the source goes inside it, and -L follows symbolic links in the source. The sandbox path is a container path, so /tmp/in.txt in the capture sits outside the workspace mount and is deleted with the sandbox.
The one rule with a security history is copy-out. Release v0.38.0, published 2026-08-06, "Fixed a destination-escape flaw in sbx cp copy-out (CVE-2026-17106)" rel-sbx v0.38.0.
$ sbx cp $CAPTURE/work/in.txt m101-demo:/tmp/in.txt
[exit 0]
$ sbx exec m101-demo sh -c 'cat /tmp/in.txt; printf out > /tmp/out.txt'
in
[exit 0]
$ sbx cp m101-demo:/tmp/out.txt $CAPTURE/work/out.txt
[exit 0]
$ cat $CAPTURE/work/out.txt
out
[exit 0]
$ sbx cp m101-demo:/tmp/out.txt m101-demo-2:/tmp/x
error: copying between sandboxes is not supported
try: sbx cp --help
[exit 1]sbx ports
ports: sbx ports SANDBOX lists the published ports, --publish SPEC adds one, and --unpublish SPEC removes one. The spec is [[HOST_IP:]HOST_PORT:]SANDBOX_PORT[/PROTOCOL], and "If HOST_PORT is omitted, an ephemeral port is allocated automatically" help-sbx sbx ports. Publishing "starts a stopped sandbox before creating the host binding" help-sbx sbx ports, as exec does.
The rule to learn is the address family. "When publishing without a PROTOCOL, tcp4 is used" help-sbx sbx ports, so the host binding is 127.0.0.1 alone, and you "Publish tcp explicitly to bind both families" help-sbx sbx ports. This changed in v0.42.0, when "a published port no longer listens on ::1 unless you name the protocol explicitly" rel-sbx v0.42.0. The 2025 plugin bound both loopbacks, as conflict C39 records.
PROTOCOL | Host binding when HOST_IP is omitted |
|---|---|
| none | 127.0.0.1 as tcp4, or tcp6 on ::1 when HOST_IP is an IPv6 address |
tcp, udp | 127.0.0.1 and ::1, or 127.0.0.1 alone when the sandbox is IPv4-only |
tcp4, udp4 | 127.0.0.1 |
tcp6, udp6 | ::1 |
The capture publishes 18081:8080, gets 127.0.0.1:18081 -> 8080/tcp4, and curl answers 200 on 127.0.0.1 and fails with exit 7 on ::1.
$ sbx ports m101-demo --publish 18081:8080
Published 127.0.0.1:18081 -> 8080/tcp4
[exit 0]
$ sbx ports m101-demo
HOST IP HOST PORT SANDBOX PORT PROTOCOL
127.0.0.1 18081 8080 tcp4
[exit 0]$ curl -sS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18081/README.md
200
[exit 0]
$ curl -sS -o /dev/null -w '%{http_code}\n' 'http://[::1]:18081/README.md'
curl: (7) Failed to connect to ::1 port 18081 after <n> ms: Couldn't connect to server
000
[exit 7]--publish 8080/tcp then binds both families, each on an ephemeral host port, and sbx ports --json lists every binding as host_ip, host_port, sandbox_port, and protocol.
$ sbx ports m101-demo --publish 8080/tcp
Published 127.0.0.1:<port> -> 8080/tcp
Published [::1]:<port> -> 8080/tcp
[exit 0]
$ sbx ports m101-demo --json
…
{
"host_ip": "::1",
"host_port": "<port>",
"sandbox_port": 8080,
"protocol": "tcp"
}
]
[exit 0]Unpublish without a protocol removes the mapping "whether it was published with that same default or as dual-stack tcp" help-sbx sbx ports, and the capture ends with sbx ports m101-demo --json printing []. -p on sbx run applies at creation only, so change ports later with sbx ports. sbx ls prints the bindings in its PORTS column, as 127.0.0.1:18081->8080/tcp4 in the capture. In the cloud a published port gets a public URL and UDP is refused, as sbx --cloud shows.
Sources:help-sbx sbx exec, sbx cp, sbx ports, sbx run (research/sources/help-sbx.md); rel-sbx v0.38.0, v0.39.0, v0.42.0 (research/sources/sbx-releases.md); docs-sbx usage (research/sources/docs-sandboxes.md); conflicts C31, C39 (research/conflicts-register.md); capture/out/04-env.txt, 04-guest.txt, 04-workspace.txt, 05-ports.txt, 05-ports.json, 05-ports-curl.txt, 05-ports-ls.txt, 05-ports-tcp.txt, 05-ports-unpublish.txt, 06-cp.txt, 07-template-save.txt
sbx template save and sbx template load
A template is a snapshot in the sandbox runtime's image store, reused with --pull never -t TAG, exported as a tar, and loaded on another host.
You spent an hour inside m101-demo installing a toolchain, and tomorrow a second sandbox needs the same tools without the hour. The capture leaves a marker file at /opt/marker-from-demo, saves the sandbox as m101-tpl:v1, and creates m101-from-tpl from it. When you finish this section, you can save a template, start a sandbox from it, and carry it to another host as a tar.
What a template holds
Template: a saved snapshot of a sandbox's container filesystem, stored as an image: "Templates are saved snapshots of sandboxes that can be reused to create new sandboxes with: sbx run --pull never -t TAG AGENT [WORKSPACE]" help-sbx sbx template. The docs draw the line between image and kit: "A template contains image content; the agent kit still supplies runtime settings such as credentials and network rules" docs-sbx Saving a sandbox as a template. Mounted filesystems are not in it, because "A saved template isn't a backup of the whole sandbox" docs-sbx Saving a sandbox as a template, and that excludes host workspaces and the Docker store at /var/lib/docker. Everything else in the container filesystem is in it, including any key an agent wrote to a file. Keep credentials in sbx secret and the proxy instead. Agent configuration files such as /home/agent/.claude/settings.json "are always recreated when a sandbox is created" docs-sbx Limitations, so a change there does not survive either.
Every built-in agent starts from docker/sandbox-templates:<variant>, and m101-demo pulled docker/sandbox-templates:shell-docker. Variants with a -docker suffix "include Docker Engine for building and running containers inside the sandbox" docs-sbx Choose a template, which is why docker version answers inside m101-demo. With platform.images.useDHI set, "the default template docker/sandbox-templates:claude-code-docker becomes dhi/sbx-templates:claude-code-docker" docs-sbx platform.images.useDHI, and the tag stays the same.
sbx template save
sbx template save SANDBOX TAG needs a stopped sandbox. The capture tries it on the running m101-demo, reads cannot save a running sandbox, stops it, and saves m101-tpl:v1. The image "is stored in the sandbox runtime's image store" help-sbx sbx template save, and that store is separate from the image store of the Docker daemon on the host docs-sbx Load a template. The 2025 plugin's docker sandbox save loaded the image into the host daemon by default, and --output was the way to a file, as conflict C32 records.
$ sbx template save m101-demo m101-tpl:v1 --output $CAPTURE/work/m101-tpl.tar
Sandbox m101-demo is running and must be stopped before saving. Stop it now? (y/N): error: cannot save a running sandbox; stop it first with:
try: sbx stop m101-demo
[exit 1]
$ sbx stop m101-demo
Sandbox 'm101-demo' stopped; state preserved. Restart with: sbx run --name m101-demo
[exit 0]
$ sbx template save m101-demo m101-tpl:v1 --output $CAPTURE/work/m101-tpl.tar
Snapshotting image in sandbox ...
Exporting image to $CAPTURE/work/m101-tpl.tar ...
Exported to $CAPTURE/work/m101-tpl.tar
Save complete. To use the image as a template:
sbx run --pull never -t docker.io/library/m101-tpl:v1 AGENT [WORKSPACE]
[exit 0]sbx template ls --json then lists three images: the base docker.io/docker/sandbox-templates:shell-docker, the new docker.io/library/m101-tpl:v1 with flavor shell, and docker.io/sandboxes-swap/m101-demo:bf19e491. The last one is an image the runtime keeps for the stopped sandbox, and it is gone once every sandbox is removed. Each record has id, repository, tag, flavor, created_at, and size, and the capture masks the id, so this manual does not state its format.
{
"images": [
…
{
"id": "<id>",
"repository": "docker.io/library/m101-tpl",
"tag": "v1",
"flavor": "shell",
"created_at": "<ts>",
"size": "<n>"
},
…
]
}Reuse with --pull never -t
Reuse is sbx run --pull never -t TAG AGENT [WORKSPACE], and the help explains the flag: "Use --pull never to use the saved image without trying to pull it from a registry" help-sbx sbx template save. The capture uses sbx create --pull never -t m101-tpl:v1 shell --name m101-from-tpl, reads Checking image instead of Pulling image, and finds /opt/marker-from-demo inside the new sandbox. Name the agent the template was built for, because a Claude template run with codex prints a warning that the sandbox "may not work correctly" docs-sbx Limitations.
$ sbx create --pull never -t m101-tpl:v1 shell --name m101-from-tpl
sandbox m101-from-tpl
agent shell
workspace none · no workspace bind mount
image m101-tpl:v1
cpu 10
memory 32 GiB
Checking image
✓ Image ready
✓ Created sandbox m101-from-tpl
…
$ sbx exec m101-from-tpl sh -c 'ls -l /opt/marker-from-demo'
-rw-r--r-- 1 root root 0 <date> /opt/marker-from-demo
[exit 0]sbx template inspect is "Cloud-only in v1: requires --cloud" help-sbx sbx template inspect. The local attempt in the capture exits 1 with a hint to run docker image inspect -- m101-tpl:v1. That command reads the host daemon's store, which the runtime does not share, so use sbx template ls for a local template. sbx template rm m101-tpl:v1 --force prints Removed: m101-tpl:v1, and sbx reset clears the cached images too.
Export and load
--output FILE on save also writes a tar. work/m101-tpl.tar has 24 entries and starts with blobs/sha256/, the layout of an OCI image. On the other host, sbx template load FILE loads "an image from a tar file into the sandbox runtime's image store" help-sbx sbx template load, and sbx run --pull never -t m101-tpl:v1 shell starts from it. The capture did not load the tar on a second machine, so figure 2.2 draws that step from the help text. The same path imports an image you built yourself: docker image save writes the tar and sbx template load imports it, so the image "doesn't need to be reachable from a registry at sandbox creation time" docs-sbx Load a template.
In the cloud, sbx --cloud template load FILE NAME needs --cpus and --memory-mib that together name a billable shape help-sbx sbx template load. The shapes run from micro, 1 vCPU and 2048 MiB, to xl, 16 vCPU and 32768 MiB. --capture-mode all adds memory and a microVM checkpoint to the disk capture. Tier C was not recorded, so this manual shows no cloud listing, and sbx --cloud cites the help text instead.
Sources:help-sbx sbx template, sbx template save, sbx template load, sbx template inspect, sbx template rm (research/sources/help-sbx.md); docs-sbx usage, Saving a sandbox as a template, Load a template, Template caching, Choose a template, platform.images.useDHI (research/sources/docs-sandboxes.md); conflict C32 (research/conflicts-register.md); capture/out/04-create.txt, 04-guest.txt, 07-template-save.txt, 07-template-ls.json, 07-template-ls.txt, 07-template-run.txt, 07-template-inspect.txt, 07-template-rm.txt, 07-template-tar-head.txt, 99-final-state.txt
sandboxd, sbx diagnose, and sbx reset
One daemon owns every sandbox, its socket and log live under one state directory, and two commands check that daemon and wipe it.
sbx ls hangs, or a command fails with ensure daemon: daemon exited unexpectedly, and you need to know which process answers, where it writes, and how to start over. Every verb in this part talks to one host daemon, sandboxd, over a Unix socket. When you finish this section, you can find the socket and the log, read the thirteen checks of sbx diagnose, and say what sbx reset deletes.
sandboxd
sandboxd: the host daemon behind every local sbx verb, managed with sbx daemon start, stop, restart, status, and log-level set. The settings commands "use the local daemon to read evaluated values and manage overrides, starting it if necessary" help-sbx sbx settings, and the first probe of sbx settings list printed Starting sandboxd daemon... before its table. sbx daemon start -d runs it in the background. --policy allow-all, balanced, or deny-all initializes the global network policy at the same time, which sbx policy init covers. sbx daemon log-level set TARGET LEVEL changes one log category, proxy, general, or all.
sbx daemon status --json names the socket and the log, both under ~/Library/Application Support/com.docker.sandboxes/sandboxes/sandboxd/ on the recording Mac, where $HOME is a capture mask.
{
"status": "running",
"socket": "$HOME/Library/Application Support/com.docker.sandboxes/sandboxes/sandboxd/sandboxd.sock",
"logs": "$HOME/Library/Application Support/com.docker.sandboxes/sandboxes/sandboxd/daemon.log"
}The log is JSON lines with time, level, msg, version, and often a runtime, the sandbox name, and the lifecycle section reads two of them. Since v0.43.0 "The local daemon now verifies connecting operating-system users on Unix sockets and Windows named pipes" rel-sbx v0.43.0, and since v0.46.0 "Starting a second daemon against a state directory already in use fails with an error instead of disrupting the running daemon" docs-sbx release-notes.
The state directory
State directory: the tree where sandboxd keeps sandboxes, images, policies, and its socket. On macOS it is ~/Library/Application Support/com.docker.sandboxes/, on Windows %LOCALAPPDATA%\DockerSandboxes, and on Linux the state "is spread across three directories" docs-sbx Removing all state: ~/.local/state/sandboxes/, ~/.cache/sandboxes/, and ~/.config/sandboxes/. Release v0.29.0 added the SANDBOXES_STORAGE_ROOT override rel-sbx v0.29.0, and v0.31.0 moved "the state directory symlink from /tmp to ~/.sbx/run/" rel-sbx v0.31.0. The daemon binds its containerd socket under that symlink. The sbx kit builder history help pages, captured in a shell that could not bind sockets, end in "failed to create unix socket on $HOME/.sbx/run/d/containerd/containerd.sock.ttrpc" help-sbx sbx kit builder history.
Two more trees matter: the shared skills store at sandboxes/agent-skills under the state directory docs-sbx agent-skills, and the audit log under ~/Library/Logs/com.docker.sandboxes/sandboxes/auditkit/, where "Files are named audit-<utc-timestamp>-<process-uuid>-<seq>.jsonl" docs-sbx audit. The 2025 plugin kept its VMs under ~/.docker/sandboxes/vm/ and its images under ~/.docker/sandboxes/image-cache/, which the migration section maps. Figure 2.3 draws the macOS tree, and the settings table lists every path for the three systems.
sbx diagnose
sbx diagnose runs thirteen checks in four groups and prints a pass, a warning, or a failure for each. Installation finds the binary, its version, the daemon, and a diagnostics bundle. Platform confirms kern.hv_support is 1 and finds mkfs.erofs with its default block size and the guest kernel page size. Storage checks the state directory, its permissions, and the free space. Connection checks the version match, the socket, the SSH client config, and the sign-in.
The page-size check exists because v0.43.0 made diagnose warn "if its default block size exceeds the sandbox kernel's page size" rel-sbx v0.43.0. On the recording Mac the block size is 4096 bytes and the guest page size is 16384 bytes, which getconf PAGESIZE inside m101-demo confirms.
{
"version": "1.0",
"checks": [
…
{
"name": "Virtualization",
"status": "pass",
"message": "supported",
"detail": "kern.hv_support is 1",
"hint": ""
},
{
"name": "mkfs.erofs",
"status": "pass",
"message": "found",
"detail": "/opt/homebrew/Caskroom/sbx/0.47.0/Sbx.app/Contents/libexec/mkfs.erofs, default block size <n> bytes, guest kernel page size <n> bytes",
"hint": ""
},
…
],
"summary": {
"pass": 13,
"warn": 0,
"fail": 0,
"skip": 0
}
}--json gives a version, a checks array with name, status, message, detail, and hint, and a summary with pass, warn, fail, and skip counts. -o github-issue formats the same report for a bug report, and --upload sends a bundle to Docker support. diagnostics.autoUpload in sbx settings list is the consent for automatic uploads.
sbx reset
sbx reset returns the install to a freshly installed state, and the help lists the steps help-sbx sbx reset:
- stop running sandboxes, with a 30 s timeout
- clear the image cache and the internal registries
- delete all sandbox state and all policies
- remove the managed SSH configuration
- clear "the Gordon assistant's sessions and history" help-sbx sbx reset
- delete stored secrets, unless
--preserve-secrets - sign out, stop the daemon, and remove the state, cache, and config directories
The docs give Gordon no role in a sandbox, so this manual quotes that line and claims nothing more, as conflict C33 rules. The docs reach for the command after an upgrade: "A newer version of sbx upgraded the local database to a schema that older binaries don't understand" docs-sbx sbx reset, and --preserve-secrets keeps your secrets through that. The last resort is to delete the state directory by hand after sbx reset. Your workspaces stay where they are. The docs say host workspace files "remain on your host" when a sandbox is removed, and sbx reset lists only state, cache, and config directories. Commands that help text names but the help tree does not list, such as sbx mount, are collected in sbx settings list.
Sources:help-sbx sbx daemon, sbx daemon start, sbx daemon log-level set, sbx diagnose, sbx reset, sbx settings, sbx kit builder history (research/sources/help-sbx.md); rel-sbx v0.29.0, v0.31.0, v0.43.0 (research/sources/sbx-releases.md); docs-sbx release-notes, Removing all state, agent-skills, audit, usage (research/sources/docs-sandboxes.md); research/sources/probes/sbx-diagnose.txt, sbx-settings-list.txt; help-legacy-docker-sandbox.md (docker sandbox reset); conflict C33 (research/conflicts-register.md); capture/out/02-daemon-status.json, 02-daemon-status.txt, 02-diagnose.json, 02-diagnose.txt, 04-auto-stop.txt, 04-guest.txt
sbx settings list
Every setting has a default, a source, a user override, and a restart flag, some have an environment alias, and sbx settings list --json is the only complete list.
Your company routes every connection through a proxy, and your first sbx run cannot pull a template. The fix is a setting, proxy. The questions are where to set it, whether the daemon must restart, and which value wins when an environment variable disagrees. When you finish this section, you can read one record of sbx settings list --json, change a setting, and know when sbx daemon restart is needed.
One record
Setting: a key such as proxy with a type, a default, an evaluated value, and a source. sbx settings list --json printed 31 records on the recording Mac, each with key, type, default, value, source, and description, and some with env_var or requires_restart. The source is one of three: "The SOURCE column shows where the value came from (default, envvar, or override)" help-sbx sbx settings list. The type is bool, int, float, string, or json, and sbx settings set KEY VALUE parses the value by that type help-sbx sbx settings set. Thirteen of the 31 keys have no env_var, among them proxy, mcp.forceLocalGateway, and skills.defaultMode, so the environment cannot set them and only an override can.
[
…
{
"default": [
{
"identityRegexp": "^.*@docker\\.com$",
"issuer": "https://accounts.google.com"
}
],
"description": "JSON array of trusted signer policies (key-based {\"key\":path} or keyless {\"issuer\":...,\"identity\":...}). Defaults to Docker employee identities (*@docker.com via https://accounts.google.com).",
"env_var": "DOCKER_SANDBOXES_KIT_TRUSTED_SIGNERS",
"key": "kit.trustedSigners",
"source": "default",
"type": "json",
…
},
…
{
"default": "",
"description": "Upstream proxy for sandbox, daemon, and supported CLI host egress (URL, PAC source, \"system\", or \"direct\"; empty = automatic: HTTP(S)_PROXY if set, otherwise the host OS proxy).",
"key": "proxy",
"requires_restart": true,
"source": "default",
"type": "string",
"value": ""
},
…
{
"default": "",
"description": "Upstream proxy for sandbox egress only (overrides proxy).",
"env_var": "DOCKER_SANDBOXES_PROXY",
"key": "proxy.sandbox",
"requires_restart": true,
"source": "default",
"type": "string",
"value": ""
},
…
{
"default": "readonly",
"description": "Default for an omitted --skills flag or sbx.yaml `skills:` key: \"off\", \"readonly\", or \"readwrite\". Applies to sandboxes created after the change; existing sandboxes' mounts are never retroactively changed.",
"key": "skills.defaultMode",
"source": "default",
"type": "string",
"value": "readonly"
},
…
{
"default": "shell",
"description": "Built-in agent used for SSH auto-created sandboxes.",
"env_var": "DOCKER_SANDBOXES_SSH_DEFAULT_AGENT",
"key": "ssh.defaultAgent",
"requires_restart": true,
"source": "default",
"type": "string",
"value": "shell"
},
…
]The proxy record shows the shape of a value. It takes a URL, a PAC source, system, or direct, and an empty string means automatic: HTTP(S)_PROXY if set, otherwise the host OS proxy. proxy.sandbox narrows that to sandbox egress only, and it is the one with an environment alias, DOCKER_SANDBOXES_PROXY.
Precedence and restart
"Environment variables take precedence over user overrides" help-sbx sbx settings set, so an exported DOCKER_SANDBOXES_PROXY beats sbx settings set proxy.sandbox. Above both sits the organization: "Administrator constraints apply to saved overrides. A conflicting value is rejected" help-sbx sbx settings set. sbx settings unset KEY removes the override, and "Administrator policy remains in effect" help-sbx sbx settings unset. After that "the setting evaluates from its environment variable, remote default, or built-in default" help-sbx sbx settings unset, and the remote default is a fourth origin that the SOURCE column does not name.
"Most changes take effect within about five seconds" help-sbx sbx settings, and the rest need sbx daemon restart. The table marks them in its RESTART column, and its footer says who needs the restart.
Some fields were truncated; use --no-trunc or --json to see them in full.
RESTART=yes: existing daemon-side consumers require `sbx daemon restart`. Supported CLI clients and new sandboxes use their current proxy settings immediately.Fifteen keys carry "requires_restart": true on the recording Mac: the four proxy* keys, the three no_proxy* keys, the six ssh.* keys, tls.allowNegativeSerial, and mcp.forceLocalGateway. The other sixteen apply within the five seconds. Some of them, such as skills.defaultMode and sandbox.disk.dockerVolume, say in their description that only sandboxes created after the change see the new value.
Where the CLI and the docs disagree
The docs settings page and the CLI list different keys, and the settings table prints only the keys the CLI returned. Four keys are in the CLI and not in the docs: ssh.autoCreate, ssh.defaultAgent with default shell, ssh.defaultTemplate, and ssh.workspaceRoot, as conflict C20 records. Four keys are in the docs and not in the CLI output: feature.model, feature.sandbox-gpu, feature.udp-egress, and diagnostics.autoUploadErrorCooldownInDays, as conflict C19 records. The last one has a documented default of 1 docs-sbx diagnostics.autoUploadErrorCooldownInDays. The capture ran sbx settings get only on skills.defaultMode and kit.allowedSources, so what get answers for a feature.* key is open.
One default disagrees. The docs tell you to run sbx settings set platform.allowExperimentalFeatures true before feature.model docs-sbx Enable model selection, which implies false. The capture prints true with source default, as conflict C18 records.
{
"default": true,
"description": "Allow experimental features.",
"env_var": "DOCKER_SANDBOXES_ALLOW_EXPERIMENTAL_FEATURES",
"key": "platform.allowExperimentalFeatures",
"source": "default",
"type": "bool",
"value": true
}The same gap covers commands. The model.providers description names sbx run --model <model> --provider <id>, and the docs show sbx run --name <SANDBOX_NAME> --model <MODEL_NAME> --provider <PROVIDER_ID> docs-sbx configuration/models, but sbx run --help in v0.47.0 lists neither flag. The secret help text names "mounts added later with sbx mount" help-sbx sbx secret set, and the help tree has no sbx mount page. sbx ssh proxy, sbx policy approval, and --usb are named in the sources that conflict C16 lists, and none of them has a help page either. This manual documents what --help prints, and the command table is that list.
The kit defaults
Six kit.* keys decide which kits a sandbox accepts, and all six print default as their source. kit.allowedSources is ["docker.io/"], and kit.allowLocalKits and kit.allowExtractedAgents are true. kit.requireSignature and kit.ignoreTransparencyLog are false, and kit.trustedSigners trusts any @docker.com identity through https://accounts.google.com. Kit signing shows what a signature check does with them.
Sources:help-sbx sbx settings, sbx settings list, sbx settings set, sbx settings unset, sbx secret set (research/sources/help-sbx.md); docs-sbx configuration/settings, configuration/models (research/sources/docs-sandboxes.md); conflicts C16, C18, C19, C20 (research/conflicts-register.md); research/sources/probes/sbx-settings-list.txt; capture/out/02-settings.json, 02-settings.txt, 02-settings-get.txt, 02-settings-experimental.json
sbx policy, sbx secret, sbx mcp
Five layers isolate the agent, and each layer has its own commands: the hypervisor, the workspace, the policy and its proxy, the secret store, and the MCP gateway.
sbx diagnose and the microVM boundary
The agent runs as a container inside a guest kernel on the host hypervisor, with a private Docker Engine, and nothing it does reaches the host daemon.
You started sbx run claude on a repository, and the agent is running docker build inside. On the host, docker ps shows nothing new, and you want to know where those images went and what else the agent can reach. sbx diagnose is the first command to run, because it names every layer between the agent and your machine.
When you finish this section, you can read sbx diagnose, name the five layers between the agent and the host, and name the doors through them.
What sbx diagnose checks
sbx diagnose: a read-only check of the installation, the platform, the storage, and the connection to the daemon, with --output json|github-issue and --upload help-sbx sbx diagnose. On the capture Mac it ran 13 checks, and all passed:
Platform
✓ Virtualization — supported
kern.hv_support is 1
✓ mkfs.erofs — found
/opt/homebrew/Caskroom/sbx/0.47.0/Sbx.app/Contents/libexec/mkfs.erofs, default block size <n> bytes, guest kernel page size <n> bytes
…
Connection
✓ Version match — v0.47.0
✓ Socket — responsive
✓ SSH client config — not configured
✓ Authentication — authenticated
…
13 passedkern.hv_support is 1 is the macOS flag for hardware virtualization. The guest runs on the host hypervisor: Hypervisor.framework on macOS, Windows Hypervisor Platform on Windows, and KVM on Linux (conflict C3). The mkfs.erofs line names the guest root filesystem format and its page size, which getconf PAGESIZE inside reports as 16384 (conflict C45). The last line is the sign-in check. sbx runs without Docker Desktop, and it still refuses to create a sandbox without a Docker account (conflict C15).
The FAQ gives Docker's reasons: "Tie sandboxes to a real person" and "Authenticate against Docker infrastructure" docs-sbx FAQ. It also lists the product's own egress hosts, starting with login.docker.com.
Five layers, from the kernel up
Inside m101-demo, the capture ran uname, id, docker version, and mount:
Linux m101-demo 7.0.14 #1 SMP PREEMPT Mon Sep 21 06:45:23 UTC 2026 aarch64 GNU/Linux
…
PRETTY_NAME="Ubuntu 26.04.1 LTS"
…
uid=1000(agent) gid=1000(agent) groups=1000(agent),27(sudo),1001(docker)
16384
…
Server: Docker Engine - Community
Engine:
Version: 29.8.1
…
containerd:
Version: v2.3.5
…
bind-<id> on /etc/resolv.conf type virtiofs (ro,relatime)
host on $CAPTURE/fixtures/repo type virtiofs (rw,nosuid,nodev,relatime)The docs name five isolation layers: hypervisor, network, Docker Engine, workspace, and credential docs-sbx Isolation layers. The listing shows four of them from inside. The guest kernel is 7.0.14 on aarch64 with 16 KiB pages, on Ubuntu 26.04.1. A build that assumes 4 KiB pages fails here, as one Docker Captain reported on 2026-05-26 blog 2026-05-26 B16. The agent is user agent, uid 1000, in the sudo and docker groups, and the docs place the control elsewhere:
Docker Engine 29.8.1 and containerd v2.3.5 in the listing belong to the VM, and so does every image the agent builds. The docs state the consequence in one sentence: "The agent has no path to your host Docker daemon." docs-sbx Docker Engine isolation
The host side has no Docker API for the sandbox either. Kevin Wittek said on 2026-04-23 that the Moby API is not offered from the host talk 2026-04-23 T06. REST clients that expect it do not work against a sandbox, and sbx exec is the way in, as sbx exec, cp, and ports shows.
VMM: the virtual machine monitor that starts the guest. Docker wrote its own and said on Hacker News on 2026-08-10 that it is not Firecracker (conflict C4). The claim that it builds on libkrun stays unverified. The names the product exposes are few. sandboxd is the daemon, whose socket sbx daemon status prints. containerd runs inside the guest, EROFS appears in the diagnose line, virtiofs in the mount table, and nerdbox in the release assets.
The doors through the boundary
Kevin Wittek named four entry points into a sandbox on 2026-09-04: bind mounts, network, secret injection, and MCP talk 2026-09-04 T02. The capture shows each as one line in the guest. The workspace is host on $CAPTURE/fixtures/repo type virtiofs (rw,...), a mount at the same path as on the host. /etc/resolv.conf is a second, read-only virtiofs bind from the host, so the resolver is the host's. The network door is HTTPS_PROXY=http://gateway.docker.internal:3128 in the environment (capture/out/04-env.txt). The secret door is ANTHROPIC_API_KEY=proxy-managed, a sentinel.
The MCP door is MCP_GATEWAY_URL=http://mcp-gateway.docker.internal/mcp. A fifth line, SSH_AUTH_SOCK=/run/ssh-agent.sock with SSH_AUTH_SOCK_GATEWAY=gateway.docker.internal:3129, is the forwarded SSH agent, a door that the talk did not name. The docs turn that forwarding on by default docs-sbx Credential isolation, and secrets returns to it. Figure 3.1 stacks the layers and draws the doors.
The VM has limits of its own. --memory defaults to 50% of host memory, clamped to 512 MiB to 32 GiB help-sbx sbx create. --cpus 0 means all host CPUs, at most 16 on Linux arm64 help-sbx sbx create. m101-policy got cpu 10 and memory 32 GiB on the capture Mac (capture/out/09-create.txt).
Sources:help-sbx sbx diagnose, sbx create, sbx daemon start (research/sources/help-sbx.md); docs-sbx Security model, Isolation layers, Default security posture, Architecture, FAQ (research/sources/docs-sandboxes.md); research/conflicts-register.md rows C3, C4, C15, C45; talk T02 (2026-09-04), T06 (2026-04-23), blog B16 (2026-05-26), HN01 (2026-08-10) from research/plan.md; capture/out/02-diagnose.txt, 02-diagnose.json, 02-daemon-status.txt, 04-guest.txt, 04-env.txt, 09-create.txt
--clone and /run/sandbox/source
Bind mode gives the agent your files, :ro and mountless modes give it less, and --clone gives it a private clone whose commits come back through a git remote.
You let an agent work on a repository overnight. In the morning git diff is clean, and a new file sits in .git/hooks/pre-commit, where no diff ever shows it. The three workspace modes decide how much of your tree the agent can write. With --clone it can write none of it.
When you finish this section, you can pick a workspace mode for a repository and fetch an agent's commits out of a clone-mode sandbox.
Three workspace modes
Direct mount: the workspace path mounted inside the VM at the same absolute path, read-write help-sbx sbx create claude. In m101-demo the agent's pwd is $CAPTURE/fixtures/repo, the same string as on the host, and git log lists the host commits c7dfd56 second and 086df68 first (capture/out/04-workspace.txt). The mount is virtiofs, as the microVM boundary showed, and the file synchronization of the legacy plugin is gone (conflict C5). Extra paths follow the first one, and :ro makes one read-only: "a read-only argument may name a single file, which holds that one path out of reach inside a workspace the sandbox can otherwise write" help-sbx sbx create claude.
Mountless: no path on sbx create, so the VM has no host bind mount. The agent works in the template's working directory, /home/agent/workspace for Docker's templates docs-sbx Workspace isolation. m101-policy was created that way, and sbx create printed workspace none · no workspace bind mount (capture/out/09-create.txt). sbx run without a path mounts the current directory instead, which sbx run, create, stop, and rm warns about.
Direct mode needs a review step. The docs list what the agent can edit, including Git hooks, CI configuration, package.json scripts, and .claude/settings.json, and they warn that hooks "don't appear in git diff output" docs-sbx Workspace isolation. A Docker Captain described this covert channel on 2026-05-26, and it is the reason clone mode exists blog 2026-05-26 B16.
Clone mode
--clone: a create-time flag that makes the agent "Run the agent on a private in-container clone of the host Git repository (mounted read-only)" help-sbx sbx create. It replaced --branch in v0.31.0, and --branch now fails with --branch is no longer supported; use --clone instead rel-sbx v0.31.0 (conflict C6). The flag is a no-op when you re-attach to an existing clone-mode sandbox help-sbx sbx run. The capture created m101-clone with it:
$ sbx create --clone shell $CAPTURE/fixtures/repo-clone --name m101-clone
✓ Git repository detected: $CAPTURE/fixtures/repo-clone
…
Git daemon: git://127.0.0.1:<port>/repo-clone
Remote: sandbox-m101-clone
✓ Created sandbox m101-clone
mount $CAPTURE/fixtures/repo-clone → /run/sandbox/source (ro, source)
…
SANDBOX AGENT STATUS PORTS WORKSPACE
m101-clone shell running 127.0.0.1:<port>->9418/tcp4 $CAPTURE/fixtures/repo-cloneInside, three facts stand out (capture/out/08-clone-inside.txt). The clone sits at the same path, $CAPTURE/fixtures/repo-clone, on /dev/vde, an ext4 volume, and its origin is /run/sandbox/source. The source mount is host on /run/sandbox/source type virtiofs (ro,nosuid,nodev,relatime), and touch /run/sandbox/source/x fails with Read-only file system. A commit made inside, 81e11df from sandbox, cannot be pushed to origin either: git push ends with remote unpack failed: unable to create temporary object directory (capture/out/08-clone-commit.txt). The docs name the limit: clone mode "protects your host repository from modification" docs-sbx Clone mode, and inspection stays open, so an untracked .env under the repository is readable inside.
Getting commits back
The git-daemon in the VM listens on 9418, and sbx publishes it on a loopback port. The CLI then writes a remote into the host repository's .git/config:
$ git -C $CAPTURE/fixtures/repo-clone remote -v
sandbox-m101-clone git://127.0.0.1:<port>/repo-clone (fetch)
sandbox-m101-clone git://127.0.0.1:<port>/repo-clone (push)
…
file:.git/config remote.sandbox-m101-clone.fetch=+refs/heads/*:refs/remotes/sandbox-m101-clone/*
file:.git/config remote.sandbox-m101-clone.fetch=+refs/heads/*:refs/sandboxes/m101-clone/*
…
$ git -C $CAPTURE/fixtures/repo-clone fetch sandbox-m101-clone
…
c7dfd56..81e11df main -> sandbox-m101-clone/main
c7dfd56..81e11df main -> refs/sandboxes/m101-clone/main
…
81e11df from sandbox
c7dfd56 second
086df68 firstTwo fetch refspecs mean one fetch updates two places: refs/remotes/sandbox-m101-clone/main and refs/sandboxes/m101-clone/main. sbx rm removes the remote and the daemon and prints the recovery command git branch <local-name> refs/sandboxes/m101-clone/<branch>. In the capture, refs/sandboxes/m101-clone/main survived the removal (capture/out/08-clone-rm.txt). sbx stop stops the daemon, and a restart assigns a new port and rewrites the remote URL docs-sbx Use Git with sandboxes. The /root/.config/git/attributes warning in the fetch output comes from the daemon's user inside the VM and changes nothing. sbx kit add keeps the clone "via a named workspace volume" help-sbx sbx kit add, which is the ext4 device above, and sbx rm "cleans up any Git worktrees" help-sbx sbx rm.
Figure 3.2 follows one commit from the clone to the host refs.
Sources:help-sbx sbx create, sbx create claude, sbx run, sbx rm, sbx kit add (research/sources/help-sbx.md); docs-sbx Isolation layers, Architecture, Use Git with sandboxes, Usage (research/sources/docs-sandboxes.md); rel-sbx v0.31.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C5, C6; blog B16 (2026-05-26) from research/plan.md; capture/out/04-workspace.txt, 08-clone-create.txt, 08-clone-inside.txt, 08-clone-commit.txt, 08-clone-host.txt, 08-clone-rm.txt, 09-create.txt
sbx policy init, allow network, deny network, and check network
A global preset plus allow and deny rules in two scopes decide every connection, deny wins, and check network asks the same authorizer without sending anything.
Your agent reports that npm install failed with a connection error, and the sandbox keeps no shell history to say why. The question is which rule decided. sbx policy answers it in two halves: ls for the rules that exist, and check network for the decision one host would get.
When you finish this section, you can initialize a policy, add a rule in the right scope, and predict the decision for any host.
The preset and the global policy
policy init: sets "the initial global policy, not a per-sandbox default" help-sbx sbx policy init. It runs once, before the first sandbox, with allow-all, balanced, or deny-all, and sbx policy reset or sbx daemon start --policy starts over help-sbx sbx daemon start. The capture host chose balanced on v0.47.0, and sbx policy ls then showed one policy, local-policy, with network: 194 allow (capture/out/03-policy-ls.txt). The wide JSON lists eight rules with created_via: default: six network groups and two filesystem rules that allow every path. The list changes between releases without a changelog (conflict C23), so this is the list as recorded on 2026-10-08:
"id": "default-ai-services",
"name": "default-ai-services",
"policy_id": "local-policy",
"scope": "global",
"applies_to": "all",
"resource_type": "network",
"decision": "allow",
"resources": [
"api.anthropic.com:443",
"statsig.anthropic.com:443",
"platform.claude.com:443",
…
"**.openai.com:443",
…
"id": "default-package-managers",
…
"registry.npmjs.org:443",
…
"pypi.org:443",
…
"provenance": {
"created_via": "default"
},
"actions": [
"net:connect:tcp"
]Rule grammar and scope
Allow rule: a comma-separated list "of hostnames, domains, IP addresses, or CIDR prefixes" help-sbx sbx policy allow network. The forms are example.com, *.example.com, **.example.com, api?.example.com, api[12].example.com, example.com:443, [2001:db8::1]:443, a CIDR, and ** for every host. A lone * and an escaped glob are rejected help-sbx sbx policy allow network. An allow rule covers TCP unless --protocol says otherwise, and a deny rule covers TCP and UDP help-sbx sbx policy deny network. The capture shows both: the allow came back as (example.com [tcp]) and the deny as (example.com [tcp,udp]) (capture/out/09-allow.txt, 09-deny.txt). The deny also warned that UDP egress stays off until feature.udp-egress is on.
Two scopes have existed since v0.29.0 (conflict C41): global, the default, and local, which --sandbox scopes to one sandbox help-sbx sbx policy allow network. A rule added with --sandbox m101-policy printed Rule added to policy local (scope: sandbox:m101-policy): <uuid>, and sbx policy inspect shows its scope, layer, origin, and provenance:
"id": "<uuid>",
"name": "<uuid>",
"policy_id": "<uuid>",
"scope": "sandbox:m101-policy",
"applies_to": "sandbox:m101-policy",
"resource_type": "network",
"decision": "allow",
"resources": [
"example.com"
],
"origin": "scoped",
"layer": "local",
"status": "active",
"editable": true,
"sandbox_id": "m101-policy",
"provenance": {
"created_via": "added"
},
"actions": [
"net:connect:tcp"
]--deny-network HOST on create or run adds the same kind of per-sandbox deny at creation. The help says why that is safe under governance: "a local deny can only narrow, never widen, egress" help-sbx sbx create. A kit adds rules with created_via: provisioned, and sbx policy ls --source org on the capture host printed No policies match the selected filters. (capture/out/03-policy-org.txt).
Deny wins: "Deny rules take precedence over allow rules for the same hostname or CIDR. An allowed hostname isn't checked against CIDR rules for its resolved IP address." help-sbx sbx policy deny network
The CLI refuses a deny that conflicts with an allow in the same scope: deny: "example.com" conflicts with existing allow rule "<uuid>" (capture/out/09-deny-conflict.txt). The capture removed the allow with sbx policy rm network --sandbox m101-policy --id <uuid> --force before the deny was accepted. Under organization governance only organization allow rules grant access, while local and kit deny rules still apply docs-sbx Precedence, and org policies prints that table.
policy ls filters with --source local|org|kit, --decision, --type, --created-via default|added|provisioned|approval, --protocol, --wide for RULE_ID, and --json help-sbx sbx policy ls. A wide row carries METHOD and PATH columns. The docs fill them with --method and --path flags. The v0.47.0 help of sbx policy allow network lists neither flag, so this manual marks local HTTP rules as unverified. The proxy returns to them.
check network
policy check network: a read-only call that "evaluates the same daemon-side policy authorizer used by sandbox network enforcement" help-sbx sbx policy check. A host without a port is evaluated with port 443, and the command "evaluates network authorization, not HTTP method or path" help-sbx sbx policy check network. It exits 1 on a denial. The capture ran it on example.com in the m101-policy context before any rule, after the allow, and after the deny:
{
"action": "net:connect:tcp",
"allowed": false,
"context": "sandbox:m101-policy",
"deny_kind": "explicit",
"governance": {
"active": false
},
"origin": "local",
"reason": "Denied by local rule",
"resource_type": "net:domain",
"resource_value": "example.com:443",
"rule": "local:<uuid>",
"target": "example.com:443",
"type": "network"
}After the allow, the same call returned "allowed": true with no rule field (capture/out/09-check-allowed.json). The global check of api.anthropic.com did the same under "context": "global" (capture/out/03-policy-check-anthropic.json). The two denials differ in deny_kind: implicit names no rule, and explicit names local:<uuid>. An implicit denial is also what a sandbox turns into an approval request, which the next section shows. Figure 3.3 walks the request through the same order.
Sources:help-sbx sbx policy init, sbx daemon start, sbx policy allow network, sbx policy deny network, sbx policy ls, sbx policy inspect, sbx policy rm network, sbx policy check, sbx policy check network, sbx create (research/sources/help-sbx.md); docs-sbx Policy concepts, Local policy, Monitoring policies (research/sources/docs-sandboxes.md); research/conflicts-register.md rows C23, C41; capture/out/03-policy-init.txt, 03-policy-ls.txt, 03-policy-org.txt, 03-policy-balanced.json, 03-policy-check-anthropic.json, 09-allow.txt, 09-deny.txt, 09-deny-conflict.txt, 09-inspect-rule.json, 09-rm-allow.txt, 09-check-verbose.json, 09-check-allowed.json, 09-check-deny.json, 12-policy-kit.txt
sbx policy log and the proxy
Every connection leaves through a host proxy, which answers a blocked HTTPS request with its own certificate and writes one policy log row per host with the matching rule.
The agent says that https://example.com timed out. It did not time out: the proxy answered it in 90 bytes. Every connection from a sandbox passes a proxy on the host. sbx policy log keeps one row per host with the decision, the proxy path, and the rule.
When you finish this section, you can read a policy log row, name its proxy path, and tell a blocked host from a slow one.
Two proxies and one CA
Inside m101-demo, the environment carries the proxy address and its certificate:
HTTPS_PROXY=http://gateway.docker.internal:3128
HTTP_PROXY=http://gateway.docker.internal:3128
…
NODE_EXTRA_CA_CERTS=/etc/ssl/certs/ca-certificates.crt
NODE_USE_ENV_PROXY=1
NO_PROXY=localhost,127.0.0.1,::1,gateway.docker.internal
…
PROXY_CA_CERT_B64=<base64>
…
REQUESTS_CA_BUNDLE=/etc/ssl/certs/ca-certificates.crt
…
SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crtThe address is gateway.docker.internal:3128, not the host.docker.internal:3128 of the legacy plugin (conflict C26). The docs describe two paths: "Agents use a forward proxy for HTTP and HTTPS; other TCP traffic is forwarded transparently. Both paths enforce network access policies." docs-sbx Networking
The first prototype was only an environment variable, and an agent bypassed it with no_proxy, as Kevin Wittek said on 2026-01-14 talk 2026-01-14 T01. The capture sent a request to gateway.docker.internal:18080, a name in NO_PROXY, so curl skipped the forward proxy and opened a plain TCP connection. The transparent path caught that connection, and the log recorded it as transparent and blocked (capture/out/10-policy-log.json).
TLS interception: the proxy's own certificate authority, carried as PROXY_CA_CERT_B64 and merged into /etc/ssl/certs/ca-certificates.crt rel-sbx v0.35.0. The capture shows where it is used:
$ sbx exec m101-policy sh -c 'curl -sS --max-time 10 -I https://example.com; echo exit=$?'
HTTP/1.1 200 OK
HTTP/1.1 403 Forbidden
Content-Length: 90
Content-Type: text/plain
exit=0
[exit 0]
$ sbx exec m101-policy sh -c 'curl -sS --max-time 10 https://example.com; echo; echo exit=$?'
Approval required for example.com:443.
Review and respond with:
sbx policy approval ls
exit=0
[exit 0]HTTP/1.1 200 OK is the proxy accepting the CONNECT. The 403 Forbidden that follows arrived inside the TLS session, under a certificate the sandbox trusts, and its 90-byte body is the approval message. After sbx policy allow network --sandbox m101-policy example.com, the same request got HTTP/1.0 200 Connection established and then HTTP/2 200 with server: cloudflare (capture/out/09-allowed.txt).
The log classed that one as forward-bypass, a tunnel without inspection and without credential injection docs-sbx Monitoring policies. So the ruling for C26 reads: the proxy terminates TLS only when it has to speak. That is a blocked host or a host with a bound credential, and it tunnels the rest. The v0.47.0 notes name both paths, a "non-MITM CONNECT tunnel" and "the transparent proxy's late handshake check" rel-sbx v0.47.0.
Reading policy log
policy log: shows "which hosts were allowed or blocked by the proxy, along with the matching rule, proxy type, and request count" help-sbx sbx policy log. It takes [SANDBOX], --json, --limit, and --type, and the help admits that "filesystem logs are not supported yet" (conflict C28). The JSON has two arrays:
{
"blocked_hosts": [
{
"host": "example.com:443",
"vm_name": "m101-policy",
"proxy_type": "forward",
"rule": "denied: rule \"local:<uuid>\" matched op(action=net:connect:tcp, resource=net:domain:example.com:443)",
"last_seen": "<ts>",
"since": "<ts>",
"count_since": "<n>",
"reason": "Denied by local rule"
},
…
"allowed_hosts": [
…
{
"host": "example.com:443",
"vm_name": "m101-policy",
"proxy_type": "forward-bypass",
"rule": "",
"last_seen": "<ts>",
"since": "<ts>",
"count_since": "<n>"
},Each row has host, vm_name, proxy_type, rule, reason, last_seen, since, and count_since. A blocked row names the operation, op(action=net:connect:tcp, resource=net:domain:example.com:443), and either no applicable policies or the rule that matched. An allowed row in this capture carries an empty rule, so the log says that a host passed, and check network says why. PROXY takes five values, forward, forward-bypass, transparent, network, and browser-open docs-sbx Monitoring policies. Figure 3.4 animates both requests.
The 403 body names sbx policy approval ls. The docs describe approval ls, inspect, and respond. Under balanced and deny-all, a request no rule matches "asks for your approval instead of being denied outright" docs-sbx Local policy. The v0.47.0 help tree has no sbx policy approval command, while sbx policy ls --created-via approval exists (conflict C16).
What a host rule cannot see
A deny covers more than TCP (conflict C27). A 2026-05-26 post said UDP and ICMP are blocked and cannot be allowed blog 2026-05-26 B16. Since v0.33.0 a sandboxed process cannot resolve a domain that policy denies, loopback names excepted, and outgoing ICMP stays blocked rel-sbx v0.33.0. Since v0.45.0 UDP follows policy behind the experimental setting feature.udp-egress, and DNS resolution stops when no rule permits it rel-sbx v0.45.0. The ruling: a deny covers TCP, UDP, and the name lookup itself, and ICMP cannot be allowed. The guest's /etc/resolv.conf is a read-only bind from the host.
A host allow permits every request to that host, with any method, path, and body. The balanced preset allows github.com:443 and **.github.com:443 (capture/out/03-policy-balanced.json). Docker staff said on 2026-04-07 and on 2026-08-10 that an issue body or a gist on an allowed host is not blocked talk 2026-04-07 T05.
The product's answer is the HTTP rule. A kit's network-policy@2 entry can deny hosts: [api.github.com] with methods: [DELETE] kitcap network-policy@2. A network allow is a ceiling that HTTP rules carve into docs-sbx HTTP method and path, and neither check network nor policy log evaluates a method or a path docs-sbx Local policy. Figure 3.5 puts the two rules side by side.
Upstream proxies are a separate setting: proxy, proxy.sandbox, proxy.daemon, and the no_proxy family, with SOCKS5 since v0.35.0 rel-sbx v0.35.0. The settings table lists the keys.
Sources:help-sbx sbx policy log, sbx policy allow network (research/sources/help-sbx.md); docs-sbx Architecture, Isolation layers, Default security posture, Monitoring policies, Local policy, Network access policies, Policy concepts, Upstream proxy (research/sources/docs-sandboxes.md); kitcap network-policy@2 (research/sources/kit-capabilities.md); rel-sbx v0.33.0, v0.35.0, v0.45.0, v0.47.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C16, C26, C27, C28; talk T01 (2026-01-14), T05 (2026-04-07), blog B16 (2026-05-26), HN01 (2026-08-10) from research/plan.md; capture/out/04-env.txt, 09-blocked.txt, 09-allowed.txt, 09-allow.txt, 09-policy-log.json, 09-policy-log-after.json, 09-policy-log-deny.json, 09-policy-log.txt, 10-policy-log.json, 03-policy-balanced.json
sbx secret set, sbx secret import, and sbx secret set-custom
A secret never enters the VM: the agent sees a sentinel, the host keychain holds the value, and the proxy swaps it into requests to the bound host.
You exported ANTHROPIC_API_KEY in ~/.zshrc, as a 2026 tutorial said, and the agent inside the sandbox still has no key. Since v0.35.0 the host environment is not read rel-sbx v0.35.0 (conflict C7). -e KEY passes a plain variable, not a secret, and the three sbx secret commands are the only way in.
When you finish this section, you can store a service secret, bind a custom one to a host, and read what the proxy sent.
Service secrets
Service secret: a value stored under one of 13 service names, anthropic, copilot, cursor, devin, droid, github, google, groq, mistral, nebius, openai, openrouter, xai help-sbx sbx secret set, two more than the docs table (conflict C21). The value comes from -t, from stdin, from --command, or from --ref with a 1Password op:// reference or an AWS Secrets Manager ARN. --oauth is the fifth source, and locally it is "openai/global only" help-sbx sbx secret set (conflict C22). Anthropic OAuth comes from the agent's own login instead. A command helper runs "from a fresh temporary directory on the host" help-sbx sbx secret set since v0.46.0, so a relative helper path no longer resolves (conflict C38). --refresh sets the cache time, default 55 minutes.
The scope is global unless --sandbox narrows it. secret ls filters with -g, --sandbox, --service, and --json, and secret rm takes --all, --placeholder, and --registry help-sbx sbx secret rm. The store is the macOS Keychain, the Windows Credential Manager, or the Linux Secret Service docs-sbx Where secrets are stored. Without a keyring, Linux uses a file under ~/.config/com.docker.sandboxes at mode 0700. Kit approvals live apart from the values, in ~/.config/sbx/credentials.yaml docs-sbx Credential bindings.
secret import: reads the host environment once, offers each variable with a last-4 preview, and takes --all, --force, and --dry-run help-sbx sbx secret import. A service that already holds an OAuth token is skipped. On the capture host nothing was exported:
$ sbx secret import --dry-run
No credential env vars detected on the host. Set e.g. OPENAI_API_KEY in your shell and re-run, or use `sbx secret set` to enter a value directly.Registry credentials are a third kind, "host-only by default", and reach a sandbox only with --all-sandboxes or --sandbox help-sbx sbx secret set.
What the agent sees
Every service variable inside a sandbox is a sentinel, and the GitHub one is shaped like a token:
ANTHROPIC_API_KEY=proxy-managed
…
GH_TOKEN=gho_sbxproxymanaged000000000000000000000
…
OPENAI_API_KEY=proxy-managed
…
SBX_CRED_ANTHROPIC_MODE=none
SBX_CRED_GITHUB_MODE=noneThree formats exist in v0.47.0 (conflict C35): proxy-managed for the eight provider keys, gho_sbxproxymanaged000000000000000000000 for GH_TOKEN, and sbx-cs-<rand> for custom secrets. The GitHub one is lower case, where a 2026-09-04 demo showed GHO_SBX_PROXY_MANAGED talk 2026-09-04 T02. No GitHub request was captured: that needs a real token.
Custom secrets and the receiver
set-custom: an experimental secret keyed to --host targets and an --env name instead of a service help-sbx sbx secret set-custom. The value comes from --value, --command, or --ref, and --placeholder sk-{rand} sets a chosen prefix. The --header and --format flags apply with --cloud only, as the help states for each (conflict C34). The capture bound M101_RECV_KEY to host.docker.internal and localhost, and started a receiver on the host at 127.0.0.1:18080. It allowed localhost:18080 for m101-secret and found M101_RECV_KEY=sbx-cs-<rand> inside (capture/out/10-env-sentinel.txt). Then it sent three plain HTTP requests:
GET /from-variable HTTP/1.1
Host: localhost:18080
User-Agent: curl/8.18.0
Accept: */*
Authorization: Bearer m101-dummy-receiver-0000
X-Demo: m101-dummy-receiver-0000
Accept-Encoding: gzip
…
GET /no-scheme HTTP/1.1
Host: localhost:18080
…
Authorization: m101-dummy-receiver-0000The receiver saw the real value in every place the placeholder appeared: inside Bearer, in X-Demo, and as the whole Authorization value. Its Host was localhost:18080: the proxy maps host.docker.internal to localhost docs-sbx Accessing host services from a sandbox. So the swap happens on plain HTTP, and it is a substring replacement. That settles conflict C34 for custom secrets. The Hacker News claim that the secret must be the whole header describes service secrets, where the kit declares the header and its format docs-sbx Services declared by kits.
The fourth request went to gateway.docker.internal:18080, a name in NO_PROXY, so it bypassed the forward proxy. The transparent proxy blocked it with Empty reply from server and swapped nothing (capture/out/10-swap-curl.txt). Figure 3.6 animates the four steps. The existing sandbox m101-demo did not receive M101_RECV_KEY (capture/out/10-env-existing.txt), so a custom variable reaches new sandboxes only. Removal is immediate: sbx secret rm --placeholder sbx-cs-<rand> -f printed Applied secret updates for <n> running sandbox(es) (capture/out/10-secret-rm.txt).
The SSH agent
One credential path does cross the boundary. "SSH agent forwarding is enabled by default" docs-sbx Credential isolation, and the guest shows SSH_AUTH_SOCK=/run/ssh-agent.sock with SSH_AUTH_SOCK_GATEWAY=gateway.docker.internal:3129 (capture/out/04-env.txt). Keys stay on the host, and any process inside can ask that agent to sign. ssh.agentForwardingEnabled in sbx settings list turns it off, followed by sbx daemon restart docs-sbx SSH agent. MCP secrets stay on the host too, under mcp:<server>:client_secret, as sbx mcp add explains.
Sources:help-sbx sbx secret, sbx secret set, sbx secret import, sbx secret ls, sbx secret rm, sbx secret set-custom (research/sources/help-sbx.md); docs-sbx Manage credentials, Isolation layers, Usage (research/sources/docs-sandboxes.md); rel-sbx v0.35.0, v0.46.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C7, C21, C22, C34, C35, C38; talk T02 (2026-09-04), HN01 (2026-08-10) from research/plan.md; capture/out/04-env.txt, 10-set-custom.txt, 10-secret-ls.json, 10-env-existing.txt, 10-secret-create.txt, 10-env-sentinel.txt, 10-swap-curl.txt, 10-receiver.log, 10-policy-log.json, 10-placeholder.txt, 10-import-dry-run.txt, 10-secret-rm.txt
sbx mcp add, sbx mcp load, and --static-mcp
MCP servers are registered on the host, served to the sandbox through one gateway endpoint, and either fixed at creation or loaded live with a tools/list_changed notice.
You want the agent to read documentation through an MCP server, and you do not want that server's token inside the VM. sbx mcp keeps the registration, the OAuth tokens, and any local server process on the host, and gives the sandbox one URL.
When you finish this section, you can register a server, pick static or dynamic mode, and load a server into a running sandbox.
Register on the host
sbx mcp add: "Register an MCP server by name. The server is validated and its specification is stored for use with sbx create/run --static-mcp." help-sbx sbx mcp add
--url takes four forms: a remote endpoint, a community-registry URL, a server.json or server.yaml manifest URL, and a dhi.io/ image reference. Other image references are rejected, and --local runs a registry OCI server on the host with docker run help-sbx sbx mcp add. The capture registered one remote server and listed it:
{
"gateway": {
"name": "LOCAL",
"local": true,
"operator": "managed by you",
"decision": "local",
"signed_in_as": "<docker-user>"
},
"servers": [
{
"name": "m101-deepwiki",
"transport": "remote http",
"status": "ready",
"type": "remote"
}
]
}The help's own registry example did not resolve on 2026-10-08: the add of https://registry.modelcontextprotocol.io/v0/servers/fetch-mcp/versions/latest failed with registry returned status 404 (capture/out/11-mcp-add-registry.txt). --command is the other input, and it runs on the host: "The process runs with your host user's full permissions" help-sbx sbx mcp add. A local server that starts a container uses host Docker, outside the engine boundary of the microVM docs-sbx Local stdio server.
OAuth metadata is discovered through RFC 9728 and RFC 8414, or supplied by hand with --oauth-authorization-server. --client-id names a pre-registered client, and dynamic registration follows RFC 7591. --scope records default scopes with a documented precedence, --resource sets the RFC 8707 indicator, and --callback-port pins the listener help-sbx sbx mcp add. Tokens stay on the host, and sbx mcp auth [server|--all], auth status, and auth rm manage them help-sbx sbx mcp auth.
The deepwiki server needs none: sbx mcp inspect reports requires_oauth: false, and auth status --all --json printed [], so no OAuth flow was captured. A confidential client's secret and any --header 'Name: ${placeholder}' value come from the secret store as mcp:<server>:client_secret and mcp:<server>:<placeholder>. A header-bearing registration is rejected on the hosted gateway help-sbx sbx mcp add.
One gateway per sandbox
MCP gateway: a host-side endpoint that "brokers access to registered MCP servers" docs-sbx MCP gateway. Inside m101-mcp the environment holds MCP_GATEWAY_URL=http://mcp-gateway.docker.internal/mcp and MCP_SENTINEL_TOKEN_NAME=proxy-managed (capture/out/11-static-inside.txt). The docs list the agents that read that URL at start: Claude Code, Codex, Devin, Gemini, Kiro, and OpenCode docs-sbx Prerequisites. Docker Agent is absent from that list. A plain shell sandbox reached the gateway with curl all the same (capture/fixtures/mcp-probe.sh):
…
Content-Type: text/event-stream
…
Mcp-Session-Id: <session>
…
message
<event-id>
{"jsonrpc":"2.0","id":1,"result":{"capabilities":{"logging":{},"prompts":{"listChanged":true},"resources":{"listChanged":true},"tools":{"listChanged":true}},"instructions":"### m101-deepwiki…"protocolVersion":"2025-11-25","serverInfo":{"name":"mcp-gateway-m101-mcp","version":"0.1.0"}}}The gateway names itself mcp-gateway-m101-mcp, one per sandbox, answers protocol version 2025-11-25, and declares tools.listChanged: true. sbx mcp ls adds a GATEWAY column, LOCAL, managed by you, with signed_in_as in the JSON help-sbx sbx mcp ls. mcp.forceLocalGateway, default false, and SBX_MCP_URL=none select the local data plane when an account would otherwise use the hosted gateway docs-sbx mcp.forceLocalGateway.
The docs separate it from the Desktop product: "You don't need the Docker Desktop MCP Toolkit to use sbx mcp" docs-sbx MCP gateway. Network policy does not apply to a server the host registered, as Kevin Wittek said on 2026-09-04 talk 2026-09-04 T02. A server the agent starts inside the VM is subject to it. DMR, Compose models, and the MCP gateway service covers the Toolkit gateway.
Static and dynamic
Static mode: --static-mcp a,b at create or run fixes the set once at creation help-sbx sbx create. The static sandbox exposed the three deepwiki tools plus code-mode and mcp-exec, and no discovery tool. The dynamic sandbox started with discovery tools only, and sbx mcp load m101-deepwiki --sandbox m101-mcp-dyn added the three:
| Sandbox and moment | tools/list names | Capture |
|---|---|---|
m101-mcp, static | ask_wiki_question code-mode mcp-exec read_wiki_contents read_wiki_structure | capture/out/11-gateway-probe-static.txt |
m101-mcp-dyn, before load | code-mode mcp-add mcp-config-set mcp-exec mcp-find | capture/out/11-gateway-probe-dynamic-before.txt |
m101-mcp-dyn, after load | ask_wiki_question code-mode mcp-add mcp-config-set mcp-exec mcp-find read_wiki_contents read_wiki_structure | capture/out/11-gateway-probe-dynamic-after.txt |
$ sbx mcp load m101-deepwiki --sandbox m101-mcp-dyn
MCP server "m101-deepwiki" loaded into sandbox "m101-mcp-dyn" (live)The help promises the notice: "Connected agents see the new server's tools immediately via the standard MCP tools/list_changed notification" help-sbx sbx mcp load. The probe opened a new session for each run, so it recorded the two lists and not the notification itself. The built-in tools belong to the gateway: mcp-exec, code-mode, mcp-find, mcp-add, mcp-config-set, and <server>-authorize for OAuth servers docs-sbx Built-in gateway tools. In Cedar MCP policies those are MCP::Primordial resources and a server's tools are MCP::Tool resources docs-sbx Built-in gateway tools, which org policies returns to.
sbx mcp catalog was removed in v0.45.0 rel-sbx v0.45.0, and the sbx mcp enable of the product page never existed (conflict C11). Figure 3.7 draws the store, the two gateways, and the load. For the protocol itself, read the MCP lesson.
Sources:help-sbx sbx mcp, sbx mcp add, sbx mcp auth, sbx mcp inspect, sbx mcp load, sbx mcp ls, sbx create (research/sources/help-sbx.md); docs-sbx MCP gateway, Architecture, Security model, Settings (research/sources/docs-sandboxes.md); rel-sbx v0.45.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C11, C91; talk T02 (2026-09-04), T23 (2026-09-29) from research/plan.md; capture/fixtures/mcp-probe.sh; capture/out/11-mcp-add.txt, 11-mcp-add-registry.txt, 11-mcp-ls.json, 11-mcp-inspect.json, 11-mcp-auth-status.json, 11-static-create.txt, 11-static-inside.txt, 11-gateway-initialize.http, 11-gateway-probe-static.txt, 11-dynamic-create.txt, 11-gateway-probe-dynamic-before.txt, 11-load.txt, 11-gateway-probe-dynamic-after.txt, 11-mcp-rm.txt
Kits and sbxenv.yaml
A kit declares what a sandbox contains and can reach, and an environment file declares the sandbox and its secrets for approval before anything runs.
`# syntax=docker/sandbox-kit:3` and kit.yaml
A v3 kit is an ordinary OCI image whose manifest annotation carries a strict YAML descriptor, which sbx resolves when it creates the sandbox.
Your team wants every Codex sandbox to carry the same three tools, the same two allowed hosts, and the same instructions. A kit declares the tools, the hosts, and the credentials in one file, and the file travels as an image.
When you finish this section, you can read a v3 descriptor line by line and say what the frontend writes into the image. You can also tell which commands build, inspect, and run it.
One image, one annotation
Kit: one OCI image whose manifest annotation vnd.docker.sandbox.kit.descriptor carries the kit's declarations, while its layers carry the content kitspec §1. "a Kit pulls, inspects, and FROMs with stock tooling, and an engine that does not read the annotation runs it as an ordinary image" kitspec §1.
Workload: a kit whose layers are a root filesystem and whose image config carries the launch command. A composition has exactly one kitspec §1.
Mixin: a kit whose layers are an overlay applied on a workload's filesystem, zero or more per composition kitspec §1. A third kind, set, exists only while authoring: "publishing derives workload or mixin from the Kits it lists" kitspec §4.
The descriptor, line by line
The capture kit wrote one v3 descriptor, a shell workload with an inline recipe and a single network grant:
# syntax=docker/sandbox-kit:3
schemaVersion: "3"
kind: workload
displayName: hello-kit
description: A shell workload with one network grant, in the v3 descriptor form.
version: "1.0.0"
licenses: [Apache-2.0]
build: |
FROM docker/sandbox-templates:shell
COPY HELLO.md /home/agent/HELLO.md
capabilities:
- type: com.docker.sandbox/network-policy@2
config:
runtime:
allow:
- example.comThe frontend is "dispatched by the descriptor's first line, # syntax=docker/sandbox-kit:3" kitspec §1.1. schemaVersion must be exactly the string "3", and kind is workload, mixin, or set kitspec §4. displayName, description, version, and licenses are optional display metadata. No name field exists, because identity is the reference that a kit is consumed by kitspec §1.
The build: field "carries literal Dockerfile text" kitspec §3.2. capabilities is a list of entries with type and config kitspec §7. Here one network-policy@2 entry allows example.com in the runtime phase, the agent's steady state kitcap network-policy@2.
Decoding is strict: "Any unrecognized field anywhere in the document is an error" kitspec §1.2. The reason is policy: "A misspelled key silently ignored would be a policy silently absent" kitspec §1.2. Figure 4.1 puts the file beside the manifest that the build pushed.
What the frontend publishes
docker buildx build ./my-kit -f ./my-kit/my-kit.yaml -t docker.io/<NAMESPACE>/my-kit:1.0.0 --push builds and publishes a kit docs-sbx Publish an image. The capture kit ran it against a local registry with a docker-container builder (28-builder.txt, 28-buildx.txt), and the registry returned this manifest:
"mediaType": "application/vnd.oci.image.manifest.v1+json",
…
"annotations": {
"org.opencontainers.image.description": "A shell workload with one network grant, in the v3 descriptor form.",
"org.opencontainers.image.licenses": "Apache-2.0",
"org.opencontainers.image.title": "hello-kit",
"org.opencontainers.image.version": "1.0.0",
"vnd.docker.sandbox.kit.built-by": "{\"name\":\"docker/sandbox-kit\",\"version\":\"3.0.0-m.8\",\"revision\":\"129be2ff45e8f9463450eb3cf04ddcb52c2b76e5\"}",
"vnd.docker.sandbox.kit.capabilities": "com.docker.sandbox/network-policy@2",
"vnd.docker.sandbox.kit.descriptor": {
"schemaVersion": "3",
…
"provides": [
"deb/adduser@3.153",
…
"... 397 more derived provides entries"
],
…
"vnd.docker.sandbox.kit.schema-version": "3"The descriptor annotation is "The published descriptor as compact JSON" kitspec §9.3, with the build: text kept. For a workload, the frontend also reads the package database and adds one deb/ entry per installed package kitspec §9.6. schema-version and capabilities repeat two fields, and the org.opencontainers.image.* keys copy the display fields kitspec §9.3. The build staged the sources at /usr/share/sandbox/kit/kit, after the stem of kit.yaml kitspec §10.
Which commands accept a v3 kit
Conflict C13 asks which generation each command accepts, and the captures settle it:
$ sbx kit validate ./fixtures/kits/hello-kit
error: kit ./fixtures/kits/hello-kit is a v3 source kit and this load path has no kit builder configured; artifact validation failed
[exit 1]$ sbx kit inspect ./fixtures/kits/hello-kit --json
→ build kit ./fixtures/kits/hello-kit (sbx-kit-src:kit-<id>)
error: build kit ./fixtures/kits/hello-kit: exit status 1 ERROR: failed to build: OCI exporter is not supported for the docker driver. Switch to a different driver, or turn on the containerd image store, and try again. …
[exit 1]The ruling: sbx kit validate, pack, push, and pull are v1 and v2 tooling. pack refused the directory too, for lack of a spec.yaml (13-kit-v3-pack.txt). The documentation agrees: "Use Buildx for v3 kits. The sbx kit pack, push, and pull commands are for v1 and v2 kits" docs-sbx Publish an image.
A v3 kit is consumed by sbx run and --kit, since "sbx run and sbx create now accept sandbox kit references as the agent positional" rel-sbx v0.42.0. The pushed hello-kit still never reached a sandbox, because kit.allowedSources defaults to ["docker.io/"] (C40, 02-settings.txt) rel-sbx v0.34.0:
$ sbx create localhost:15000/m101/hello-kit:v1 fixtures/repo-kit --name m101-v3
error: resolve kits: kit "localhost:15000/m101/hello-kit:v1": kit "localhost:15000/m101/hello-kit:v1" cannot be installed — its source is not in your allowlist; current kit.allowedSources: docker.io/; …
try: sbx settings set kit.allowedSources '["docker.io/","localhost:15000/m101/"]'
…
[exit 1]The capture kit never changes sbx settings, so no capture shows sbx running a v3 kit. Two rules bound the mix. "V3 kits cannot be combined with v1 or v2 kits in the same sandbox" docs-sbx Version compatibility. And "The built-in agent names, such as claude and codex, select v2 kits" docs-sbx Version compatibility. Docker publishes its v3 workloads as docker/sbx-kit-*, such as docker.io/docker/sbx-kit-codex:0.155.1 docs-sbx Run a kit.
Where the run and the pages disagree
Five results of the run contradict the specification or the documentation, and rows C113 to C117 of the register cover them. The table prints both sides:
| The page says | The capture shows | Files |
|---|---|---|
docker buildx build -f kit.yaml --push publishes the kit docs-sbx Publish an image | the default docker driver of Desktop 4.94.0 pushed a Docker schema 2 manifest with no annotations | 28-buildx-docker-driver.txt, 28-manifest-docker-driver.json |
| the frontend "promotes all four onto the image index whenever the export produces one" kitspec §9.3 | the index has an attestation manifest and no annotations, and the arm64 manifest carries all eight | 28-index.json, 28-manifest.json |
the floating docker/sandbox-kit:3 never moves for a milestone (the README of the specification) | :3 resolved to the milestone 3.0.0-m.8 of 2026-10-02 | 28-buildx.txt |
"During development, you can pass a local source directory to sbx instead" docs-sbx Publish an image | sbx kit inspect of the directory failed: the OCI exporter is not supported for the docker driver | 13-kit-v3-inspect.txt, 28-kit-inspect-source.txt |
| source-form builds "run inside a shared builder sandbox named sbx-kit-builder" help-sbx sbx kit builder | that build ran in the host Docker daemon, and the builder status recorded after it reads "not created" | 13-kit-v3-inspect.txt, 13-kit-builder-status.txt |
An image without the descriptor annotation is not a kit to any consumer kitspec §10, so build with a docker-container builder. A consumer that finds no annotation on the index reads the platform manifest kitspec §9.3. The specification calls itself experimental, with a final version targeted for Q4 2026.
The default size of a kit volume stays in dispute (C12). The v0.39.0 notes say 512 MB rel-sbx v0.39.0, the research notes say 20 GiB, and this manual prints both. kit-tck validate judges a published artifact, kit-tck inspect reads it back, and this manual ran neither. Section 7.4 lists the claims it checked instead.
Sources:kitspec §1, §1.1, §1.2, §3.2, §4, §7, §9.3, §9.6, §10 (research/sources/SPEC-v3-at-v3.0.0-m.8.md); kitcap network-policy@2 (research/sources/kit-capabilities.md); research/sources/kit-spec-extras.md (README and RELEASES.md of docker/sandbox-kit-spec, for the frontend tag and the milestone dates); docs-sbx Kits, Use kits, Build and distribute kits (research/sources/docs-sandboxes.md); help-sbx sbx kit builder (research/sources/help-sbx.md); rel-sbx v0.34.0, v0.39.0, v0.42.0 (research/sources/sbx-releases.md); research/conflicts-register.md rows C12, C13, C40, and C113 to C117; capture/README.md (the findings of K28); capture/fixtures/kits/hello-kit/kit.yaml; capture/out/02-settings.txt, 13-kit-v3-validate.txt, 13-kit-v3-pack.txt, 13-kit-v3-inspect.txt, 13-kit-builder-status.txt, 28-builder.txt, 28-buildx.txt, 28-buildx-docker-driver.txt, 28-manifest.json, 28-manifest-docker-driver.json, 28-index.json, 28-kit-inspect-source.txt, 28-kit-inspect.txt, 28-run-kit.txt
`com.docker.sandbox/*` capabilities
A kit grants itself nothing: each capability is typed and versioned, the resolver unions one workload with its mixins, and any widening on update stops for approval.
The gh mixin a colleague published reaches api.github.com with your token. You want to know exactly which requests it can make with that token, and what changes when version 2 arrives. Both answers are in its capabilities list.
When you finish this section, you can read that list, predict the merged grant set, and say which update will stop and ask.
Typed, versioned requests
Capability: one typed request in a kit's capabilities list, "Everything the Kit needs but cannot supply itself" kitspec §7. The host answers each one: granted, refused, or prompted.
Each entry has a type of the form <namespace>/<name>@<version>, an optional display name, and an optional flag kitspec §7. Its config is decoded strictly for the types the specification defines, so an unknown key is an error kitspec §7. The version names the config schema, so network-policy@1 and @2 both exist and a descriptor states one of them kitspec §7.1. "Policy-shaped types are singletons" kitspec §7.1, while instance-shaped types repeat once per thing requested, such as credential@1 per service and phase.
Unknown types are allowed by design: "An unknown type is the extension point working as designed" kitspec §7.3. A required unknown type fails resolution, and an optional one is skipped and recorded.
The nineteen types at the pin
| Type | Shape | What it asks for |
|---|---|---|
network-policy@1 | singleton | hosts the sandbox can reach, per phase |
network-policy@2 | singleton, exclusive with @1 | hosts plus HTTP methods and paths |
credential@1 | per service and phase | one service the workload authenticates to |
ssh-agent@1 | per phase | what the forwarded SSH agent signs |
volume@1 | per path | persistent or tmpfs storage |
host-mount@1 | per path | a host directory, sharing the storage key with volume@1 |
port@1 | per container port and transport | a published port |
usb-device@1 | instance | a USB device match |
resources@1 | singleton | CPU, memory, and GPU limits, a constraint rather than a grant |
privileged@1 | singleton, no config | a privileged container |
long-running@1 | singleton, no config | keep running with no session attached |
lifecycle@1 | singleton | install hooks, startup hooks, files, the interactive argv |
agent-context@1 | singleton | the instruction file for the agent |
agent-sessions@1 | singleton | headless prompt and resume verbs |
agent-skills@1 | per path | where the agent reads skills |
agent-skill@1 | per effective name | one bundled skill |
git-identity@1 | singleton, no config | the runtime's git name and email |
kit-registry@1 | singleton, no config | reach the runtime's own kit registry, where builds push and pull |
sbx@1 | singleton, no config | launch the workload a particular way, with the identity the image states |
The table is kitspec §7.2 at the v3.0.0-m.8 tag. The main branch adds a twentieth, com.docker.sandbox/agent-interactive-sessions@1, as a singleton kitspec-main §7.2, unreleased at the pin. The newest type in sbx itself arrived with the pinned release: "Kits can declare com.docker.sandbox/long-running@1 to keep local sandboxes running after all sessions disconnect" rel-sbx v0.47.0.
network-policy@2 and credential@1
A network-policy@2 entry is a plain host string or an object with hosts, methods, and paths kitcap network-policy@2. A host string grants the connection for any protocol, as @1 did. An entry with methods or paths is bounded: it grants those HTTP requests and nothing else on that host. Deny entries are "Entries to refuse. Deny wins" kitcap network-policy@2. A refused request gets a status, because a runtime "MUST answer a request these entries refuse with HTTP status 403" kitcap network-policy@2. The proxy that enforces this is the one in the policy log section.
A credential@1 entry names a service and a phase, "never where the secret lives" kitcap credential@1. With apiKey.proxyManaged, "The real value stays on the host" kitcap credential@1. Every inject domain must appear in the allow list of the same phase kitcap credential@1. The binding that answers it is the one the secret section stores.
Resolution and the merged set
A runtime "MUST include exactly one workload Kit per composition" kitspec §5.3. It "MUST fail when two Kits provide the same normalized name" kitspec §5.3, so claude plus claude-mixin is refused. Across the set, allow entries union per phase, deny entries union per phase, and deny wins on the result kitcap network-policy@2. When one kit allows a host outright, another kit's bounded entry for that host is dropped, because the union is the unbounded grant kitspec §9.5. Figure 4.2 runs that arithmetic on three kits.
What a grant looks like inside sbx
The recording host ran a v2 mixin, so the captured rows come from permissions.network.allow rather than a v3 entry. The daemon stores the grant as a policy rule the sandbox owner cannot edit:
"name": "kit:m101-kit",
…
"scope": "sandbox:m101-kit",
…
"resource_type": "network",
"decision": "allow",
"resources": [
"example.com"
],
…
"editable": false,
…
"created_via": "provisioned",sbx policy check network example.com --sandbox m101-kit --json then answers "allowed": true (capture/out/12-check-kit.json). The rule carries created_via: provisioned and answers to --source kit, which is how the policy section tells a kit rule from one you added.
Updates and the lock
Every grant projects onto one normalized set, and "Consumers that gate updates store a Kit's surface in the lock and diff a candidate's against it" kitspec §7.4. A new version whose set stays inside the stored one can apply silently. Any widening must stop for approval. The list is explicit: a new allow entry, "a removed deny entry (the deny was part of what made the grant acceptable)" kitspec §7.4, a new credential, path, or port, or write access over a read-only skills path. optional does not change the set kitspec §7.4. Six types contribute nothing to it: resources@1, lifecycle@1, agent-context@1, agent-sessions@1, sbx@1, and long-running@1 kitspec §7.4.
Sources:kitspec §5.3, §7, §7.1, §7.2, §7.3, §7.4, §9.5 (research/sources/SPEC-v3-at-v3.0.0-m.8.md); kitspec-main §7.2 (research/sources/SPEC-v3.md, unreleased); kitcap network-policy@2, credential@1 (research/sources/kit-capabilities.md); rel-sbx v0.47.0 (research/sources/sbx-releases.md); capture/fixtures/kits/hello-kit/kit.yaml; capture/out/12-policy-kit.json, 12-policy-kit.txt, 12-check-kit.json
`sbx kit pack`, `sbx kit push --sign`, and `sbx kit verify`
The sbx kit commands package, sign, push, verify, and show provenance for v1 and v2 artifacts, and kit add appends a mixin to a running sandbox.
You wrote a v2 mixin that installs tree and drops a HELLO.md into the workspace. A colleague wants it without cloning your repository, and your security team wants to know who signed it. The sbx kit commands cover both, as long as the kit is v1 or v2.
When you finish this section, you can validate and pack a v2 kit, and describe what push --sign attaches to the manifest. You can also verify a signature and add a mixin to a running sandbox.
spec.yaml and files/
Mixin (v2): a directory with a spec.yaml whose schemaVersion is "2" and kind is mixin, plus an optional files/ tree docs-sbx Use existing kits. The capture kit wrote one:
schemaVersion: "2"
kind: mixin
name: hello-mixin
description: Installs tree and adds a greeting file to the workspace.
setup:
install:
- command: apt-get update && apt-get install -y tree
user: "0"
description: Install tree
permissions:
network:
allow:
- example.comfiles/workspace/HELLO.md holds one line, and sbx kit inspect reports it as a workspace file with mode 420 (capture/out/12-inspect.json). The v2 grammar is the one the built-in agents use, and its normative text lives in SPEC-v2.md of docker/sbx-kits-contrib docs-sbx Top-level fields.
validate and pack
sbx kit validate REFERENCE accepts a directory, a ZIP, or a git reference, and --json reports a verdict and warnings help-sbx sbx kit validate:
{
"reference": "./fixtures/kits/hello-mixin",
"kind": "directory",
"valid": true,
"warnings": []
}$ sbx kit pack ./fixtures/kits/hello-mixin -o $CAPTURE/work/hello-mixin.zip
Packed artifact to $CAPTURE/work/hello-mixin.zip
[exit 0]
$ zip listing of work/hello-mixin.zip (size, name)
0 files/
0 files/workspace/
22 files/workspace/HELLO.md
297 spec.yamlFor pack, "The directory must contain a valid spec.yaml and an optional files/ directory" help-sbx sbx kit pack. The same ZIP validates with kind: "zip" (capture/out/12-validate-zip.json). A v3 directory fails both commands, as the descriptor section showed. Figure 4.3 draws the whole path, with the steps the capture kit did not run in dashed boxes.
sign, push --sign, and pull
The capture kit ran none of these three, because it has no registry and no signing identity. The help text is the source. sbx kit sign is keyless by default, with an OIDC token from the CI provider or a browser login. "A token is never read from SIGSTORE_ID_TOKEN, so it cannot be chosen by anything that can set an environment variable" help-sbx sbx kit sign. With --key it signs with an unencrypted PEM private key. "For a local directory, a detached signature bundle is written to kit.sig.bundle next to spec.yaml" help-sbx sbx kit sign.
sbx kit push DIRECTORY REFERENCE chooses the artifact form from schemaVersion: a ZIP for "1", a tar+gzip layer with the spec in the config blob for "2" help-sbx sbx kit push. "With --sign, the pushed manifest is signed and the Sigstore bundle is attached to the kit as an OCI referrer" help-sbx sbx kit push. Signed or not, "Every push also attaches a SLSA provenance attestation as an OCI referrer" help-sbx sbx kit push, and "The provenance is unsigned unless --sign is given" help-sbx sbx kit push. --tlog-upload=false keeps a private kit out of the Rekor log. For pull, "The registry must support HTTPS" help-sbx sbx kit pull.
A published v3 image can be signed by reference with the same command, but "V3 kits shared as source directories or Git references can't be signed" docs-sbx Sign and verify kits.
verify, provenance, and the settings that admit a kit
sbx kit verify takes --key PUB for a key, or --certificate-identity with --certificate-oidc-issuer for a keyless signature. --insecure-ignore-tlog accepts a private keyless signature, and "It has no effect on key-based verification" help-sbx sbx kit verify. sbx kit provenance REFERENCE prints the attestation, and "only attestations that verify and whose subject matches the kit's own digest are reported as VERIFIED" help-sbx sbx kit provenance.
Six settings decide what a sandbox accepts, all at their defaults on the recording host:
kit.allowExtractedAgents true bool default Admit the pinned kit references that replace fo…
kit.allowLocalKits true bool default Allow installing kits from local directories or…
kit.allowedSources ["docker.io/"] json default JSON array of allowed kit source prefixes (e.g.…
kit.ignoreTransparencyLog false bool default Verify keyless kit signatures without requiring…
kit.requireSignature false bool default Require a valid signature from a trusted signer…
kit.trustedSigners [{"identityRegexp":"^.*@docker\… json default JSON array of trusted signer policies (key-base…"When kit.requireSignature is true, sbx rejects unsigned kits, signatures that don't match kit.trustedSigners, and ZIP kits" docs-sbx Require signed kits. For kit.trustedSigners, "The default policy trusts Docker employee identities ending in @docker.com, attested by Google's issuer" docs-sbx kit.trustedSigners. The settings table lists the environment variable of each key.
kit add and the builder
sbx kit add SANDBOX REFERENCE recreates the container with the mixin appended, "preserving kit-owned volumes (e.g. agent session state) across the swap" help-sbx sbx kit add. The capture shows the accepted case and the refused one:
$ sbx kit add m101-demo ./fixtures/kits/env-mixin
Recreating sandbox "m101-demo" to apply augmented kit list...
Swap container m101-demo-swap-<id> started (id=<sha256>).
Kit "env-mixin" added to sandbox "m101-demo"
[exit 0]
$ sbx exec m101-demo sh -c 'env | grep M101_FROM_KIT'
M101_FROM_KIT=yes
[exit 0]$ sbx kit add m101-demo ./fixtures/kits/hello-mixin
error: kit "hello-mixin" declares files, which the kit-add recreate flow does not yet apply; recreate the sandbox from scratch via `sbx rm` + `sbx create --kit` to use this kit
[exit 1]Sandboxes created before the recreate-aware label are refused with an error help-sbx sbx kit add, and long-running@1 cannot be added this way rel-sbx v0.47.0. sbx create shell --kit ./fixtures/kits/hello-mixin applied the same mixin in full: one install command, one workspace file, and one allow rule (capture/out/12-run-kit.txt, 12-policy-kit.txt).
Source-form builds of v3 kits "run inside a shared builder sandbox named sbx-kit-builder, created on first use" help-sbx sbx kit builder. sbx kit builder status reports it, and history passes through to docker buildx history:
$ sbx kit builder status
Builder: not created (the first source-form kit build creates it)
Builder kit: docker.io/docker/sbx-kit-builder:1
Kit registry: 127.0.0.1:5411 — reachable
[exit 0]Sources:help-sbx sbx kit add, builder, pack, provenance, pull, push, sign, validate, verify (research/sources/help-sbx.md); docs-sbx Kits v2, Build and distribute kits, kit settings (research/sources/docs-sandboxes.md); rel-sbx v0.47.0 (research/sources/sbx-releases.md); research/conflicts-register.md row C13; capture/fixtures/kits/hello-mixin/spec.yaml, capture/fixtures/kits/env-mixin/spec.yaml; capture/out/02-settings.txt, 12-inspect.json, 12-validate.json, 12-validate-zip.json, 12-pack.txt, 12-run-kit.txt, 12-policy-kit.txt, 12-kit-add.txt, 12-kit-add-files.txt, 13-kit-builder-status.txt
sbxenv.yaml and `sbx env plan`
An environment file declares the agent, kits, workspace, secrets, ports, MCP servers, and host commands, and sbx env plan prints every change before create or run applies it.
A new contributor clones your repository and needs the sandbox you use. That means the shell workload, the hello-mixin, one published port, a greeting variable, and a host command that prepares the workspace. You could send a list of flags. An environment file sends the same thing as one checked-in file, and the contributor sees a plan before anything runs.
When you finish this section, you can write an sbxenv.yaml, read its plan symbol by symbol, and say which edits wait for the next create.
The file
Environment file: a sbxenv.yaml that declares one sandbox and the host resources around it, read by the five sbx env commands help-sbx sbx env. The capture kit wrote this one:
schemaVersion: "1"
name: m101-env
agent: shell
workspace: .
args:
greeting:
default: hello
description: Greeting word
kits:
- source: ../kits/hello-mixin
env:
GREETING: ${{ env.args.greeting }}
lifecycle:
initialize:
- command: echo init
ports:
- sandbox: 8080
host: 18083| Key | What it declares |
|---|---|
schemaVersion, name, agent | the file version "1", the sandbox name, the built-in agent or kit |
args | inputs with default or required: true, read as ${{ env.args.NAME }} |
kits | mixins, as a reference or as source plus args |
workspace, additionalWorkspaces | the host directories to mount, with an object form for clone mode |
env | variables for the sandbox |
secrets, bindings, registries | credentials with value, ref, or command, and the domains they reach |
mcp.servers | MCP servers registered with the gateway |
ports | published ports, sandbox, host, protocol, hostIP |
lifecycle | host commands under initialize, postCreate, preRemove |
sandboxOptions | template, memory, cpus, skills, writableEnvFiles, and more |
The keys come from the file reference docs-sbx Top-level fields and from the command's own help. Three placeholders expand, ${{ env.args.NAME }}, ${{ env.projectDir }}, and ${{ env.fileDir }}, and a $ anywhere else is literal text help-sbx sbx env create. A relative kit path or workspace resolves against the directory of the file that declares it, so workspace: . mounts the file's own directory help-sbx sbx env. A file with no workspace mounts nothing rel-sbx v0.42.0.
The plan
sbx env plan reads the file and prints what applying it would set up. "Nothing is applied, approved, or recorded" help-sbx sbx env plan.
── ENVIRONMENT PLAN
m101-env
kits:
+ - source: $CAPTURE/fixtures/kits/hello-mixin
workspace: (present, not recorded as applied, needs your approval)
path: $CAPTURE/fixtures/env
env:
+ GREETING: hello
+ sandbox:
+ name: m101-env
+ agent: shell
ports:
+ - sandbox: 8080
+ protocol: tcp4
+ host: 18083
lifecycle:
initialize:
+ - command: echo init
+ workdir: $CAPTURE/fixtures/env
+ envFiles:
+ - $CAPTURE/fixtures/env/sbxenv.yaml
Plan: + 5 to add, ~ 0 to change, - 0 to destroy.
this plan runs commands on this machine, outside the sandbox, with your own privileges
✓ not approved yet; applying asks onceThe plan has the shape of your file, and the margin carries the verdict. Five symbols exist help-sbx sbx env. + adds, ~ changes, and - destroys. > marks a command that runs again, and ! marks a resource the environment applied and no longer declares.
A row that is unchanged and already approved is left out. "Literal secret values appear as SHA-256 digests" docs-sbx Review an environment plan. The tcp4 protocol was not in the file: published ports default to it since v0.42.0 rel-sbx v0.42.0. Figure 4.4 follows the file through the plan to the state record.
Host commands and approval
The lifecycle block runs on the host, outside the sandbox, with your own privileges. initialize runs on every create and every run, postCreate once the sandbox exists, and preRemove after you confirm sbx env rm help-sbx sbx env. Commands run through your shell from the project directory, with workdir and timeout per command. An environment with any of them "asks on every invocation, whether or not this one is what runs them" help-sbx sbx env, because approving a command also trusts the script it calls. --skip-host-commands runs none of them, --auto-approve answers yes once, and sbx settings set env.rememberHostCommands true asks again only when a command changes.
The state record and the second plan
"What was approved is recorded per environment under sbx's state directory, not next to the file" help-sbx sbx env, "so a later invocation asks only about what moved" help-sbx sbx env. The exact path was not captured, because the kit ran plan and never create. The second capture changed one input instead:
$ sbx env plan --env-arg greeting=servus ./fixtures/env
…
+ GREETING: servus
…
Plan: + 5 to add, ~ 0 to change, - 0 to destroy.Every row is still +, because nothing was applied. After a create, the same edit prints ~ GREETING: hello -> servus and nothing else, in the form the help text shows for a changed kit argument help-sbx sbx env. Changes to workspaces, kits, ports, secrets, bindings, and sandboxOptions wait for the next create. New env values reach the next session docs-sbx Update an environment.
Several PATH arguments deep-merge in order, "later files override earlier ones" help-sbx sbx env create, and lists such as ports concatenate rather than override. With no PATH, a .sbxenv.yaml in your home directory merges underneath as a base layer help-sbx sbx env create. A hidden .sbxenv.yaml in the project is no longer read.
rm, the read-only file, and --cloud
sbx env rm removes the sandbox and the secrets provisioned at its scope. Bindings stay, since they are user-wide, so "pass --prune-bindings to also remove the bindings this environment declares" help-sbx sbx env rm. The file itself is bound read-only inside the sandbox, because an agent that can edit it decides what the next plan asks about. sandboxOptions.writableEnvFiles: true lifts that help-sbx sbx env.
With --cloud, "workspace, additionalWorkspaces and clone name host directories, which a cloud sandbox cannot mount" help-sbx sbx env, and host port bindings, MCP definitions, and dynamic secret sources are rejected before any host command runs. The cloud section shows what remains.
Sources:help-sbx sbx env, sbx env create, sbx env plan, sbx env rm (research/sources/help-sbx.md); docs-sbx Sandbox environment files (research/sources/docs-sandboxes.md); rel-sbx v0.42.0 (research/sources/sbx-releases.md); capture/fixtures/env/sbxenv.yaml; capture/out/14-env-plan.txt, 14-env-plan-arg.txt
`sbx skills add` and `skills: true`
One shared skills store serves every sandbox of a supported agent read-only by default, and docker-agent reads its own skill directories, not the store, with skills: true.
You keep a pdf skill under ~/.claude/skills on your laptop. Inside a sandbox the agent has its own home directory and never sees it. The shared skills store is the one place you fill, and every sandbox for a supported agent reads it.
When you finish this section, you can fill the store, choose how each sandbox mounts it, and turn the same skills on for docker-agent.
The store
Skill: a directory with a SKILL.md whose metadata an agent reads into its system prompt, and whose body it loads when a task matches docs-agent How Skills Work.
Shared skills store: the host directory that sbx links into every sandbox for a supported agent. On the recording host it was empty:
{
"store": "$HOME/Library/Application Support/com.docker.sandboxes/sandboxes/agent-skills",
"skills": []
}On Linux the store is ~/.local/state/sandboxes/sandboxes/agent-skills, and on Windows %LOCALAPPDATA%\DockerSandboxes\sandboxes\state\agent-skills docs-sbx Import skills from the host. "Running sbx reset clears the shared store" docs-sbx Shared store behavior.
add, import, ls, rm, and update
sbx skills add <repository> installs from a Git URL or a GitHub owner/repository, and "The repository must contain one or more valid SKILL.md files" help-sbx sbx skills add. --skill NAME picks skills by name, repeatable or comma-separated, so sbx skills add anthropics/skills --skill pdf installs one. The capture kit did not record an add, because the step needs GitHub. sbx skills update refreshes only skills that add installed, and sbx skills rm asks before it removes one help-sbx sbx skills update.
sbx skills import copies skills already installed on the host, checking six directories in order, and the first copy of a duplicate name wins help-sbx sbx skills import:
| Host source | Agent | Mount target in the sandbox |
|---|---|---|
~/.agents/skills | Codex and Devin | /home/agent/.agents/skills |
~/.claude/skills | Claude Code | /home/agent/.claude/skills |
~/.config/opencode/skills | OpenCode | not listed in the docs table |
~/.copilot/skills | Copilot | /home/agent/.copilot/skills |
~/.cursor/skills | Cursor | /home/agent/.cursor/skills |
~/.factory/skills | Droid | /home/agent/.factory/skills |
The order comes from the help text and the mount targets from the docs docs-sbx Import skills from the host. "Imported skills are available to Claude, Codex, Copilot, Cursor, Droid, and OpenCode" help-sbx sbx skills import. sbx skills import arrived in v0.37.0, and add, update, and rm in v0.42.0 rel-sbx v0.42.0.
How a sandbox sees the store
By default "the store's entries are linked into the agent's skills directory read-only, which stays writable so kits can install skills beside them" help-sbx sbx skills. "Linking happens at container start" help-sbx sbx skills, so an edit to an existing skill is live, while "adding a store entry reaches a running sandbox only on its next start" help-sbx sbx skills. Removing a skill breaks its link at once.
--skills off|readonly|readwrite on sbx run or sbx create chooses the mode per sandbox, and readwrite mounts the store over the directory so the sandbox's writes are shared. The default is readonly, or the skills.defaultMode setting, which the recording host had at its default (capture/out/02-settings.txt). The three-way flag replaced --no-share-skills in v0.43.0 rel-sbx v0.43.0. The capture kit did not list that directory from inside a sandbox, so the link form is not shown.
In a v3 kit the agent declares the path itself, with agent-skills@1, because "A runtime cannot know where an arbitrary agent reads skills" kitcap agent-skills@1. A runtime "MUST default an omitted mode to readonly" kitcap agent-skills@1. The effective access is the narrower of the host setting and the kit's mode. Raising a path to readwrite is a widening that stops for approval, as the capabilities section explains.
docker-agent and skills: true
docker-agent does not read the store. It scans its own directories, and two of them, ~/.claude/skills/ and ~/.agents/skills/, are also import sources. "Docker Agent scans standard directories for SKILL.md files" docs-agent How Skills Work, and "Skill metadata (name, description) is injected into the agent's system prompt" docs-agent How Skills Work:
| Path | Search |
|---|---|
~/.codex/skills/ | recursive |
~/.claude/skills/ | immediate children only |
~/.agents/skills/ | recursive |
.claude/skills/ | the current directory only |
.github/skills/ | each directory from the git root to the current one |
.agents/skills/ | each directory from the git root to the current one |
In the agent file, skills: true loads every discovered skill, a list restricts it, and false turns it off docs-agent Filtering Skills. A list item that is local or an http:// or https:// URL is a source, and any other string is a skill name. "A name that doesn't match any discovered skill is logged as a warning at startup but is otherwise ignored" docs-agent Filtering Skills. The agent needs the filesystem toolset to read skill files, as the toolsets section lists. context: fork in a skill's front matter "tells the agent to run the skill in an isolated sub-agent instead" docs-agent Running a Skill as a Sub-Agent.
In a docker-agent sandbox
With docker-agent run --sandbox, the host paths above are invisible from the VM. Docker Agent builds a kit before the sandbox starts, "bind-mounted read-only into the VM at the same path" docs-agent Auto-Kit. Every SKILL.md found on the host "is copied under <kit>/skills/<skill-name>/" docs-agent What gets staged, every text file passes a secret redaction step, and --no-kit turns the staging off. The end-to-end section shows the printed summary of what was staged.
Sources:help-sbx sbx skills, sbx skills add, sbx skills import, sbx skills update (research/sources/help-sbx.md); docs-sbx Share agent skills, skills.defaultMode (research/sources/docs-sandboxes.md); docs-agent Skills, Sandbox Mode Auto-Kit (research/sources/docs-docker-agent.md); kitcap agent-skills@1 (research/sources/kit-capabilities.md); rel-sbx v0.37.0, v0.42.0, v0.43.0 (research/sources/sbx-releases.md); capture/out/02-settings.txt, 15-skills-ls.json, 15-skills-ls.txt
docker-agent run and the Agent File
An agent file is agents, models, toolsets, and the rules between them, and docker-agent run drives the loop, records it, and replays it.
docker-agent run, new, and doctor
docker-agent run loads an agent file and drives the loop with or without a terminal, doctor says which model auto would pick, and new needs a terminal.
A colleague sends you an agent file and one command to run it. You have no API key, only Docker Desktop with Model Runner, and you want proof that the file runs before you open a chat window.
When you finish this section, you can run an agent file with no terminal, read its event stream, and ask doctor which model auto picks.
What run takes
Agent reference: the first argument of docker-agent run. It is a .yaml, .yml, or .hcl file, a registry reference, an alias, or nothing help-agent docker-agent run. With nothing, run uses docker-agent.yaml, docker-agent.yml, or docker-agent.hcl from the current directory, or else a built-in default agent docs-agent CLI Reference. coder is a second built-in agent, and alias add saves a name for a file or a reference together with run options such as --safety help-agent docker-agent alias add.
A registry reference behaves like a file. 24-run-ref.txt runs localhost:15000/m101/agent:v1 and prints the same answer as the local file, as share push and share pull shows. Each further argument is one user message, and the messages run as turns in order. A - reads the message from stdin help-agent docker-agent run.
--exec, --json, and --last
Headless run: run --exec, which writes to stdout and opens no TUI. --json writes one JSON event per line, and --last prints only the final answer help-agent docker-agent run. The capture runs capture/fixtures/agents/files.yaml with the same two messages in each form.
$ docker-agent run --exec --working-dir fixtures/repo --last --fake work/cassettes/18-files fixtures/agents/files.yaml 'List the files in the working directory.' 'How many lines does README.md have? Count them with a shell command.'
README.md has 1 line.
[exit 0]{"message": "List the files in the working directory.", "session_id": "<uuid>", "session_position": 0, "timestamp": "<ts>", "type": "user_message"}
{"agent_name": "root", "timestamp": "<ts>", "tool_call": {"function": {"arguments": "{\"path\": \".\"}", "name": "list_directory"}, "id": "<call-id>", "type": "function"}, "tool_definition": …
{"agent_name": "root", "content": "The working directory contains 1 file: README.md (1 line).", "message_id": "<uuid>", "session_id": "<uuid>", "timestamp": "<ts>", "type": "agent_choice"}The other event types of the file include team_info, toolset_info, tool_call_response, token_usage, and stream_stopped. The second turn shows the limit of a headless run. The model asked for shell with wc -l README.md, and the runtime raised tool_call_confirmation. With no terminal to answer it, the tool result was "The user rejected the tool call." The model then used read_file (18-transcript.txt). Permissions, --safety, and hooks explains why that call asked.
The model in every recording
The docs call ai/qwen3 "the model Docker Agent reaches for by default" docs-agent Set Up a Model. With no Model Runner, doctor resolves auto to the 8B tag ai/qwen3:latest (16-doctor.txt). On the recording Mac, two pulls of that tag ended with a digest mismatch after the last 5.03 GB blob (25-model-pull-latest.txt). Every agent file of the kit therefore names ai/qwen3:4b, the 4B tag of the same repository. It is a thinking model, so it streams a few thousand reasoning tokens before each answer. --exec prints them, and the kit cuts them to one line such as [... 299 lines of model reasoning cut by run.py ...] (18-run-exec.txt).
doctor and models
doctor reports provider credentials, whether Docker Model Runner answers, the auto pick, and, with a file, the variables that file needs. It exits non-zero on an issue help-agent docker-agent doctor. The kit ran it once with no Docker daemon and once with Model Runner up.
$ docker-agent doctor
…
Docker Model Runner
Status: unreachable: docker --config=$HOME/.docker --context=m101-no-daemon model status --json: …
…
Model auto-selection
auto -> dmr/ai/qwen3:latest
Issues
- no usable model: no provider credential was found and Docker Model Runner is unreachable; …
Error: 1 issue(s) found
[exit 1]$ docker-agent doctor
…
Docker Model Runner
Status: reachable, 1 model(s) pulled:
- docker.io/ai/qwen3:4b
Model auto-selection
auto -> dmr/docker.io/ai/qwen3:4b
No issues found.
[exit 0]With no runner, auto still names dmr/ai/qwen3:latest, and the issue line says why nothing can run. With the runner up, auto takes the one pulled model, as the provider page says: auto-selection "prefers a locally-installed model" docs-agent Docker Model Runner. models list printed one row, dmr ai/qwen3:latest, even with no daemon (16-models.txt). setup is the interactive fix. Its four paths are a provider key in ~/.config/cagent/.env, a Model Runner pull, a custom OpenAI-compatible endpoint, and the Claude Code harness help-agent docker-agent setup.
new
new asks questions and writes an agent file, and a description argument skips "the initial prompt" help-agent docker-agent new. Its --model takes anthropic, openai, google, dmr, or a custom provider, and --max-iterations defaults to 20 for DMR (conflict C66). In the capture, new with a description and no controlling terminal stopped at /dev/tty and wrote no file (17-new-agent.yaml).
$ docker-agent new --model dmr/ai/qwen3:4b 'an agent that greets the user and names one fact about Docker sandboxes' (in work/, without a controlling terminal)
…
Error: bubbletea: error opening TTY: bubbletea: could not open TTY: open /dev/tty: device not configured
[exit 1]Figure 5.1 puts the pieces of one run on one page.
run loads the agent file, alternates model calls and tool calls in one loop, and writes the answer, the events, a session, and with --record a cassette. Read from the top left. Solid plum is a model call and dashed olive a tool effect. Dotted indigo is the event stream, and solid teal is a stored record. With --fake, the cassette answers in place of Model Runner. From capture/fixtures/agents/files.yaml, capture/out/18-run-json.ndjson, and capture/out/18-cassette-head.txt.Sources:help-agent docker-agent run, setup, doctor, new, alias add (research/sources/help-docker-agent.md); docs-agent CLI Reference, Set Up a Model, Docker Model Runner (research/sources/docs-docker-agent.md, pages features/cli, getting-started/set-up-a-model, providers/dmr); conflict C66 (research/conflicts-register.md); capture/README.md; capture/fixtures/agents/files.yaml; capture/out/16-doctor.txt, 16-models.txt, 17-new.txt, 17-new-agent.yaml, 18-run-exec.txt, 18-run-last.txt, 18-run-json.ndjson, 18-transcript.txt, 18-cassette-head.txt, 24-run-ref.txt, 25-doctor.txt, 25-model-pull-latest.txt
agents, models, and providers
agents is the only required block, a model reference is provider/model or a name from models, and providers sets a reusable endpoint such as Docker Model Runner.
You open capture/fixtures/agents/files.yaml and find model: local on the agent, a local entry under models, and provider: dmr inside that entry. A second file, dmr.yaml, adds a providers block with a URL, and its run never reaches the recorder.
When you finish this section, you can follow an agent's model value to the endpoint it calls, and predict what auto picks.
The file and its schema
Agent file: a YAML or HCL document that matches agent-schema.json, the schema of Docker Agent v16. Its root allows 16 keys and requires one: "Map of agent configurations. At least one agent is required" schema agents. A misspelled top-level key fails the load, because "the parser rejects unknown top-level keys" docs-agent Configuration Overview.
version is a string from "0" to "16" in the schema enum. The same docs page still says "The current version is 15" docs-agent Configuration Overview. The capture files declare "16" and load, so this manual follows the schema. "When you load an older config, Docker Agent automatically migrates it to the latest schema" docs-agent Configuration Overview. The agent that share pull fetched from agentcatalog/pirate declares version: "2" and names openai/gpt-4.1 inline (16-share-pull-agent.yaml).
version: "16"
agents:
root:
model: local
description: Lists the files of a small repository and counts lines.
instruction: |
You answer questions about the files in the working directory.
Use the tools to look; never guess. Answer in one sentence that
names the files and gives the line count.
max_iterations: 8
toolsets:
- type: filesystem
- type: shell
models:
local:
provider: dmr
model: ai/qwen3:4b
temperature: 0Agent keys
Agent: one entry under agents, named by its key. The first agent of the team runs unless --agent names another help-agent docker-agent run. The schema AgentConfig lists the keys, and these are the ones the capture files use or rely on.
| Key | In the capture | Schema rule |
|---|---|---|
model | local, qwen, dmr/ai/qwen3, openai/gpt-4.1 | a model name or provider/model schema AgentConfig |
instruction | a block string in every file | the system prompt: a string, or a list joined with blank lines |
max_iterations | 8 in files.yaml, 3 for writer | an integer from 0 |
max_consecutive_tool_calls | not set | identical calls before the agent stops, and 0 means the default of 5 |
redact_secrets | not set | true by default: a builtin scrubs secrets on three hook events |
toolsets, sub_agents, handoffs, hooks | files.yaml, team.yaml, guarded.yaml | toolsets, delegation, permissions |
Model references
Model reference: the value of model on an agent. It takes five forms:
provider/modelinline, such asdmr/ai/qwen3ingreeter.yamldocs-agent Models.- a name under
models, such aslocal, withprovider,model, and parameters such astemperature: 0. - a
first_availablelist: "At load time, Docker Agent selects the first candidate whose credentials are configured" docs-agent Models. - an alloy: two references with a comma between them, which the runtime alternates in one conversation.
auto: "the first cloud provider with a configured credential", then a pulled Model Runner model docs-agent Set Up a Model.
run --model [agent=]provider/model replaces the reference for one run help-agent docker-agent run. On the recording Mac, doctor resolved auto to dmr/docker.io/ai/qwen3:4b with only the 4B tag pulled (25-doctor.txt). With no Model Runner, it named dmr/ai/qwen3:latest and reported no usable model (16-doctor.txt).
providers and the Model Runner endpoint
Provider definition: an entry under providers with an underlying provider (default openai), a base_url, a token_key, and defaults that its models inherit docs-agent Provider Definitions.
version: "16"
providers:
runner:
provider: dmr
base_url: http://localhost:12434/engines/llama.cpp/v1
models:
qwen:
provider: runner
model: ai/qwen3:4b
temperature: 0
…files.yaml names provider: dmr with no base_url, and then "Docker Agent auto-discovers the DMR endpoint" docs-agent Docker Model Runner. It runs docker model status --json, which 16-dry-run.txt prints inside its error when no Docker daemon answers. An eval container has no docker CLI, so the same discovery fails there, as eval shows.
With an explicit base_url, dmr.yaml answered Hello (25-run-dmr.txt). That run is live in every capture, because an explicit base_url bypasses the --record proxy and leaves the cassette empty (capture/README.md). Set base_url when discovery cannot work, and leave it out when you want a cassette.
Provider ids changed too. doctor prints fireworks-ai, togetherai, and moonshotai, and the docs Models table lists fireworks, together, and moonshot. This manual prints the ids that doctor prints (conflict C58), and the providers table has the rest. A named patch under flavors changes any of these blocks at run time with --flavor help-agent docker-agent run, and the capture ran none.
files.yaml resolves local to the dmr provider and finds the endpoint itself, while dmr.yaml resolves qwen through providers.runner to a fixed URL. Read each column top down, from the agent to the endpoint. Violet blocks are agent entries and plum blocks are model and provider entries. The bottom row is the auto decision as doctor printed it. From capture/fixtures/agents/files.yaml, capture/fixtures/agents/dmr.yaml, capture/out/16-doctor.txt, 16-dry-run.txt, and 25-doctor.txt.Sources:schema agents, AgentConfig (research/sources/agent-schema.json); docs-agent Configuration Overview, Models, Set Up a Model, Provider Definitions, Docker Model Runner (research/sources/docs-docker-agent.md, pages configuration/overview, concepts/models, getting-started/set-up-a-model, providers/custom, providers/dmr); help-agent docker-agent run (research/sources/help-docker-agent.md); conflict C58 (research/conflicts-register.md); capture/README.md; capture/fixtures/agents/files.yaml, dmr.yaml, greeter.yaml; capture/out/16-doctor.txt, 16-dry-run.txt, 16-share-pull-agent.yaml, 25-doctor.txt, 25-run-dmr.txt
toolsets and mcps
A toolset is one type from a fixed list of 27 that gives an agent tools, and mcp reaches a server by Docker reference, command, or URL.
Your agent answers questions about a repository from memory, and you suspect it never received a file tool. The agent file says filesystem, but the model sees tool names, not toolsets.
When you finish this section, you can list the tools an agent gives its model, call one without a model, and choose an MCP form.
The 27 types
Toolset: one entry under an agent's toolsets, with a type and options, that the runtime turns into one or more tools. docker-agent toolsets prints 27 types, and the schema Toolset enum has the same 27 schema Toolset.
$ docker-agent toolsets
TYPE SUMMARY
…
background_agents Dispatch work to sub-agents concurrently and collect results
…
filesystem Read, write, list, search, and navigate files and directories
…
mcp Extend agents with external tools via the Model Context Protocol
mcp_catalog Discover and activate remote MCP servers from the Docker MCP Catalog
…
plan Shared persistent scratchpad for multi-agent collaboration
…
shell Execute shell commands in the user's environment
…
[exit 0]transfer_task and handoff are not types. The docs built-in table lists both, the schema enum and the CLI list leave them out, and sub_agents and handoffs inject them (conflict C56). The same docs table lists session_plan, which the CLI does not print, and leaves out environment and file docs-agent Tool Configuration. The toolsets table follows the CLI.
From toolsets to tools
files.yaml declares two toolsets, filesystem and shell. Every toolset_info event of its run reports "available_tools": 10 (18-run-json.ndjson). The request in the cassette 18-files.yaml lists them: nine filesystem tools from directory_tree to remove_directory, then shell. The model chose list_directory and read_file from that list.
debug toolsets FILE --json prints the same names and schemas with no model help-agent docker-agent debug toolsets. For guarded.yaml it prints one tool:
[
{
"agent": "root",
"tools": [
{
"name": "shell",
"category": "shell",
…
"parameters": {
"additionalProperties": false,
"properties": {
"cmd": {
"description": "Shell command",
"type": "string"
},
"cwd": {
"description": "Working directory (default \".\")",
"type": "string"
},
"timeout": {
"description": "Timeout in seconds (default 30)",
"type": "integer"
}
},
…
"annotations": {
"idempotentHint": false,
"readOnlyHint": false,
"title": "Shell"
},
…debug tool FILE TOOL JSON calls one tool with no model turn, and "Calls have real side effects and bypass other hooks and approval checks" help-agent docker-agent debug tool. In 20-debug-tool.txt, shell with {"cmd":"echo m101-direct"} printed m101-direct. Use it to test a tool before any model calls it.
Keys every toolset shares
Tool filter: a key that narrows what a toolset gives the model. tools keeps only the listed names, and readonly keeps only tools whose annotations carry a read-only hint schema Toolset. defer hides tools until the model finds them with search_tool and add_tool. instruction replaces the toolset's built-in instructions unless the text contains {ORIGINAL_INSTRUCTIONS}. model names the model for the turn after a tool result. Since v1.148.0 a complete tool result is bounded to 50 KiB rel-agent v1.148.0.
The readOnlyHint: false on shell matters again in permissions, --safety, and hooks.
The mcp type
MCP toolset: a toolset with type: mcp that connects to one MCP server and gives its tools to the agent. The docs name three forms docs-agent Tool Configuration:
ref: docker:duckduckgoruns a catalog server in a container through the MCP Gateway.command,args, andenvstart a local process over stdio, and a missing binary is installed into~/.cagent/tools/bin/from the aqua registry.remote.urlwithtransport_typeset tostreamableorssereaches a server over the network, with optionalheaders.
lifecycle.profile sets reconnects for each mcp toolset: resilient by default, strict, or best-effort. A top-level mcps entry holds a server definition that agents reference as {type: mcp, ref: <name>} schema mcps. A top-level toolsets entry, named in use_toolsets, does the same for any type schema toolsets.
No mcp toolset ran in this edition's captures, so the three forms come from the docs and the schema. Whether ref: docker: needs Docker Desktop's MCP Toolkit stays open (conflict C73). The sandbox side of MCP is sbx mcp add, load, and --static-mcp. An agent served as an MCP server is serve mcp and serve acp.
toolsets entry becomes named tools in the model request: built-in types run in process, and mcp goes through a gateway, a child process, or a URL. Read each row left to right, from the agent file entry to the tool names in the model request. Solid rows ran in capture/out/18-run-json.ndjson, with names from the tools array of capture/cassettes/18-files.yaml.gz. Dashed rows come from the docs Tool Configuration page and were not run.Sources:schema Toolset, mcps (research/sources/agent-schema.json); docs-agent Tool Configuration (research/sources/docs-docker-agent.md, page configuration/tools); help-agent docker-agent toolsets, debug toolsets, debug tool (research/sources/help-docker-agent.md); rel-agent v1.148.0 (research/sources/docker-agent-CHANGELOG.md); conflicts C56, C73 (research/conflicts-register.md); capture/fixtures/agents/files.yaml, guarded.yaml; capture/cassettes/18-files.yaml.gz; capture/out/16-toolsets.txt, 18-run-json.ndjson, 20-debug-toolsets.json, 20-debug-tool.txt
sub_agents, transfer_task, and background_agents
transfer_task runs a sub-agent in a clean sub-session and returns its answer, handoff moves the whole session to another agent, and background agents need approval to start.
You want a coordinator that asks a writer for one sentence and then lets a reviewer answer the user. One agent file can say both things: a delegation that comes back, and a move that does not.
When you finish this section, you can choose between sub_agents, handoffs, and background_agents, and read each one in an event stream.
The team file
version: "16"
agents:
root:
model: local
description: Coordinates a writer and a reviewer.
instruction: |
You coordinate two colleagues and never write text yourself.
For every request: first call transfer_task to ask the writer for
the text, then call handoff to pass the conversation to the reviewer.
max_iterations: 6
sub_agents: [writer]
handoffs: [reviewer]
toolsets:
- type: background_agents
writer:
model: local
description: Writes one sentence on a given topic.
…
max_iterations: 3
reviewer:
model: local
description: Reviews the sentence it receives.
…All three agents use the model local, which is dmr/ai/qwen3:4b. Neither transfer_task nor handoff appears under toolsets, because the two lists inject them (conflict C56).
transfer_task: a child in a sub-session
Delegation: a call of transfer_task with agent, task, and expected_output, which sub_agents adds to the parent. "The call blocks until the sub-agent returns its result, which becomes the tool's response" docs-agent Transfer Task Tool.
user: Ask the writer for one sentence about microVMs.
root -> tool_call transfer_task {"agent": "writer", "task": "Generate one sentence about microVMs", "expected_output": "A single sentence describing microVMs"}
agent_switching: {"agent_name": "writer", "switching": true, "from_agent": "root", "to_agent": "writer"}
writer: MicroVMs are lightweight virtual machines that provide strong isolation and security for applications with minimal resource overhead.
stream_stopped (stop, normal)
sub_session_completed: {"agent_name": "root", "parent_session_id": "<uuid>", "sub_session": {"id": "<uuid>", "origin": "run", "title": "Transferred task", "messages": [{"message": {"agent_name": "", "message": {"role": "system", "content": "You are a member of a team of agents. Your goal is to complete the following task:\n\n<task>\nGenerate one sentence about microVMs\n</task>…
agent_switching: {"agent_name": "root", "switching": false, "from_agent": "writer", "to_agent": "root"}
root <- tool_call_response Transfer Task: MicroVMs are lightweight virtual machines that provide strong isolation and security for applications with minimal resource overhead.
root -> tool_call handoff {"agent": "reviewer"}
root <- tool_call_response Handoff Conversation: The agent root handed off the conversation to you. …
reviewer: APPROVED: MicroVMs are lightweight virtual machines that provide strong isolation and security for applications with minimal resource overhead.
stream_stopped (stop, normal)
user: Now hand the conversation to the reviewer.
reviewer: APPROVED: MicroVMs are lightweight virtual machines that provide strong isolation and security for applications with minimal resource overhead.
stream_stopped (stop, normal)The writer never saw the user's message. Its sub-session starts with a system message that holds <task> and <expected_output>. An implicit user message follows, "Please proceed." The sub-session ran under the writer's own max_iterations: 3, with "tools_approved": false. Then agent_switching returned control to root, and the sentence came back as the tool response.
"Unlike other tools, transfer_task is always auto-approved" docs-agent Multi-Agent Systems, and the capture agrees: the call raised no confirmation under --exec. A delegation to an agent already in the chain fails, and the depth is capped at 10 nested delegations docs-agent Transfer Task Tool. A sub_agents entry can also be a registry reference. Its tag is resolved again on every run unless you pin it to a digest schema AgentConfig.
The session move
Session move: a call of handoff with one argument, agent, which handoffs adds. The named agent "becomes the active agent and sees the full conversation history" docs-agent Multi-Agent Systems. In the capture, root called it in its next model turn, after the sentence came back. The tool response told reviewer which tools and agents it can use, and reviewer answered the user.
The second message shows the move. It went straight to reviewer, and the transcript has no root line after the move. force_handoff makes the same move on every final response without a tool call schema AgentConfig.
background_agents
Background agent: a sub-agent task that run_background_agent starts, which returns a task id at once. list_background_agents, view_background_agent, and stop_background_agent follow it, and the target must be in the caller's sub_agents docs-agent Background Agents Tool.
user: Run the writer as a background agent on the topic microVMs, wait for it, and repeat its sentence.
tool_call_confirmation: {"agent_name": "root", "tool_call": {"id": "<call-id>", "type": "function", "function": {"name": "run_background_agent", "arguments": "{\"agent\": \"writer\", \"task\": \"Write one sentence on the topic microVMs\", \"expected_output\": \"A single sentence about microVMs\"}"}}, "metadata": {"safety_label": "unknown"}}
root <- tool_call_response Run Background Agent: The user rejected the tool call.
…
stderr: Error: Agent terminated: detected 5 consecutive identical calls to run_background_agent. This indicates a degenerate loop where the model is not making progress.Under --exec with no --safety, the call itself asked for approval, and with no terminal it was rejected. The model sent the same call again until the runtime stopped the run with exit 1. The limit is max_consecutive_tool_calls, and 0 "uses the default of 5" schema AgentConfig. Inside a running task, "any tool call that would normally prompt the user for approval will be automatically denied" docs-agent Background Agents Tool. A headless coordinator therefore needs an allow rule for run_background_agent and for the tools its sub-agents call. The capture did not test that.
The kit asks for one tool call per message. A turn with two parallel calls cannot be replayed from a cassette, as session.db, sessions diff, and eval explains.
transfer_task sends writer only the task and returns its sentence to root, while the session move makes reviewer answer every later message. Read top down, one step per beat. Solid ink is a call, dashed ink a reply, and dashed amber the handoff call that changes which agent owns the session. The amber box is the writer's sub-session. From capture/out/19-transcript.txt and 19-transfer-task.json.Sources:docs-agent Multi-Agent Systems, Transfer Task Tool, Background Agents Tool (research/sources/docs-docker-agent.md, pages concepts/multi-agent, tools/transfer-task, tools/background-agents); schema AgentConfig (research/sources/agent-schema.json); conflict C56 (research/conflicts-register.md); capture/README.md; capture/fixtures/agents/team.yaml; capture/out/19-team.txt, 19-transcript.txt, 19-transfer-task.json, 19-handoff.json, 19-background.txt, 19-background.ndjson
permissions, --safety, and hooks
Deny, allow, and ask patterns, then a safety mode, then hooks decide whether a tool call runs, and the docs say none of them is a security boundary.
Your agent runs in CI with --exec, and nobody watches the terminal. You want echo to run, rm never to run, and every other call to fail closed unless its label is safe. The agent file and one flag can say that, but only for calls that go through docker-agent.
When you finish this section, you can write permissions patterns, pick a --safety mode for an unattended run, and predict what a pre_tool_use hook changes.
The same page says the restricted mode "is defense in depth against unwanted tool calls, not a security boundary" docs-agent Permissions. For isolation, run the agent in a sandbox, as docker-agent run --sandbox shows.
The guarded agent
version: "16"
agents:
root:
model: local
description: Runs shell commands under a permission list and a hook.
instruction: |
Run exactly the shell commands the user lists, one tool call per
command, in the given order. Then report each command and its
result or refusal in one line each.
max_iterations: 8
toolsets:
- type: shell
hooks:
pre_tool_use:
- matcher: shell
preempt_yolo: true
hooks:
- type: command
command: ./fixtures/hooks/log-hook.sh
env:
M101_HOOK_LOG: ./work/hook-stdin.jsonl
permissions:
allow:
- "shell:cmd=echo*"
deny:
- "shell:cmd=rm*"
…Permission pattern: a tool name glob with optional argument conditions, such as shell:cmd=rm*, in an allow, ask, or deny list. Patterns from the agent file and from settings.permissions in the user config merge, and a deny from either side wins docs-agent Permissions.
Safety mode: what the runtime does with a call that no pattern matched. It reads the call's label, safe, destructive, or unknown, and the mode decides docs-agent Permissions:
| Mode | safe | destructive | unknown |
|---|---|---|---|
strict | ask | ask | ask |
balanced | allow | ask | ask |
restricted | allow | deny | deny |
autonomous | allow | allow | allow |
--yolo is the same as --safety autonomous help-agent docker-agent run. A session that never chooses a mode keeps "the historical default: read-only tools auto-approve, everything else asks" docs-agent Permissions. That is why the wc -l call in run, new, and doctor asked although its label was safe, because shell carries readOnlyHint: false.
Three calls in two modes
The kit sent the same three messages under --safety strict and --safety restricted, one command per message, with no terminal to answer a prompt.
| Command | What decided | strict | restricted |
|---|---|---|---|
echo m101-ok | allow: shell:cmd=echo* | ran, m101-ok | ran, m101-ok |
pwd | no pattern, label safe | asked, then "The user rejected the tool call." | ran, $CAPTURE |
rm -rf work/m101-nothing | deny: shell:cmd=rm* | "Tool 'shell' is denied by permissions configuration." | the same denial |
The patterns behaved the same in both modes, and only the unmatched pwd changed. Under restricted, the mode allowed the safe label. Under strict, the runtime raised tool_call_confirmation, and the empty terminal turned it into a rejection (20-strict.txt, 20-restricted.txt).
The order, and what a hook can change
The docs give one order for every call docs-agent Permissions:
preempt_yolopre_tool_usehooks run first, and no mode or allow rule can bypass their deny or ask.- A
denypattern blocks the call. - An
allowpattern approves it. - An
askpattern prompts the user. - With no match, the safety mode applies to the call's label.
- On a mode ask, default
pre_tool_usehooks can allow, deny, or ask. - With no decision, the user is asked.
Hook: a command, builtin, model, or evaluator entry that runs at a named event schema HookDefinition. A command hook reads one JSON object on stdin and can answer with JSON on stdout. Exit code 2 blocks, and the default timeout is 60 seconds docs-agent Hooks.
{"agent_name": "root", "cwd": "$CAPTURE", "hook_event_name": "pre_tool_use", "safety_policy": "restricted", "session_id": "<uuid>", "tool_input": {"cmd": "rm -rf work/m101-nothing"}, "tool_name": "shell", "tool_use_id": "<call-id>"}log-hook.sh appends that line to a file and prints "permission_decision":"allow". The stream shows it as pre_tool_use_pre_yolo with "allowed": true on all three calls, yet rm was denied and strict still asked for pwd. From a preempting hook, "an allow verdict is advisory" schema HookMatcherConfig. Its deny or ask would have ended the call.
Two more hooks run on every call: tool_input_transform and tool_response_transform appear in each stream, although no file declares them. redact_secrets, true by default, installs a builtin on those events schema AgentConfig. A hook can also run several times at once. When one message asked for three commands, the hook ran three times at once, and two runs appended to the log file together (capture/README.md). Write each record in one write.
rm and echo, and the safety mode alone decides pwd. Read top down. Each question box is one stage of the order, and the box to its right is what the capture saw at that stage. Rose is a denial or rejection, and olive a command that ran. From capture/fixtures/agents/guarded.yaml, capture/out/20-strict.ndjson, 20-restricted.ndjson, and 20-hook-stdin.jsonl.Sources:docs-agent Permissions, Hooks (research/sources/docs-docker-agent.md, pages configuration/permissions, configuration/hooks); schema AgentConfig, HookMatcherConfig, HookDefinition (research/sources/agent-schema.json); help-agent docker-agent run (research/sources/help-docker-agent.md); capture/README.md; capture/fixtures/agents/guarded.yaml; capture/fixtures/hooks/log-hook.sh; capture/out/18-transcript.txt, 20-strict.txt, 20-strict.ndjson, 20-restricted.txt, 20-restricted.ndjson, 20-hook-stdin.jsonl
session.db, sessions diff, and eval
Every run is rows in one SQLite file, a cassette replays its model calls, sessions diff finds the first different tool call, and eval scores saved sessions in containers.
You changed one line of an agent's instruction, and you want to know whether it still does the same work. A live model words every answer differently, so you need records that you can replay and compare.
When you finish this section, you can replay a run from a cassette, compare two runs by tool calls, and score an agent with eval.
session.db
Session: "the record of a conversation, including every message, tool call, sub-agent run, and cost" docs-agent Sessions. It lives in session.db under the data directory, ~/.cagent unless --data-dir or -s points elsewhere help-agent docker-agent run. --session -1 resumes the newest session by creation time docs-agent Sessions.
$ sqlite3 work/data/session.db .tables (python sqlite3)
generated_media_blobs generated_media_manifest migrations session_items sessions sqlite_sequence
…
$ sqlite3 work/data/session.db 'pragma table_info(sessions)' (12 rows)
id created_at tools_approved input_tokens output_tokens title cost send_user_message max_iterations working_dir starred permissions agent_model_overrides custom_models_used thinking parent_id instruction_context safety_policy attributes origin
$ sqlite3 work/data/session.db 'select id, title from sessions order by created_at desc limit 5'
<uuid> | Running agent
…The table held 12 sessions: ten runs, and the two Transferred task sub-sessions of the team runs, which carry a parent id (19-transcript.txt). Each session stores its safety_policy, and session_items holds the messages. The five newest, all --exec runs, are titled Running agent, not a title made from the first message as the docs describe docs-agent Sessions.
Cassettes: --record and --fake
Cassette: a YAML file of the HTTP exchanges between docker-agent and the model. --record writes it, and --fake replays it with no model at all help-agent docker-agent run. 18-cassette-head.txt shows the format: version: 2, then interactions, each with a request to localhost:12434 and the streamed reply.
def start_recording(self, name):
if not self.record_cassettes and os.path.exists(cassette_path(name)):
return None
self.recorded.append(name)
return ["--models-gateway", f"{DMR_URL}/engines", f"--record={CASSETTE_WORK}/{name}"]
…
def replay(name):
return ["--fake", f"{CASSETTE_WORK}/{name}"]The recording found four rules that the help does not state (capture/README.md). --record takes its value only as --record=PATH, and a separate word is read as the agent reference. The path is relative to --working-dir, and .yaml is appended. With the dmr provider the recording proxy answers 400 unless the run also has --models-gateway http://localhost:12434/engines.
Replay matches each request by its body. When one turn issues two tool calls, the runtime runs them in parallel and appends the results in completion order. The next request body then differs, and the replay answers 500, "requested interaction not found". The kit therefore sends one tool call per message.
sessions diff
"Comparison is over the sequence of tool calls, not over the assistant's prose" help-agent docker-agent sessions diff, and the report stops at the first divergence.
$ docker-agent sessions diff -1 -2
Error: unknown shorthand flag: '1' in -1
[exit 1]
$ docker-agent sessions diff -- -1 -2
Comparing -1 (5 turns) against -2 (5 turns)
✅ Identical behaviour across all 5 turns.
[exit 0]
$ docker-agent sessions diff --fail-on-divergence -- -1 -3
Comparing -1 (5 turns) against -3 (6 turns)
❌ First divergence at turn 0 (after 0 matching turn(s)).
-1 called:
list_directory({"path": "."})
-3 called:
shell({"cmd": "echo m101-ok"})
Everything after this point is downstream of the divergence and is not compared.
Error: sessions diverged
[exit 1]The help's own example, sessions diff -1 -2, fails, because the parser reads -1 as a flag. Put -- before the references. -1 and -2 are two replays of files.yaml (21-two-runs.txt), with five turns each, one per model reply in 18-transcript.txt. -3 is the restricted run of guarded.yaml, which called shell first. --json printed {"turns_a": 5, "turns_b": 5, "turns_matched": 5} for the identical pair (21-sessions-diff.json).
eval
Eval session: a JSON session with a user message, the expected tool calls, and an evals block docs-agent Evaluation. The block holds relevance statements, a size, and a working_dir. eval replays each one in a container and scores tool-call F1, relevance by a judge model, and size.
The capture needed three tries with --judge-model dmr/ai/qwen3:4b. Plain, both evals failed with exec: "docker": executable file not found in $PATH, because the image docker/docker-agent:1.149.0 has no docker CLI to find Model Runner (22-eval.txt). With --models-gateway http://model-runner.docker.internal/engines, the judge check failed first with HTTP 403. The docs explain why: "the LLM judge runs on the host, not inside the eval container" docs-agent Evaluation, and the host cannot reach that name (22-eval-gateway.txt). With -e DOCKER_AGENT_MODELS_GATEWAY=…, only the containers changed, and both evals ran.
$ docker-agent eval fixtures/agents/files.yaml fixtures/evals --judge-model dmr/ai/qwen3:4b -c 1 -e DOCKER_AGENT_MODELS_GATEWAY=http://model-runner.docker.internal/engines --output work/eval-results-container-env
…
✓ Count the lines of README.md ($0.000000)
✓ size S
✓ tool calls
✓ relevance 1/1
✗ List the files in the working directory ($0.000000)
✓ size S
✓ relevance 1/1
✗ tool calls score 0.67
…
✅ Sizes: 2/2 passed (100.0%)
✅ Tool Calls: 83.3% avg F1 (2 evals)
✅ Relevance: 2/2 passed (100.0%)
…The second eval expected one list_directory call, and the model added a shell call after it, so F1 fell to 0.67 (22-eval-run-container-env.json). Both eval sessions ran with "safety_policy": "autonomous". The help defaults are -c 10 and the judge openai/gpt-5.6-terra, and the docs table says the number of CPUs and anthropic/claude-opus-5. This manual follows the help (conflict C65). The output directory held <run>.db, <run>.json, and <run>.log, without the -sessions.json file that the docs list (22-eval-results-ls-container-env.txt).
files.yaml match on all five turns, the guarded run differs at turn 0, and eval scores the same agent per tool call. Read the top tracks left to right along the turn axis, one cell per model reply. The rose cell is the first divergence, and dashed cells are not compared. The bottom rows are the eval scores. From capture/out/18-transcript.txt, 20-restricted.txt, 21-sessions-diff.txt, and 22-eval-container-env.txt.Sources:docs-agent Sessions, Evaluation (research/sources/docs-docker-agent.md, pages features/sessions, features/evaluation); help-agent docker-agent run, sessions diff, eval (research/sources/help-docker-agent.md); conflict C65 (research/conflicts-register.md); capture/README.md; capture/run.py; capture/fixtures/agents/files.yaml, guarded.yaml; capture/fixtures/evals/count-lines.json, list-files.json; capture/out/18-cassette-head.txt, 18-transcript.txt, 19-transcript.txt, 20-restricted.txt, 21-session-db.txt, 21-sessions-diff.txt, 21-sessions-diff.json, 21-two-runs.txt, 22-eval.txt, 22-eval-gateway.txt, 22-eval-container-env.txt, 22-eval-run.json, 22-eval-run-container-env.json, 22-eval-results-ls-container-env.txt
docker-agent serve and share
One agent file answers over REST and SSE, an OpenAI-compatible endpoint, MCP, ACP, and A2A, travels as a signed OCI artifact, and runs on a local model.
serve api and serve chat
serve api is the native control plane, with sessions and an SSE run stream on port 8080, and serve chat is the OpenAI-compatible subset on port 8083.
Your agent file answers well in a terminal, and now a web page and a CI job must call it over HTTP. The page wants every event of a turn, and the CI job speaks only the OpenAI chat format.
When you finish this section, you can start both servers on one agent file, run one turn through each, and read the event stream.
One agent file, five servers
The a2a, acp, api, and mcp commands moved under serve in v1.23.4 rel-agent v1.23.4, and v1.53.0 added chat rel-agent v1.53.0. The help of v1.149.0 lists all five help-agent docker-agent serve. This part serves one small file all five ways. Its one agent, root, answers every message with the word pong:
version: "16"
agents:
root:
model: local
description: Answers every message with the word pong.
instruction: Reply with exactly the word pong and nothing else.
models:
local:
provider: dmr
model: ai/qwen3:4b
temperature: 0serve api: a session, then a run
serve api: the command that "exposes your agents through a REST-style API with Server-Sent Events (SSE) streaming" docs-agent API Server. It listens on 127.0.0.1:8080, and an empty --auth-token means no authentication help-agent docker-agent serve api. --max-request-size rejects a body over 1 MiB with HTTP 413, and --session-workingdir-root confines the working_dir of new sessions help-agent docker-agent serve api.
The session database has a different default here. serve api -s writes session.db in the current directory, while run, serve a2a, and serve acp use <data-dir>/session.db help-agent docker-agent serve api (conflict C63). The capture passed -s work/api-session.db and --fake work/cassettes/23-api, so a recorded cassette gave the model answer (capture/out/23-serve-api.log).
…
"name": "pong",
"description": "Answers every message with the word pong.",
"multi": false
…
…
{}
…
"id": "<uuid>",
"origin": "run",
…
"tools_approved": false,
…
…
Accept: text/event-stream
…
"role": "user",
"content": "ping"
…
…
Content-Type: text/event-streamThe agent identifier in the path is the file name without .yaml, so pong, while the events name the agent root docs-agent API Server. The run body is a messages array with an optional model field, which sets a model override for that agent in the session docs-agent API Server. A new session has tools_approved: false. A tool call then raises tool_call_confirmation, and the client answers with POST /api/sessions/:id/resume docs-agent API Server.
The run stream
{"agent_name": "root", "available_agents": [{"description": "Answers every message with the word pong.", "model": "ai/qwen3:4b", "name": "root", "provider": "dmr"}], "current_agent": "root", "timestamp": "<ts>", "type": "team_info"}
{"agent_name": "root", "available_tools": 0, "loading": false, "timestamp": "<ts>", "type": "toolset_info"}
{"message": "ping", "session_id": "<uuid>", "session_position": 0, "timestamp": "<ts>", "type": "user_message"}
{"agent_name": "root", "session_id": "<uuid>", "timestamp": "<ts>", "type": "stream_started"}
{"agent_name": "root", "available_tools": 0, "loading": false, "timestamp": "<ts>", "type": "toolset_info"}
{"agent_name": "root", "description": "Answers every message with the word pong.", "model": "dmr/ai/qwen3:4b", "timestamp": "<ts>", "type": "agent_info"}
{"agent_name": "root", "content": "pong", "message_id": "<uuid>", "session_id": "<uuid>", "timestamp": "<ts>", "type": "agent_choice"}
{"agent_name": "root", "session_id": "<uuid>", "timestamp": "<ts>", "type": "message_added"}
{"agent_name": "root", "session_id": "<uuid>", "timestamp": "<ts>", "type": "token_usage", …}
{"agent_name": "root", "finish_reason": "stop", "reason": "normal", "session_id": "<uuid>", "timestamp": "<ts>", "type": "stream_stopped"}The docs list eight event types, from stream_started to error docs-agent API Server. The capture adds six more: team_info, toolset_info, user_message, agent_info, message_added, and token_usage. The pong agent has no tools, so no tool_call frame appears. The last request of the file reads the session back, and the stored assistant message keeps the model's reasoning_content beside the answer (capture/out/23-api.http). Figure 6.1 draws the whole exchange.
A run is one request, and the session outlives it. /steer injects messages into a running turn, /followup queues them with an optional Idempotency-Key, and /fork copies a session up to a user message docs-agent API Server. GET /api/sessions/:id/events is a session-wide stream that resumes from Last-Event-ID or ?since= docs-agent API Server. An interactive docker-agent run --listen ADDR serves the same control plane, with a fixed 1 MiB body limit and no --auth-token. The help does not list that flag, and each such run writes <data-dir>/runs/<pid>.json for discovery docs-agent API Server.
serve chat: the OpenAI-compatible subset
serve chat: the command that "exposes the agent through an OpenAI-compatible API at /v1/chat/completions and /v1/models" help-agent docker-agent serve chat. The model id is the agent name, so the capture lists root, the name inside the file, where serve api used pong, the file name.
…
Authorization: Bearer m101-chat-key
…
"id": "root",
…
"owned_by": "docker-agent",
…
…
"role": "assistant",
"content": "pong"
…
…
"message": "missing or invalid bearer token",By default the server keeps no conversation, so each request carries the full history. --conversations-max caches up to N conversations keyed by X-Conversation-Id, and --conversation-ttl evicts them after 30 minutes help-agent docker-agent serve chat. A failed turn leaves a cached conversation unchanged, so a client can retry with the same id docs-agent Chat Server. With stream: true in the body, the reply is an SSE stream of chat.completion.chunk deltas docs-agent Chat Server. --max-idle-runtimes keeps four idle runtimes per agent, and --request-timeout bounds each request at five minutes, model and tool calls included help-agent docker-agent serve chat. A listener that is not on loopback needs --api-key, --api-key-env, or --insecure-no-auth docs-agent Chat Server.
serve api | serve chat | |
|---|---|---|
| Default address | 127.0.0.1:8080 | 127.0.0.1:8083 |
| Token flag | --auth-token | --api-key or --api-key-env |
| Agent id in the capture | pong, the file name | root, the agent name |
| State | sessions in -s, default session.db | none, or a cache keyed by X-Conversation-Id |
| Tool approval | tool_call_confirmation, then resume | --safety, default restricted docs-agent Chat Server |
| Cassettes | --fake and --record | none |
Sources:help-agent docker-agent serve, serve api, serve chat (research/sources/help-docker-agent.md); docs-agent API Server, Chat Server (research/sources/docs-docker-agent.md); rel-agent v1.23.4, v1.53.0 (research/sources/docker-agent-CHANGELOG.md); research/conflicts-register.md row C63; research/plan.md, Part 6; capture/README.md; capture/fixtures/agents/pong.yaml; capture/out/23-serve-api.log, 23-api.http, 23-api-run.sse, 23-chat.http
serve mcp and serve acp
serve mcp exposes each agent as one MCP tool over stdio or streaming HTTP on port 8081, and serve acp speaks the Agent Client Protocol to an editor over stdio.
Your MCP client already calls tools, and your editor already talks to coding agents. Neither of them reads an agent file. serve mcp turns the agent into a tool for the first, and serve acp turns it into an editor agent for the second.
When you finish this section, you can serve an agent as an MCP tool, read its schema, and open an ACP session from an editor.
serve mcp: one agent, one tool
serve mcp: "Start an MCP server that exposes the agent via the Model Context Protocol. By default, uses stdio transport" help-agent docker-agent serve mcp. --http switches to streaming HTTP on 127.0.0.1:8081. Without -a, every agent of a team file becomes its own tool, and --tool-name renames the tool when one agent is exposed help-agent docker-agent serve mcp.
--attach exposes the session of a running TUI, found by pid, address, or session id. --mcp-keepalive works over stdio only, while --auth-token, --insecure-no-auth, and --safety work with --http only help-agent docker-agent serve mcp. Over HTTP the safety policy defaults to restricted docs-agent MCP Mode.
The capture ran serve mcp fixtures/agents/pong.yaml --http --listen 127.0.0.1:8081 --tool-name pong, with the file of serve api and serve chat:
…
"method": "initialize",
"params": {
"protocolVersion": "2026-07-28",
…
Content-Type: text/event-stream
…
message
{"jsonrpc":"2.0","id":1,"result":{"capabilities":{"logging":{},"tools":{"listChanged":true}},"protocolVersion":"2025-11-25","serverInfo":{"name":"docker agent","version":"v1.149.0"}}}
…
{"jsonrpc":"2.0","id":2,"result":{"ttlMs":0,"cacheScope":"public","tools":[{"annotations":{"destructiveHint":false,"idempotentHint":true,"openWorldHint":false,"readOnlyHint":true,…"inputSchema":{…"properties":{"message":{"description":"the message to send to the agent","type":"string"}},"required":["message"],"type":"object"},"name":"pong","outputSchema":{…
…
{"jsonrpc":"2.0","id":3,"result":{"content":[{"type":"text","text":"{\"response\":\"pong\"}"}],"structuredContent":{"response":"pong"}}}The client asked for revision 2026-07-28, which v1.129.0 added with "stateless Streamable HTTP transport" rel-agent v1.129.0. The server answered 2025-11-25, so read the version from the result, never from your request. No response carried an Mcp-Session-Id header. Each answer came as one SSE message event, and the notification got 202 Accepted (capture/out/23-mcp.http).
The tool takes one string, message, and returns structuredContent.response, with the same text as a JSON string in content. Its annotations mark it read-only and idempotent, and its title is the agent's description. The tools/call answer came as one event after the whole turn, with no progress notification before it. The kit README adds that the server answered on / as well as /mcp (capture/README.md). The docs register the stdio form in Claude Code with claude mcp add --transport stdio, followed by -- docker agent serve mcp and the agent reference docs-agent MCP Mode.
serve acp: JSON-RPC lines to an editor
serve acp: a server that "communicates over stdio (standard input/output)" docs-agent ACP. The editor spawns docker-agent serve acp FILE, writes JSON-RPC requests to its stdin, and reads responses and notifications from its stdout. Sessions persist in <data-dir>/session.db unless -s names another file help-agent docker-agent serve acp. A team file works as well, and the docs say its sub-agents need no change for ACP docs-agent ACP.
>> {"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {"protocolVersion": 1, …}}
<< {"id": 1, "jsonrpc": "2.0", "result": {"agentCapabilities": {"auth": {"logout": {}}, "loadSession": true, …, "agentInfo": {"name": "docker agent", "title": "docker agent", "version": "v1.149.0"}, …, "protocolVersion": 1}}
>> {"jsonrpc": "2.0", "id": 2, "method": "session/new", "params": {"cwd": "$CAPTURE", "mcpServers": []}}
<< {"jsonrpc": "2.0", "method": "session/update", "params": {"sessionId": "<uuid>", "update": {"availableCommands": [{"description": "Summarize and compact session history", …, "name": "compact"}, {"description": "Display current context token usage and session cost", "name": "usage"}], "sessionUpdate": "available_commands_update"}}}
<< {"id": 2, "jsonrpc": "2.0", "result": {"configOptions": [{"category": "mode", "currentValue": "default", …}], "modes": {…, "currentModeId": "default"}, "sessionId": "<uuid>"}}initialize returns agentCapabilities: loadSession, MCP servers over HTTP and SSE, audio, image, and embedded-context prompts, and the session operations close, delete, list, and resume. It offers one auth method, host-credentials, which v1.144.0 added together with session deletion rel-agent v1.144.0. Before the session/new result, a session/update notification lists the slash commands. Release v1.143.0 announced /compact, /usage, and /new rel-agent v1.143.0, and the captured list has no new.
The session starts in mode default, which auto-approves read-only tools and asks for the rest. The mode list adds default to the four values of --safety. A client changes the mode with session/set_config_option or the older session/set_mode rel-agent v1.144.0. The capture stops at session/new, so it shows no prompt turn over ACP. The docs sketch the host side with a method agent/run and call that code pseudocode docs-agent ACP. Use the method names of the capture.
For the protocol under serve mcp, read the MCP transports lesson.
Sources:help-agent docker-agent serve mcp, serve acp (research/sources/help-docker-agent.md); docs-agent MCP Mode, ACP (research/sources/docs-docker-agent.md); rel-agent v1.129.0, v1.143.0, v1.144.0 (research/sources/docker-agent-CHANGELOG.md); research/plan.md, Part 6; capture/README.md; capture/fixtures/agents/pong.yaml; capture/out/23-mcp.http, 23-acp.jsonl
serve a2a
serve a2a publishes an A2A 1.0 agent card on port 8082 and answers SendMessage with a completed task, and its card breaks three REQUIRED rules of A2A 1.0.1.
A planner on another team speaks A2A and wants to give your agent work. It will fetch a card, send one message, and read a task, as the A2A 101 manual teaches in SendMessage.
When you finish this section, you can serve an agent over A2A, call it with 1.0 names, and predict where a strict client objects.
The card at the well-known path
serve a2a: "Start an A2A server that exposes the agent via the Agent-to-Agent protocol" help-agent docker-agent serve a2a. It listens on 127.0.0.1:8082, and -a "defaults to the team's first agent" help-agent docker-agent serve a2a, where the legacy plugin named root (conflict C64). --auth-token covers the card and every invocation, and a listener off loopback needs it or --insecure-no-auth docs-agent A2A Protocol. Sessions default to the restricted safety policy docs-agent A2A Protocol.
The Docker Agent docs name neither the card path nor the protocol version (conflict C71). The capture found the card at /.well-known/agent-card.json. /.well-known/agent.json answered 404, and so did a POST to / (capture/README.md).
{
"supportedInterfaces": [
{
"url": "http://127.0.0.1:8082/invoke",
"protocolBinding": "JSONRPC",
"protocolVersion": "1.0"
}
],
"capabilities": {
"streaming": true
},
"defaultInputModes": [],
"defaultOutputModes": [],
"description": "Answers every message with the word pong.",
"name": "pong",
"skills": [
{
"description": "Answers every message with the word pong.",
"id": "pong_",
"name": "",
"tags": [
"llm",
"docker agent"
]
}
],
"version": "v1.149.0"
}The card names the file, pong, where serve chat used the agent name root. Its one skill has the id pong_ and the description of the agent. The response also carried Access-Control-Allow-Origin: *, although the run set no --cors-origin and the help says "empty disables CORS" help-agent docker-agent serve a2a.
SendMessage and GetTask
The card points to one JSON-RPC interface, http://127.0.0.1:8082/invoke, at protocol version 1.0. The capture sent SendMessage there with A2A-Version: 1.0 and no configuration:
…
A2A-Version: 1.0
…
"method": "SendMessage",
"params": {
"message": {
"messageId": "m101-0001",
"role": "ROLE_USER",
"parts": [
{
"text": "ping"
}
…
"result": {
"task": {
…
"artifacts": [
{
"artifactId": "<uuid>",
"metadata": {
"adk_partial": true
},
"parts": [
{
"data": {},
…
{
"artifactId": "<uuid>",
"parts": [
{
"text": "pong"
…
"status": {
"state": "TASK_STATE_COMPLETED",
…
"method": "GetTask",
"params": {
"id": "<uuid>"
}The send returned a task already in TASK_STATE_COMPLETED, and GetTask returned the same task with its history (capture/out/23-a2a.http). The answer pong sits in the second artifact, and history keeps only the user message. The server made one live call to ai/qwen3:4b, because serve a2a has no --fake. The card sets streaming: true, so SendStreamingMessage is on offer, but the capture did not call it. Figure 6.3 draws the exchange.
Sessions persist in the session database since v1.59.0, and a client resumes one through /invoke with its A2A context rel-agent v1.59.0. A context id that collides with another session is rejected, and that session stays unchanged docs-agent A2A Protocol. Sessions that older binaries stored are labelled run and cannot be resumed over A2A docs-agent A2A Protocol.
Where it differs from A2A 1.0.1
The A2A 101 manual of this repository states the 1.0.1 rules: AgentCard fields for the card and SendMessage for the call. The table holds each captured value against those rules.
What serve a2a returned | A2A 1.0.1, as A2A 101 states it | Result |
|---|---|---|
the card at /.well-known/agent-card.json | the well-known path of spec §8.2 | match |
one interface: /invoke, JSONRPC, protocolVersion: "1.0" | url, protocolBinding, and protocolVersion are REQUIRED, and the version is Major.Minor | match |
SendMessage, GetTask, ROLE_USER, TASK_STATE_COMPLETED | PascalCase method names and the 1.0 enum names | match |
a send with no configuration returned a terminal task | a send blocks until a terminal or interrupted state | match |
skills[0].name is an empty string | each skill needs an id, a name, a description, and one tag | differs: no skill name |
defaultInputModes and defaultOutputModes are empty lists | a REQUIRED list holds at least one element | differs: two empty lists |
version is v1.149.0 | version is the agent's own release | differs: the docker-agent release |
task metadata keyed by https://google.github.io/adk-docs/a2a/a2a-extension/ | an extension is declared in capabilities.extensions, and its data is keyed by its URI | differs: the card declares no extension |
an artifact whose one part is an empty data object, marked adk_partial | an artifact needs a unique artifactId and at least one part | allowed, but empty |
The adk_* keys come from the A2A library the server is built on, which v1.129.0 moved to adka2a/v2 rel-agent v1.129.0. The docs list four limitations, among them "A2A artifact support not yet integrated" docs-agent A2A Protocol (conflict C70). The capture contradicts that one, because the answer arrives as an artifact. The other three cover tool events, memory, and sub-agents, and the pong agent has none of them, so the capture cannot test them.
The a2a toolset is the client side
An agent file calls a remote A2A agent with a type: a2a toolset. It takes a url, an optional name, and headers, such as a bearer token for --auth-token docs-agent A2A Tool. The schema adds allow_private_ips, to "Opt in to dialling non-public IP addresses (valid for type 'fetch', 'api', 'openapi', 'a2a', and remote MCP toolsets)" schema allow_private_ips. Without it, the client refuses loopback on the direct path, so a call to 127.0.0.1:8082 needs allow_private_ips: true. The capture did not run that toolset. For the protocol itself, read the A2A lesson.
Sources:help-agent docker-agent serve a2a (research/sources/help-docker-agent.md); docs-agent A2A Protocol, A2A Tool (research/sources/docs-docker-agent.md); schema allow_private_ips (research/sources/agent-schema.json); rel-agent v1.59.0, v1.129.0 (research/sources/docker-agent-CHANGELOG.md); research/conflicts-register.md rows C64, C70, C71; research/plan.md, Part 6; manuals/a2a-101/sections/2-1-agentcard-fields.md, 4-1-send-message.md, 2-2-discovery-and-supported-interfaces.md, 3-3-artifact-and-chunks.md, 5-1-json-rpc.md, 6-2-extensions.md; capture/README.md; capture/fixtures/agents/pong.yaml; capture/out/23-agent-card.json, 23-a2a.http
DMR, Compose models, and the docker/mcp-gateway service
A local model is an unauthenticated OpenAI-compatible endpoint on port 12434, Compose binds it to a service as two variables, and four things named gateway stay distinct.
Every model call in Part 5 and in this part ran with no API key and no cloud account. The model was ai/qwen3:4b on Docker Model Runner, on the same Mac as the agent.
When you finish this section, you can reach the runner from the host, a container, and a sandbox, and give it to a Compose service.
Docker Model Runner on port 12434
Docker Model Runner (DMR): the Docker component that pulls models from Docker Hub, an OCI registry, or Hugging Face docs-dmr Docker Model Runner. It serves them over OpenAI, Anthropic, and Ollama compatible APIs docs-dmr DMR REST API. Since Desktop 4.71.0, "Docker Model Runner is now disabled by default and must be explicitly enabled in Settings" docs-desktop 4.71.0. docker desktop enable model-runner --tcp <port> turns on host TCP docs-dmr DMR REST API, and the capture used port 12434.
{
"running": true,
"backends": {
…
"llama.cpp": "Running: llama.cpp b9879-metal (sha256:<digest>) 72874f5",
…
},
"kind": "Docker Desktop",
"endpoint": "http://model-runner.docker.internal/v1/",
"endpointHost": "http://localhost/exp/vDD4.40/v1/"
}"The Model Runner API is not authenticated" docs-dmr Docker Model Runner. Any client that reaches the port can pull, load, and run models. The base URL depends on where the caller runs, and figure 6.5 maps the five that this manual met:
| Caller | Base URL | Evidence |
|---|---|---|
| a process on the host | http://localhost:12434 | capture/out/25-models.json |
| a container on Docker Desktop | http://model-runner.docker.internal | capture/out/26-compose-up.txt |
| a container on Docker Engine | http://172.17.0.1:12434 | docs-dmr DMR REST API, not captured |
| a process on the host, through the Docker socket | http://localhost/exp/vDD4.40 | capture/out/25-model-status.json, endpointHost |
docker-agent inside a sandbox, through its proxy | http://host.docker.internal:12434 | capture/out/27-inside-run.txt |
After the base URL come path prefixes and whole endpoints. The OpenAI paths sit under /engines/v1/, and /engines/llama.cpp/v1/ names the engine docs-dmr DMR REST API. The Anthropic table lists /anthropic/v1/messages, while its examples call /v1/messages (conflict C79). No capture tested either path. Ollama clients use /api/. The capture also met the /v1/ root, which status reports as endpoint.
The 8B tag ai/qwen3:latest failed to pull on the recording Mac, as docker-agent run, new, and doctor explains (capture/out/25-model-pull-latest.txt). With only the 4B tag pulled, doctor resolves auto to dmr/docker.io/ai/qwen3:4b (capture/out/25-doctor.txt). "By default model-runner unloads idle models after a few minutes" docs-agent Docker Model Runner, and provider_opts.keep_alive changes that time. provider_opts.context_size sets the context window through the runner's _configure endpoint docs-agent Docker Model Runner. No capture ran docker model configure, so conflict C77 stays open.
An agent file can also name the runner in a providers: block. fixtures/agents/dmr.yaml sets base_url: http://localhost:12434/engines/llama.cpp/v1, and its run answered Hello (capture/out/25-run-dmr.txt). An explicit base_url bypasses the --record proxy, so that run is live on every capture (capture/README.md).
Compose models
name: m101
models:
llm:
model: ai/qwen3:4b
services:
mcp-gateway:
image: docker/mcp-gateway
use_api_socket: true
command: ["--transport=streaming", "--port=8811"]
printer:
image: alpine:3.22
depends_on: [mcp-gateway]
models:
llm:
endpoint_var: LLM_URL
model_var: LLM_MODEL
…models: the top-level Compose element that declares a model, which a service binds by name. The short syntax derives LLM_URL and LLM_MODEL from the name llm, and endpoint_var and model_var choose the names. Compose v2.38 or later is required (research/sources/docs-compose-models.md). docker compose up pulled and configured the model before it started a container:
llm Pulling
llm Pulled
llm Configuring
llm Configured
…
printer-1 | LLM_MODEL=ai/qwen3:4b
printer-1 | LLM_URL=http://model-runner.docker.internal/v1/
printer-1 | {"object":"list","data":[{"id":"docker.io/ai/qwen3:4b","object":"model","created":0,"owned_by":"docker","dmr":{}}]}
printer-1 | wget: server returned error: HTTP/1.1 401 UnauthorizedCompose injected the /v1/ root, where the plan, read from the model-runner source, expected /engines/v1/. GET ${LLM_URL}models answered from inside the container all the same. Compose cannot start sandboxes (conflict C96), and no capture ran models: on the private engine inside a sandbox (conflict C95).
The docker/mcp-gateway service
The mcp-gateway service runs the open source MCP gateway with the host's Docker API socket, so that it can start MCP servers as containers. In the capture it found no profile and no server, listed 0 tools, and added its own management tools such as mcp-find and mcp-add. It printed Gateway URL: http://localhost:8811/mcp and a bearer token, and the request of printer without that token got 401 (capture/out/26-compose-up.txt). Figure 6.6 draws the stack.
Four things named gateway
| Name in this manual | What it is | Address in the captures |
|---|---|---|
| models gateway | the address set by --models-gateway or DOCKER_AGENT_MODELS_GATEWAY, to "Route all provider traffic through a models gateway URL" docs-agent A2A Protocol | http://localhost:12434/engines, which --record with dmr needs (capture/README.md) |
| sbx MCP gateway | one host-side gateway per sandbox, set up by sbx mcp, separate from the MCP Toolkit docs-sbx MCP gateway | http://mcp-gateway.docker.internal/mcp inside the VM |
| MCP gateway service | docker mcp gateway run or the docker/mcp-gateway image, mcp v0.44.1 on the recording host | http://localhost:8811/mcp inside the service |
| hosted MCP gateway | the gateway of Docker AI Governance, "an invite-only feature" docs-mcp MCP Gateway | not captured |
The Toolkit gateway runs each MCP server in its own container. "Containers for MCP tools are limited to 2 GB" docs-mcp MCP Toolkit, and each one gets 1 CPU. A docker-agent toolset with ref: docker:<name> takes that route, to "Run MCP servers as secure Docker containers via the MCP Gateway" docs-agent MCP Tool. The sbx MCP gateway is sbx mcp add, sbx mcp load, and --static-mcp, and the MCP gateways lesson covers the pattern.
Sources:docs-dmr Docker Model Runner, DMR REST API, Get started (research/sources/docs-model-runner.md); docs-desktop 4.71.0 (research/sources/docs-desktop-release-notes.md); docs-agent Docker Model Runner, A2A Protocol, MCP Tool (research/sources/docs-docker-agent.md); docs-mcp MCP Gateway, MCP Toolkit (research/sources/docs-mcp.md); docs-sbx MCP gateway (research/sources/docs-sandboxes.md); research/sources/docs-compose-models.md; research/sources/mcp-gateway-README.md; research/conflicts-register.md rows C77, C79, C82, C84, C88, C91, C94, C95, C96; research/plan.md, Part 6; capture/README.md; capture/fixtures/agents/dmr.yaml, files-sandbox.yaml; capture/fixtures/compose/compose.yaml; capture/out/25-model-status.json, 25-models.json, 25-model-ls.txt, 25-model-pull-latest.txt, 25-doctor.txt, 25-run-dmr.txt, 26-compose-config.txt, 26-compose-up.txt, 26-compose-down.txt, 27-inside-run.txt, 11-static-inside.txt
Operating It
The same sandbox runs in the cloud, an organization narrows it with policies and reads its audit log, old commands map to new ones, and the claims are checked.
sbx --cloud run, sbx move, and sbx ttl
A cloud sandbox is the same microVM on Docker's compute, billed by the second in five shapes, with a 24 hour TTL ceiling and its own secrets and policies.
The agent in m101-demo is halfway through a long build, your laptop goes into a bag, and the sandbox stops with it. sbx move m101-demo --to cloud captures its filesystem and starts it again on Docker's compute, where a clock ends it and the lid does not. The build process does not move with it, so the agent runs the build again. When you finish this section, you can start a sandbox in the cloud, move one in either direction, and set the clock that ends it.
This edition recorded no cloud run, because a cloud sandbox costs money and that run was not approved. Every fact below comes from the help text and the documentation.
The --cloud flag and what it hides
Cloud sandbox: a sandbox that sbx creates through the Cloud Sandboxes API instead of the local sandboxd, selected with the global --cloud flag help-sbx sbx.
Sandbox Commands:
attach Attach to a cloud sandbox, starting it first if it is stopped
cp Copy files or directories between a sandbox and the host
create Create a sandbox for an agent
exec Execute a command inside a sandbox
ls List sandboxes
move Move a sandbox between local and cloud
ports Manage sandbox port publishing
rm Remove one or more sandboxes
run Run an agent in a sandbox
stop Stop one or more sandboxes without removing them
ttl Inspect or extend a cloud sandbox's TTL
…
volume Manage persistent volumes (cloud-only)The cloud tree drops daemon, prune, settings, and skills, and rm --all is refused with --cloud help-sbx sbx rm (conflict C47). Sandboxes, templates, secrets, volumes, and network policy live in a separate cloud store docs-sbx Compare local and cloud sandboxes.
You need sbx 0.45.1 or later and a pay-as-you-go plan on a Personal or Pro account. The docs page says 0.45.0 and the launch blog says 0.45.1 (conflict C36), and Team and Business accounts are not mentioned (conflict C103). sbx --cloud diagnose checks sign-in, the cloud API, and account access without a local daemon docs-sbx Use cloud sandboxes. Docker Offload, a remote daemon for Docker Desktop, is a different product with no sandbox path (conflicts C98 and C99).
Shapes, prices, and quotas
Without --cpus and --memory a cloud sandbox gets 2 CPUs and 4 GiB help-sbx sbx create. The pair must name one of five billable shapes help-sbx sbx template load. Prices come from the launch blog of 2026-09-24, and the docs tree prints none (conflict C101).
| Shape | vCPU | Memory | Price per hour, blog of 2026-09-24 |
|---|---|---|---|
micro | 1 | 2048 MiB | $0.07 |
small, the default | 2 | 4096 MiB | $0.14 |
medium | 4 | 8192 MiB | $0.28 |
large | 8 | 16384 MiB | $0.56 |
xl | 16 | 32768 MiB | $1.12 |
The blog meters compute per second, and the docs confirm half of that: "Compute isn't billed while it is stopped." docs-sbx Use cloud sandboxes The $250 credit in the same blog is promotional. An account starts with 10 concurrent sandboxes, 50 stored sandboxes, 100 volumes, 100 secrets, and 3 images in preparation docs-sbx-api Account quotas.
sbx --cloud run and attach
sbx --cloud run claude --name cloud-project creates the sandbox and attaches to its agent, or restarts the named sandbox when it exists help-sbx sbx run. Ctrl-\ detaches, and sbx --cloud attach cloud-project joins the session again help-sbx sbx attach. sbx --cloud ports cloud-project --publish 8080 returns a public HTTPS URL, so the service behind it needs its own authentication docs-sbx Use cloud sandboxes. Secrets come from sbx --cloud secret set anthropic, and rules from sbx --cloud policy init, which can be run again help-sbx sbx policy init. HTTP method rules, --protocol, and governance profiles are refused in the cloud docs-sbx Manage cloud network policy.
sbx ttl and --on-timeout
TTL: the time a cloud sandbox lives before its timeout action, one hour by default and at most 24 hours from creation help-sbx sbx ttl.
sbx --cloud ttl +2h cloud-project extends the expiration under that ceiling and never shortens it help-sbx sbx ttl. --on-timeout chooses stop, restart, or delete help-sbx sbx create. Omit it, and the server stops a sandbox it can resume and deletes the rest. restart needs a --ttl of at least one hour, and a volume-backed sandbox must use delete docs-sbx Use cloud sandboxes. Since v0.47.0 a stopped sandbox reports stopped, and its clock restarts on resume rel-sbx v0.47.0. A volume is a snapshot taken when the sandbox exits, and the last sandbox to exit overwrites it help-sbx sbx volume.
sbx move
sbx move m101-demo --to cloud captures the filesystem as a template image, uploads it, and creates a sandbox with a new id, named moved-m101-demo plus a suffix help-sbx sbx move. Kit network rules and the local environment travel inside the image. The host bind mount, the secrets in the store, local policy rules, host port bindings, and running processes stay behind help-sbx sbx move. Credential files written by an interactive sign-in are ordinary files, so they travel unless you delete them first docs-sbx Move a sandbox.
HTTP method rules prompt before the move, because the cloud cannot apply them, and --force answers the prompt and keeps the warning. The destination expires after the server's default of one hour unless --ttl and --on-timeout say otherwise. Moving to local can stage up to 32 GiB in the host's temporary directory help-sbx sbx move. The docs also warn that docker exec inside a cloud sandbox can read the VM filesystem instead of the container's (conflict C104). This edition did not confirm it.
Sources:help-sbx sbx, sbx create, sbx run, sbx attach, sbx rm, sbx move, sbx ttl, sbx volume, sbx policy init, sbx template load (research/sources/help-sbx.md); research/sources/help-sbx-cloud.md; docs-sbx Cloud sandboxes, Authenticate cloud agents, Compare local and cloud sandboxes, Move a sandbox, Manage cloud network policy, Use cloud sandboxes (research/sources/docs-sandboxes.md); docs-sbx-api Compute sizes and limits, Account quotas (research/sources/docs-sandboxes-api.md); rel-sbx v0.45.0, v0.45.1, v0.47.0 (research/sources/sbx-releases.md); blog 2026-09-24 and research/conflicts-register.md rows C36, C47, C98, C99, C101, C103, C104; capture/out/01-help-sbx.txt
sbx policy ls --source org, --profile, and the audit JSONL
An organization writes policies in Docker Home, the daemon pulls them within five minutes, local rules can only narrow, and every decision is written to a rotating JSONL file.
Your security team asks which hosts the agents reached last week, who allowed each one, and whether a developer could have widened the list. On one laptop the answer is sbx policy log, and across a company it is an organization policy with an audit log behind it. When you finish this section, you can say what an organization policy changes on a developer machine and read one audit record.
What an organization can set
Organization policy: a named set of rules written in Docker Home, applied to every local sandbox of the organization or of chosen teams docs-sbx Organization policies. Three kinds exist: Network access, Filesystem access, and MCP access. A network rule is HTTP, one destination with methods and path patterns, or All traffic over TCP, UDP, or both, with Allow or Deny docs-sbx Organization policies. A network policy can require approval, so each destination it allows waits once for the developer's confirmation. An MCP policy is Cedar in the MCP namespace, default deny, where a forbid beats every permit docs-sbx Policy concepts. An organization holds at most 100 policies of 250 rules and 400 KB each docs-sbx Policy concepts.
A change reaches developer machines within 5 minutes, and sbx policy reset forces the pull at the price of every local rule and recorded approval docs-sbx Organization policies. Network rules apply to the next request, filesystem rules only when a sandbox is created, and MCP registration rules at the next sbx mcp add. Sign-in enforcement lists allowedOrgs in a managed com.docker.sbx profile on macOS, the key HKLM\SOFTWARE\Policies\Docker\SBX on Windows, or /etc/docker-sbx/config.json on Linux docs-sbx Sign-in enforcement. sbx login then revokes a credential from any other account. All of this is a separate paid subscription, Docker AI Governance, and local use stays free docs-sbx FAQ.
What an organization cannot do
Governance narrows and never widens. When a policy is active, only organization allow rules grant access, and deny rules from every source still apply docs-sbx Policy concepts. Local and kit allow rules are inactive. The help text for --deny-network says the same:
A local deny even beats an organization approval requirement, so the request is blocked and no prompt appears docs-sbx Policy concepts.
| Rule | Evaluated under organization governance |
|---|---|
| Organization allow | yes |
| Organization deny | yes |
| Local allow | no |
| Local deny | yes |
| Kit-defined allow | no |
| Kit-defined deny | yes |
Governance covers local sandboxes only, and a cloud sandbox uses its own account and sandbox policy, as the cloud section describes docs-sbx Governance. Audit records hold metadata and never prompt content, agent output, or parameter values docs-sbx AI Governance Audit Logs.
This capture has no organization
$ sbx policy ls --source org
No policies match the selected filters.
[exit 0]
$ sbx policy ls --include-inactive
POLICY SOURCE APPLIES TO SUMMARY
local-policy local all network: 194 allow; filesystem read: 1 allow; filesystem write: 1 allow
[exit 0]{
"profiles": [],
"policy_rules_unavailable": false
}No Governance: Managed by <org> line appears, and the docs name that line as the test for an active policy docs-sbx Local audit logs. Every policy check in this capture reports "governance": {"active": false} (capture/out/09-check-verbose.json). Profile: a named group of organization policies that a developer assigns with --profile at creation, listed by sbx policy profile ls help-sbx sbx policy profile ls. When no rule matches and no organization governs the machine, the proxy asks for approval. Its 403 body names sbx policy approval ls (capture/out/09-blocked.txt), a command the v0.47.0 help tree does not list (conflict C16).
The audit record
Audit record: one JSON object per policy decision or daemon session event, written by the daemon and never by the CLI docs-sbx Local audit logs. Records exist since v0.32.0 rel-sbx v0.32.0. The daemon writes them only for a signed-in user with an AI Governance license under an enforced organization policy. This capture therefore produced none, and the directory listing planned as capture K32 was not recorded (conflict C37). The sample record in the docs carries the same reason string that sbx policy log printed for m101-policy:
"blocked_hosts": [
{
"host": "example.com:443",
"vm_name": "m101-policy",
"proxy_type": "forward",
"rule": "no applicable policies for op(action=net:connect:tcp, resource=net:domain:example.com:443)",
…
"reason": "No matching allow rule (default deny)"{
"audit_event_id": "95e7257f-93c9-4f29-bde7-88830e2dae80",
"timestamp": "2026-05-28T19:15:00.728933Z",
"schema_version": "1.82.0",
"category": "AUDIT_CATEGORY_EVALUATION",
"decision": "AUDIT_DECISION_DENY",
…
"resource_id": "example.com:443",
"os": "macos",
"app_version": "v0.31.0",
"client_name": "sbx",
"hostname": "host-machine",
"deny_reason": [
"no applicable policies for op(action=net:connect:tcp, resource=net:domain:example.com:443)"
],
"action_type": "network_egress",
"network_egress": { "protocol": "tcp" },
"agent": "claude"
}On macOS the files are audit-<utc-timestamp>-<process-uuid>-<seq>.jsonl under ~/Library/Logs/com.docker.sandboxes/sandboxes/auditkit/ docs-sbx Local audit logs. The daemon finalizes a .tmp file into .jsonl every 5 minutes, 1000 events, or 50 MiB, and never deletes one. category is management, evaluation, or execution, and decision is one of five values from AUDIT_DECISION_ALLOW to AUDIT_DECISION_APPROVAL_DENY docs-sbx Audit record reference. action_type names the payload, such as network_egress, http_request with method, host, port, and path, or tool_invocation.
client_name is sbx and hostname names the machine, which is how a GitHub Actions run with runtime: docker-sbx appears blog 2026-08-21. Docker Cloud delivery is on by default and keeps events searchable for 90 days, with CSV export up to 1 000 000 rows docs-sbx View and export audit events. From 0.39.0 on it also forwards to Splunk Cloud, Dynatrace, Datadog, or Sumo Logic docs-sbx SIEM forwarding.
Sources:docs-sbx Governance, Organization policies, Policy concepts, Local policy, Monitoring policies, AI Governance Audit Logs, Local audit logs, Audit record reference, Configure audit delivery, SIEM forwarding, View and export audit events, Sign-in enforcement, FAQ (research/sources/docs-sandboxes.md); help-sbx sbx create, sbx policy ls, sbx policy profile ls (research/sources/help-sbx.md); rel-sbx v0.32.0, v0.39.0 (research/sources/sbx-releases.md); blog 2026-08-21 and research/conflicts-register.md rows C16, C28, C37, C41; capture/out/03-policy-org.txt, 03-policy-profile-ls.txt, 09-check-verbose.json, 09-policy-log.json, 09-blocked.txt
From docker sandbox and cagent to sbx and docker-agent
Two renames and one CLI restructure changed every command in the 2025 and early 2026 material, and this section maps each old line to its current form.
A blog post from 2026-02-23 tells you to run docker sandbox run claude and to point the proxy at host.docker.internal:3128. A talk from 2025-09-24 shows cagent run agent.yaml and cagent push. On a machine with Docker Desktop 4.94.0, sbx 0.47.0, and docker-agent 1.149.0, neither command exists. When you finish this section, you can map any command from that material to its current form and name the version that changed it.
Two tracks and three removals
The sandbox product had four lives. The plan's timeline dates a container-based docker sandbox run to Docker Desktop 4.50.0 on 2025-11-06, and the preview blog followed on 2025-11-25 blog 2025-11-25. Desktop 4.58.0 replaced it with microVMs on 2026-01-26 docs-desktop 4.58.0, and Desktop 4.61.0 bundled plugin v0.12.0 on 2026-02-18 docs-desktop 4.61.0. That plugin is the one whose help this section quotes. The standalone sbx binary had its first public tag, v0.21.0, on 2026-03-31 rel-sbx v0.21.0. Desktop 4.80.0 then ended the plugin on 2026-06-29:
No docker sbx command exists in any help tree, and the binary is sbx (conflict C1).
The agent product had three names. It launched as cagent with a blog on 2025-09-18 blog 2025-09-18, and Desktop 4.49.0 bundled it on 2025-10-23. The Desktop notes for 4.49.0 now read "Docker Agent is now available through Docker Desktop." docs-desktop 4.49.0, under the current name. The agent docs say "In Docker Desktop versions 4.49 through 4.62, this feature was called cagent." docs-agent Installation, and the docker agent plugin arrived with v1.23.3 on 2026-02-16 rel-agent v1.23.3.
Version 1.23.4 restructured the commands three days later, and v1.30.0 finished the rename on 2026-03-09 rel-agent v1.30.0. The docs date Docker Agent in Desktop to 4.63, while the first release note that names it is 4.64.0 with v1.27.1 (conflict C49). Desktop 4.81.0 removed the deprecated cagent binary on 2026-07-06 docs-desktop 4.81.0. Desktop 4.94.0 bundles Docker Agent v1.144.0 docs-desktop 4.94.0, five releases behind the v1.149.0 binary this manual pins. Many Desktop notes state no bundled version at all (conflict C50).
The old command trees
Management Commands:
create Create a sandbox for an agent
network Manage sandbox networking
Commands:
exec Execute a command inside a sandbox
ls List VMs
reset Reset all VM sandboxes and clean up state
rm Remove one or more sandboxes
run Run an agent in a sandbox
save Save a snapshot of the sandbox as a template
stop Stop one or more sandboxes without removing them
version Show sandbox version information
…
--bypass-cidr string Bypass MITM proxy for an IP range in CIDR
notation (can be specified multiple times)
--bypass-host string Bypass MITM proxy for a domain or IP (can be
specified multiple times)
…
--policy allow|deny Set the default policyCore Commands:
new Create a new agent configuration
run Run an agent
share Share agents
Advanced Commands:
alias Manage aliases
eval Run evaluations for an agent
serve Start an agent as a server
…
--sandbox Run the agent inside a Docker sandbox (requires Docker Desktop with sandbox support)
…
--template string Template image for the sandbox (passed to docker sandbox create -t)
…
--yolo Automatically approve all tool calls without promptingBeside them, the v0.47.0 tree in capture/out/01-help-sbx.txt groups policy, secret, template, and mcp under Management Commands, and network and save are gone. The v1.149.0 tree in capture/out/01-help-docker-agent.txt adds setup, doctor, models, toolsets, sessions, debug, sandbox, and serve chat, which Part 5 covers.
The command map
| Old command or habit | Where it appears | Stopped being true | Do this now |
|---|---|---|---|
docker sandbox run <agent> | Desktop docs, blogs, videos | Desktop 4.80.0, 2026-06-29 | sbx run <agent> |
docker sandbox run --mount-docker-socket kiro | re:Invent blog, 2025-12-12 | Desktop 4.58.0, 2026-01-26 | nothing: every sandbox has a private Docker Engine |
--load-local-template | Desktop 4.58 to 4.60 | Desktop 4.61, 2026-02-18 | sbx template load FILE, then --pull never -t TAG |
--pull-template missing, the default | plugin v0.12.0 | v0.21.0, 2026-03-31 | --pull always, missing, or never, default always |
docker sandbox create cagent . | plugin v0.12.0, blog 2026-03-11 | Desktop 4.80.0 | sbx create docker-agent ., and cagent remains an alias |
docker sandbox network proxy S --allow-host api.example.com | legacy docs, blog 2026-02-23 | Desktop 4.80.0 | sbx policy allow network api.example.com --sandbox S |
--bypass-host, --bypass-cidr | plugin v0.12.0 | Desktop 4.80.0 | no replacement in the v0.47.0 help, see capture/out/09-bypass-grep.txt |
docker sandbox network log --json | legacy docs | Desktop 4.80.0 | sbx policy log S --json |
docker sandbox save S TAG into host Docker | plugin v0.12.0, blog 2026-02-23 | Desktop 4.80.0 | sbx template save S TAG -o FILE into the sandbox runtime's own store |
docker sandbox exec -d | plugin v0.12.0 | sbx | sbx exec -d is "not supported" |
names with _, +, or . | plugin v0.12.0 | v0.43.0, 2026-09-15 | 2 to 63 characters, letters, digits, hyphens, periods, and no periods with --cloud |
a proxy set by hand at host.docker.internal:3128 | blogs 2026-02-23 and 2026-05-26 | sbx | nothing: the daemon sets gateway.docker.internal:3128, see capture/out/04-env.txt |
API keys in ~/.zshrc, then restart Desktop | tutorials before 2026-07 | v0.35.0, 2026-07-10 | sbx secret set SERVICE or sbx secret import |
sbx run claude --branch | blog 2026-05-26, videos | v0.31.0, 2026-05-28 | sbx run claude --clone |
kit v1 grammar, schemaVersion: "1", network.allowedDomains | blog 2026-08-03 | v2 on 2026-09-09, v3 on 2026-09-24 | v2 spec.yaml for sbx kit, v3 kit.yaml with # syntax=docker/sandbox-kit:3 |
sbx mcp catalog | docs before 0.45.0 | v0.45.0, 2026-09-21 | sbx mcp add --url URL |
sbx mcp enable github-official | product page, 2026-10-07 | never existed | sbx mcp add, then --static-mcp or sbx mcp load |
-p 3000:8080 binds IPv4 and IPv6 | before v0.42.0 | v0.42.0, 2026-09-07 | tcp4 by default, write 3000:8080/tcp for both |
| kits from any registry | before v0.34.0 | v0.34.0, 2026-06-26 | add the prefix to kit.allowedSources |
a relative --command ./helper secret source | docs before v0.46.0 | v0.46.0, 2026-09-28 | an absolute path outside writable sandbox mounts |
| Windows 10 | early 2026 | v0.35.0, 2026-07-10 | Windows 11 with Windows Hypervisor Platform |
Toolkit gateway at host.docker.internal:8811, five manual steps | re:Invent blog, 2025-12-12 | v0.38.0, 2026-08-06 | sbx mcp add and sbx mcp load, the Toolkit gateway is separate |
| "one sandbox per workspace" | blog 2026-03-11 | sbx names default to <agent>-<workdir> | --name for a second sandbox on one directory, see capture/out/04-second-sandbox.txt |
cagent run agent.yaml, cagent new, cagent exec | 2025 blogs and talks | v1.23.4, 2026-02-19, and Desktop 4.81.0 | docker agent run, docker agent new, docker agent run --exec |
cagent push, cagent pull, cagent acp, api, mcp, a2a | 2025 blogs | v1.23.4 | docker agent share push and pull, docker agent serve acp, api, mcp, a2a |
cagent config, feedback, build, catalog | 2025 to early 2026 | v1.23.4 | removed, and the catalog is the Hub namespace agentcatalog/* |
cagent version, brew install cagent | blog 2025-11-13 | v1.30.0 and Desktop 4.81.0 | docker agent version, brew install docker-agent, winget install Docker.Agent |
| "Docker Agent v2.x renamed the CLI" | a third-party tutorial | never true | the latest tag is v1.149.0 |
docs.docker.com/ai/cagent/ | blog 2026-03-11, old links | v1.30.0 | docs.docker.com/ai/docker-agent/ |
CAGENT_MODELS_GATEWAY, CAGENT_CONFIG_DIR, CAGENT_PPROF_ADDR | older docs and issues | v1.30.0, still accepted | DOCKER_AGENT_MODELS_GATEWAY, DOCKER_AGENT_CONFIG_DIR, DOCKER_AGENT_PPROF_ADDR |
docker/sandbox-templates:cagent | legacy agent page | the rename | docker/sandbox-templates:docker-agent |
docker cp the agent binary into a sandbox, then agent run dev-team.yaml inside | blog 2026-03-11 | the docker-agent template | sbx run docker-agent . or docker agent run --sandbox agent.yaml |
--yolo | everywhere | still accepted | --safety autonomous, of which --yolo is the alias |
| the eval judge is a paid Anthropic model | v1.32.4 help | v1.147.0, 2026-10-05 | default openai/gpt-5.6-terra, or any provider/model |
--sandbox "requires Docker Desktop with sandbox support" | v1.32.4 help | sbx | --sbx defaults to true, and --sbx=false has no target after Desktop 4.80.0 |
| mcp-gateway v0.22.0 in the GitHub Action | the action's defaults | the repository is at v0.44.1 | pin the gateway version yourself |
The rows restate the register's stale-advice list S1 to S30 and the plugin rows C107 to C112, which the sources section prints in full.
What kept the old name
The rename did not touch the directories. docker-agent --help still defaults --data-dir to ~/.cagent, --config-dir to ~/.config/cagent, and --cache-dir to ~/Library/Caches/cagent help-agent docker-agent. The v1.32.4 plugin printed the same three values, and the debug log stays at ~/.cagent/cagent.debug.log (conflict C54). The CAGENT_* variables are still read next to DOCKER_AGENT_* rel-agent v1.30.0.
The OCI annotation became io.docker.agent.version with the old name kept rel-agent v1.23.3. The Hub image moved from docker/cagent to docker/docker-agent, and the image keeps cagent as a symlink rel-agent v1.30.0. In sbx, cagent remains an alias of the docker-agent create subcommand help-sbx sbx create docker-agent.
Sandbox state did not migrate. The plugin kept VM state under ~/.docker/sandboxes/vm/ and an image cache under ~/.docker/sandboxes/image-cache/, as its own reset help says. sbx keeps its socket and log under ~/Library/Application Support/com.docker.sandboxes/sandboxes/sandboxd/ (capture/out/02-daemon-status.txt), and no page describes a migration (conflict C17). No capture lists the leftover directories.
On Desktop 4.94.0, docker info still lists the plugin as sandbox v0.13.0 (29-docker-plugins.txt), and each of its commands prints the removal notice:
$ docker sandbox --help
"docker sandbox" is deprecated and has been removed.
Please migrate to Docker Sandboxes: https://www.docker.com/products/docker-sandboxes
[exit 1]The notice names the product page, not the docker sbx of the 4.80.0 note. On the same host, the bundled docker agent version prints v1.144.0 and the Homebrew docker-agent version prints v1.149.0 (29-docker-agent-plugin.txt).
Sources:docs-desktop 4.49.0, 4.58.0, 4.61.0, 4.64.0, 4.80.0, 4.81.0, 4.94.0 (research/sources/docs-desktop-release-notes.md); rel-agent v1.23.3, v1.23.4, v1.30.0, v1.147.0 (research/sources/docker-agent-CHANGELOG.md); rel-sbx v0.21.0, v0.31.0, v0.35.0, v0.42.0, v0.43.0, v0.45.0 (research/sources/sbx-releases.md); docs-agent Installation (research/sources/docs-docker-agent.md); help-sbx sbx create docker-agent (research/sources/help-sbx.md); help-agent docker-agent (research/sources/help-docker-agent.md); research/sources/help-legacy-docker-sandbox.md; research/sources/help-legacy-docker-agent.md; blog 2025-09-18, 2025-11-25, and research/conflicts-register.md rows C1, C17, C49, C50, C54, C107 to C112, S1 to S30; research/plan.md (the timeline); capture/out/01-help-sbx.txt, 01-help-docker-agent.txt, 02-daemon-status.txt, 04-env.txt, 04-second-sandbox.txt, 09-bypass-grep.txt, 29-docker-sandbox.txt, 29-docker-plugins.txt, 29-docker-agent-plugin.txt
kit-tck, sbx diagnose, and the claims this manual checked
Every claim this manual repeats from Docker's pages or the community was checked against a capture or marked unverified, and this section is the ledger.
A vendor page says that credential values never enter the VM docs-sbx Security model. A Hacker News thread from 2026-08-10 says the proxy re-signs every certificate, and a podcast from 2026-09-26 says nobody audited the boundary. You want to know which of those three someone checked. When you finish this section, you can say which claims this manual verified, with which file, and which it only reports.
The ledger
Each row names the claim and its source, the verdict, and the capture file behind the verdict. A verdict of confirmed means a recorded file shows the behaviour. Contradicted means a file shows the opposite. Disputed means two sources disagree and no capture settles it, and unverified means no capture was planned or recorded.
| Claim and where it is made | Verdict | Evidence |
|---|---|---|
| credential values never enter the VM (docs, Security model) | confirmed | 10-env-sentinel.txt has M101_RECV_KEY=sbx-cs-<rand> and ANTHROPIC_API_KEY=proxy-managed, and 10-receiver.log shows the host receiver got Bearer m101-dummy-receiver-0000 |
| the secret is injected only for a bound host, and the whole header is replaced (HN 2026-08-10, C34) | confirmed, both sides | 10-swap-curl.txt: 204 on host.docker.internal:18080 and 000 on gateway.docker.internal:18080, while 10-receiver.log shows the swap with and without Bearer, and in X-Demo |
| outbound TCP is blocked unless a rule allows it (docs, Default security posture) | confirmed | 09-blocked.txt: 403, and 09-check-verbose.json: "deny_kind": "implicit" |
| each sandbox has its own Docker Engine with no path to the host daemon (docs, Security model) | confirmed in part | 04-guest.txt: Docker Engine 29.8.1 server inside, user agent in group docker, and no probe for a host path was run |
--clone commits reach the host through sandbox-<name> (help, sbx create) | confirmed | 08-clone-host.txt: git fetch sandbox-m101-clone brought 81e11df from sandbox, and 08-clone-commit.txt: a push to /run/sandbox/source is rejected |
| the guest uses 16 KiB pages on Apple silicon (blog 2026-05-26, C45) | confirmed | 04-guest.txt: getconf PAGESIZE prints 16384 |
the proxy is gateway.docker.internal:3128 (issue #12, C26) | confirmed | 04-env.txt: HTTPS_PROXY=http://gateway.docker.internal:3128, and PROXY_CA_CERT_B64 set |
| HTTPS is intercepted and re-signed (issue #12, HN 2026-08-10) | not checked | 09-allowed.txt shows the CONNECT tunnel and a cloudflare reply, and no certificate chain was read |
| one MCP gateway endpoint per sandbox (docs, Architecture) | confirmed | 11-static-inside.txt: MCP_GATEWAY_URL=http://mcp-gateway.docker.internal/mcp, and 11-gateway-tools-static.json: ask_wiki_question |
27 toolset types, and transfer_task is not one of them (schema, C56) | confirmed | 16-toolsets.txt: 27 rows, neither implicit name |
the sbx kit commands are v1 and v2 tooling (C13) | confirmed | 13-kit-v3-validate.txt: "is a v3 source kit and this load path has no kit builder configured", and 12-validate.txt: VALID |
a kit signature verifies with sbx kit sign and verify (docs, Build and distribute kits) | not checked | no sign or verify step was recorded, only 13-kit-builder-status.txt |
| a local model runs with no key (docs, Use local and hosted models) | confirmed | 25-doctor.txt: every provider credential not set, Docker Model Runner reachable with docker.io/ai/qwen3:4b, and 25-run-dmr.txt answers Hello live. 18-cassette-head.txt shows the recorded requests going to localhost:12434, and 27-inside-run.txt answers from inside the VM. The 8B ai/qwen3:latest failed to pull (25-model-pull-latest.txt) |
| local use needs a Docker sign-in (docs, FAQ, disputed in HN 2026-08-10 and issue #321, C15) | confirmed in part | 02-diagnose.txt lists Authentication as a check, the kit's prerequisites include sbx login, and help-sbx-cloud.md notes a TLS attempt to login.docker.com on every --help |
sbx policy approval is not a command (the v0.47.0 help tree, C16) | unverified | 09-blocked.txt: the 403 body says "Review and respond with: sbx policy approval ls", and no capture ran that command |
| the VMM is libkrun (talk 2026-01-14, HN 2026-08-10, C4) | unverified | 00-sbx-version.json names no runtime, and 04-guest.txt shows kernel 7.0.14 and nothing more |
| a sandbox starts in tens of milliseconds (talk 2026-02-10, C29) | not checked | the timing runs planned as K35 were not recorded |
| an allowed host can carry data out (HN 2026-08-10, talk 2026-04-07) | disputed | both sides are printed in the proxy section, and no capture tested it |
| a microVM HTTP API, a Kubernetes runtime, Warp Oz (blog 2026-05-26, blog 2026-08-21, talk 2026-03-31) | unverified | no Docker page confirms any of them (C105, C106) |
| an independent audit of the boundary exists | none found | the podcast of 2026-09-26 says none exists, and the corpus has none (T24) |
sbx ls hangs on macOS (#163), Windows start fails (#350) | not reproduced | 99-final-state.txt: sbx ls answered at the end of the run |
| a sandbox stops itself after the last session ends (not in the docs) | observed | 04-auto-stop.txt: "auto-stop grace period expired, stopping runtime" |
sbx diagnose, the only self-check
sbx diagnose: the one command that checks an installation, in the four groups that capture/out/02-diagnose.txt shows: Installation, Platform, Storage, and Connection. It takes --json, --output github-issue for a report to paste into an issue, and --upload to send diagnostics to Docker support help-sbx sbx diagnose.
Platform
✓ Virtualization — supported
kern.hv_support is 1
✓ mkfs.erofs — found
/opt/homebrew/Caskroom/sbx/0.47.0/Sbx.app/Contents/libexec/mkfs.erofs, default block size <n> bytes, guest kernel page size <n> bytes
…
Connection
✓ Version match — v0.47.0
✓ Socket — responsive
✓ SSH client config — not configured
✓ Authentication — authenticated
─────────────────────────────────────────
13 passedAll 13 checks passed, and the JSON form in 02-diagnose.json carries the same rows with a summary of pass, warn, fail, and skip counts. The Platform group is where the microVM boundary section reads the hypervisor and the page size. sbx --cloud diagnose runs a different set, sign-in, the cloud API, and account access, and needs no daemon docs-sbx Use cloud sandboxes.
kit-tck, not run in this edition
kit-tck: the conformance suite of the Sandbox Kit Specification v3, with two independent halves, one for a published kit artifact and one for a runtime. The conformance page of the specification repository describes both (research/sources/kit-spec-extras.md). kit-tck validate <reference> judges a published artifact against the publishing and OCI layout rules, and the BuildKit frontend runs the same checks before it exports. kit-tck runtime --adapter <path> drives an adapter, an executable with verbs such as capabilities, create, exec, stop, start, recreate, and rm.
The suite judges what the runtime does through those verbs. An adapter exits 2 when it refuses a request by policy, and any other non-zero code is a failure. A runtime that cannot provide a required capability must therefore refuse the kit.
task tck:kit REF=docker.io/me/sbx-kit-gh:1.0.0 # is this artifact a conforming Kit?
task tck:runtime ADAPTER=./my-adapter # does this runtime behave as the pages require?
…
kit-tck validate docker.io/me/sbx-kit-gh:1.0.0 --verbose
kit-tck validate docker.io/me/sbx-kit-gh:1.0.0 --format jsonReleased kit-tck binaries are attached to each GitHub release of the specification for Linux, macOS, and Windows. kit-tck inspect reads a kit's descriptor and recipe without judging it. The suite places three known values in the adapter's environment, KIT_TCK_HOST_SENTINEL, KIT_TCK_BOUND_SECRET, and KIT_TCK_SKILL_NAME. A leak of host environment or of a bound secret is then visible.
This edition did not run it. The capture kit built a v3 artifact and pushed it to a local registry (28-buildx.txt, 28-manifest.json), but no kit-tck step ran against it.
The agent regression check
For agent runs, the planned check is docker-agent sessions diff --fail-on-divergence, which compares two recorded sessions "over the sequence of tool calls, not over the assistant's prose" and stops at the first divergence help-agent docker-agent sessions diff. The capture kit ran it on recorded sessions (21-sessions-diff.txt). Two replays of one run are identical across five turns, and a run that called shell first diverges at turn 0 and exits 1. A relative reference such as -1 needs -- before it, or the parser reads it as a flag. The sessions section explains the session store.
Sources:docs-sbx Security model, Default security posture, Architecture, FAQ, Use cloud sandboxes, Build and distribute kits, Use local and hosted models (research/sources/docs-sandboxes.md); help-sbx sbx diagnose, sbx create (research/sources/help-sbx.md); help-agent docker-agent sessions diff (research/sources/help-docker-agent.md); research/sources/kit-spec-extras.md (README Conformance, docs/spec/conformance.md); research/sources/help-sbx-cloud.md; research/conflicts-register.md rows C4, C13, C15, C16, C26, C29, C34, C45, C56, C105, C106 and the community rows of research/plan.md; capture/out/00-sbx-version.json, 02-diagnose.txt, 02-diagnose.json, 04-auto-stop.txt, 04-env.txt, 04-guest.txt, 08-clone-commit.txt, 08-clone-host.txt, 09-allowed.txt, 09-blocked.txt, 09-check-verbose.json, 10-env-sentinel.txt, 10-receiver.log, 10-swap-curl.txt, 11-gateway-tools-static.json, 11-static-inside.txt, 12-validate.txt, 13-kit-builder-status.txt, 13-kit-v3-validate.txt, 16-toolsets.txt, 18-cassette-head.txt, 21-sessions-diff.txt, 25-doctor.txt, 25-model-pull-latest.txt, 25-run-dmr.txt, 27-inside-run.txt, 28-buildx.txt, 28-manifest.json, 99-final-state.txt
Reference
Lookup tables for every name in the manual.
sbx commands
Every verb of sbx 0.47.0 in the four groups of sbx --help, with its purpose, the flags that matter, its scope, and the section that explains it.
The root help of sbx 0.47.0 lists 31 commands in four groups: Sandbox, Management, Experimental, and Other help-sbx sbx. The vendored tree holds 115 help pages, one for each command and subcommand that answered --help. The tables below keep the order of the root help. Scope says where a verb runs: local sandboxes through sandboxd, cloud sandboxes through sbx --cloud, or both. sbx --cloud --help hides daemon, prune, settings, and skills, and attach, ttl, and volume run only with --cloud (conflict C47).
Three flags apply to every command help-sbx sbx.
| Global flag | Meaning |
|---|---|
--cloud | Dispatch to the Docker Cloud Sandboxes API instead of the local sandboxd |
-D, --debug | Enable debug logging |
-h, --help | Print the help of the command |
Sandbox commands
| Command | Purpose | Flags that matter | Scope | Section |
|---|---|---|---|---|
attach SANDBOX | Attach a terminal to a cloud sandbox, and start it first when it is stopped help-sbx sbx attach | --detach-keys (default Ctrl-\) | cloud only | §7.1 |
cp SRC DST | Copy a file or a directory between a sandbox and the host, with one side written as SANDBOX:PATH help-sbx sbx cp | -L, --follow-link | both | §2.2 |
create AGENT|SANDBOX_KIT [PATH...] | Create a sandbox for one of the 11 agents or from a kit reference, without attaching help-sbx sbx create | --name (default <agent>-<workdir>), --clone, --cpus (0 is auto), -m, --memory (default 50% of host memory, 512 MiB to 32 GiB), -e, --env, --env-file, --deny-network, -p, --publish, --pull always|missing|never (default always), --skills off|readonly|readwrite (default readonly), --static-mcp, -t, --template, --profile, --kit (experimental), --kit-arg, --kit-args-file, -q | both | §2.1 |
create <agent> [PATH...] | The eight subcommands claude, codex, cursor, devin, docker-agent, gemini, opencode, shell, with the same flags, and cagent as an alias of docker-agent help-sbx sbx create docker-agent | the flags of create | both | §2.1 |
create with --cloud | Create a cloud sandbox with no host workspace, 2 CPUs and 4 GiB unless sized help-sbx sbx create | --allow-network, --image-ref, --on-timeout stop|restart|delete, --platform, --ttl, -v, --volume (experimental) | cloud only | §7.1 |
exec SANDBOX COMMAND | Run a command inside a sandbox, and start a stopped one first help-sbx sbx exec | -i, -t, -u, --user, -w, --workdir, -e, --env-file, --privileged, --detach-keys. -d is not supported. -d, --user, and --privileged are rejected with --cloud | both | §2.2 |
ls | List sandboxes with agent, status, published ports, and workspace help-sbx sbx ls | --json, -q | both | §2.1 |
move SANDBOX | Move a sandbox between the host and the cloud as a filesystem image help-sbx sbx move | --to local|cloud, --name (default moved- plus the source name), --ttl (15s to 24h), --on-timeout stop|delete, -f | both | §7.1 |
ports SANDBOX | List, publish, or unpublish sandbox ports help-sbx sbx ports | --publish, --unpublish, --json. A port without a protocol binds tcp4 | both | §2.2 |
prune | Remove all stopped sandboxes and their sandbox-scoped secrets help-sbx sbx prune | --dry-run, --filter until=TIMESTAMP, -f, --json | local only | §2.1 |
rm [SANDBOX...] | Remove sandboxes, their containers, git worktrees, state, and scoped secrets help-sbx sbx rm | --all (disabled with --cloud), -f | both | §2.1 |
run [AGENT|SANDBOX_KIT] [PATH...] [-- AGENT_ARGS...] | Run an agent and create the sandbox when it does not exist help-sbx sbx run | the flags of create, plus -d, --detached, --rm, --name to re-attach, --new (cloud), --detach-keys (cloud). --clone works only at creation | both | §2.1 |
stop SANDBOX... | Stop sandboxes and keep their state help-sbx sbx stop | none | both | §2.1 |
ttl [+DURATION] SANDBOX | Print or extend the TTL of a cloud sandbox under its 24 hour ceiling help-sbx sbx ttl | --json | cloud only | §7.1 |
The 11 agent names that run and create accept are claude, codex, copilot, cursor, devin, docker-agent, droid, gemini, kiro, opencode, and shell help-sbx sbx run. Only eight of them have a create subcommand, because copilot, droid, and kiro are public kits (conflict C9).
Management commands
| Command | Purpose | Flags that matter | Scope | Section |
|---|---|---|---|---|
daemon start | Start sandboxd help-sbx sbx daemon start | -d, --detach, --policy allow-all|balanced|deny-all | local only | §2.4 |
daemon stop, daemon restart | Stop or restart sandboxd help-sbx sbx daemon | none | local only | §2.4 |
daemon status | Print the daemon state, its socket, and its log path help-sbx sbx daemon status | --json | local only | §2.4 |
daemon log-level set <target> <level> | Set the log level of proxy, general, or all help-sbx sbx daemon log-level set | none | local only | §2.4 |
diagnose | Run the 13 installation checks help-sbx sbx diagnose | --json, -o json|github-issue, --upload | both | §2.4 and §7.4 |
mcp add <name> (--url | --command) | Register an MCP server from a remote endpoint, a registry URL, a manifest URL, a dhi.io image, or a host command help-sbx sbx mcp add | --url, --command, --args, --dir, --local, --header, --client-id, --oauth-authorization-server, --scope, --no-scope, --resource, --callback-port, --skip-auth, --skip-ssrf-check, --disable-http2 | both | §3.6 |
mcp auth [server-name] | Authorize or reauthorize remote servers through the hosted control plane help-sbx sbx mcp auth | --all, --scope, --no-scope, --verbose, --format text|json, --json | both | §3.6 |
mcp auth rm, mcp auth status | Remove hosted OAuth credentials, or show their status without starting OAuth help-sbx sbx mcp auth status | --all, -f (rm only), --format, --json | both | §3.6 |
mcp inspect <name> | Show one registration, its headers, and the effective resource value help-sbx sbx mcp inspect | --json | both | §3.6 |
mcp load <name> --sandbox S | Attach a registered server to a running sandbox, with a tools/list_changed notification to the agent help-sbx sbx mcp load | --sandbox (required) | both | §3.6 |
mcp ls | List registered servers under the gateway that serves them help-sbx sbx mcp ls | --json, -q | both | §3.6 |
mcp rm <name> | Remove a registration help-sbx sbx mcp rm | -f | both | §3.6 |
policy init <allow-all|balanced|deny-all> | Set the global network policy once, before the first sandbox help-sbx sbx policy init | --sandbox (cloud only) | both | §3.3 |
policy allow network RESOURCES | Add an allow rule for hosts, domains, IP addresses, or CIDR prefixes, TCP by default help-sbx sbx policy allow network | --sandbox, --protocol tcp|udp | both | §3.3 |
policy deny network RESOURCES | Add a deny rule, TCP and UDP by default, which wins over allow help-sbx sbx policy deny network | --sandbox, --protocol tcp|udp | both | §3.3 |
policy check network TARGET | Ask the daemon authorizer whether a host and port is allowed, without sending anything help-sbx sbx policy check network | --sandbox, --protocol, --verbose, --json | both | §3.3 |
policy inspect <policy-or-rule> | Show one policy or one rule with its RULE_ID and removal command help-sbx sbx policy inspect | --json | both | §3.3 |
policy log [SANDBOX] | Show which hosts the proxy allowed or blocked, with rule, proxy type, and count help-sbx sbx policy log | --json, --limit, -q, --type all|network|filesystem (filesystem logs are not supported yet) | both | §3.4 |
policy ls [SANDBOX] | List policies, one overview row per policy, or the rules of one sandbox help-sbx sbx policy ls | --wide, --json, --source local|org|kit, --decision allow|deny, --type, --created-via default|added|provisioned|approval, --profile, --protocol, --include-inactive | both | §3.3 and §7.2 |
policy profile ls | List the profiles that remote governance policies provide help-sbx sbx policy profile ls | --json, -q | both | §7.2 |
policy reset | Delete the local policy store and stop the daemon, or delete the cloud account policy help-sbx sbx policy reset | -f | both | §3.3 |
policy rm network | Remove a rule by RULE_ID or by resource help-sbx sbx policy rm network | --id, --resource, --sandbox, -f | both | §3.3 |
reset | Return sbx to a freshly installed state, including secrets and the sign-in help-sbx sbx reset | -f, --preserve-secrets | local only | §2.4 |
secret set [SERVICE] | Store a service secret for one of 13 services, a dynamic source, or a registry credential help-sbx sbx secret set | -t, stdin, --ref, --command, --refresh (default 55m), --no-verify, --show-error, --oauth (openai only locally), --sandbox, -f, --registry, --username, --password-stdin, --registry-auth-endpoint, --all-sandboxes | both | §3.5 |
secret import [SERVICE] | Import secrets found in host environment variables, with a last four character preview help-sbx sbx secret import | --all, --force, --dry-run | both | §3.5 |
secret ls | List stored secrets across global and sandbox scopes help-sbx sbx secret ls | -g, --sandbox, --service, --json, -q | both | §3.5 |
secret rm [SERVICE] | Remove a secret, a registry credential, or every secret help-sbx sbx secret rm | --sandbox, --all, --all-sandboxes, --registry, -f. The examples also show --placeholder, which the flag list omits | both | §3.5 |
secret set-custom | Store a secret for a service sbx does not know, behind a placeholder the proxy swaps help-sbx sbx secret set-custom | --host (repeatable), --env, --value, -t, --command, --ref, --placeholder ({rand} suffix), --refresh, --sandbox. Cloud only: --header, --format, --name | both, experimental | §3.5 |
settings get <key> | Print one evaluated value help-sbx sbx settings get | --json | local only | §2.5 |
settings list | List every setting with value, type, source, restart flag, and description help-sbx sbx settings list | --json, --no-trunc | local only | §2.5 and R.3 |
settings set <key> <value>, settings unset <key> | Write or remove a user override help-sbx sbx settings set | none | local only | §2.5 |
template save SANDBOX TAG | Save a snapshot into the sandbox runtime image store help-sbx sbx template save | -o, --output (also export a tar), --capture-mode disk|all (cloud), -d, --description (cloud) | both | §2.3 |
template load FILE [NAME] | Load a tar into the image store, or upload it as a cloud template help-sbx sbx template load | --cpus and --memory-mib (required with --cloud, must name a billable shape), --capture-mode, --description | both | §2.3 |
template ls, template rm TAG|ID|NAME | List or remove template images help-sbx sbx template ls | --json, -q, -f | both | §2.3 |
template inspect NAME|ID | Show the full metadata of one cloud template help-sbx sbx template inspect | --json | cloud only in v1 | §2.3 |
tui | Open the interactive dashboard help-sbx sbx tui | none | both | §1.1 |
volume create|inspect|ls|rm | Manage persistent cloud volumes, saved as a snapshot when a sandbox exits help-sbx sbx volume | --json, -q, -f | cloud only | §7.1 |
sbx with no command opens interactive mode, and sbx tui opens the dashboard by name (conflict C44) help-sbx sbx.
Experimental commands
Each of these prints the line EXPERIMENTAL: this command may change or be removed in future releases. at the top of its help help-sbx sbx kit.
| Command | Purpose | Flags that matter | Scope | Section |
|---|---|---|---|---|
env create [PATH...] | Provision the secrets and bindings of an sbxenv.yaml and create the sandbox help-sbx sbx env create | -y, --auto-approve, --clone, --env-arg, --env-args-file, --kit-arg, --kit-args-file, --name, --skip-host-commands | both, experimental | §4.4 |
env run [PATH...] | Create when needed, then attach to the environment sandbox help-sbx sbx env run | the flags of env create, plus -d, --detached | both, experimental | §4.4 |
env plan [PATH...] | Print what applying the file would change, and change nothing help-sbx sbx env plan | --clone, --env-arg, --kit-arg, --name, --skip-host-commands | both, experimental | §4.4 |
env exec [PATH...] -- COMMAND | Run a command in the environment sandbox help-sbx sbx env exec | the flags of exec, plus --env-arg, --env-args-file, --name | both, experimental | §4.4 |
env rm [PATH...] | Remove the environment sandbox and the secrets it provisioned help-sbx sbx env rm | -f, --prune-bindings, --env-arg, --name, --skip-host-commands | both, experimental | §4.4 |
kit add SANDBOX REFERENCE | Recreate a sandbox with a mixin appended to its kit list help-sbx sbx kit add | --kit-arg, --kit-args-file | both, experimental | §4.3 |
kit builder status|rm | Show or remove the sbx-kit-builder sandbox and its build cache help-sbx sbx kit builder | -f (rm) | local, experimental | §4.3 |
kit builder history export|inspect|logs|ls|rm|trace | Run the matching docker buildx history command inside the builder sandbox help-sbx sbx kit builder history | pass-through | local, experimental | §4.3 |
kit inspect REFERENCE | Load and display a kit from a directory, ZIP, OCI reference, or git repository help-sbx sbx kit inspect | --json, --kit-arg, --kit-args-file | both, experimental | §4.2 |
kit pack DIRECTORY | Validate a spec.yaml directory and package it as a ZIP help-sbx sbx kit pack | -o, --output (default <name>.zip) | both, experimental | §4.3 |
kit provenance REFERENCE | Print the SLSA provenance attached to an OCI kit, UNSIGNED or VERIFIED help-sbx sbx kit provenance | --key, --certificate-identity, --certificate-identity-regexp, --certificate-oidc-issuer, --certificate-oidc-issuer-regexp, --insecure-ignore-tlog, --json | both, experimental | §4.3 |
kit pull REFERENCE | Pull a v1 ZIP or v2 tar.gz kit artifact from an HTTPS registry help-sbx sbx kit pull | -o, --output | both, experimental | §4.3 |
kit push DIRECTORY REFERENCE | Package and push a spec.yaml kit, and attach SLSA provenance as an OCI referrer help-sbx sbx kit push | --sign, --key, --identity-token, --identity-token-file, --tlog-upload (default true) | both, experimental | §4.3 |
kit sign REFERENCE | Sign a kit with a cosign-compatible Sigstore signature, keyless by default help-sbx sbx kit sign | --key, --identity-token, --identity-token-file, --tlog-upload | both, experimental | §4.3 |
kit validate REFERENCE | Check that a directory, ZIP, or git reference is a valid kit help-sbx sbx kit validate | --json, --kit-arg, --kit-args-file | both, experimental | §4.3 |
kit verify REFERENCE | Verify a signature on a directory, a git reference, or an OCI kit help-sbx sbx kit verify | --key, the four --certificate-* flags, --insecure-ignore-tlog, --json | both, experimental | §4.3 |
setup | Detect agent secrets in the host environment and import the accepted ones help-sbx sbx setup | none | local, experimental | §3.5 |
setup ssh, setup ssh remove | Write or remove the SSH client config that makes ssh <name>.sbx work help-sbx sbx setup ssh | --alias (default *.sbx) | local, experimental | §2.4 |
skills add <repository> | Install skills from a git repository or owner/repo shorthand into the shared store help-sbx sbx skills add | -s, --skill, -f | local, experimental | §4.5 |
skills import | Import skills from six agent directories on the host help-sbx sbx skills import | --dry-run, -f | local, experimental | §4.5 |
skills ls, skills rm <skill>..., skills update [skill]... | List, remove, or refresh installed skills help-sbx sbx skills ls | --json, -q, -f | local, experimental | §4.5 |
sbx kit ls is not a command, and the probe answered unknown command (conflict C46). The six kit builder history pages in the vendored tree hold an error instead of help text. The daemon could not bind its socket under the capture sandbox help-sbx sbx kit builder history ls.
Other commands
| Command | Purpose | Flags that matter | Scope | Section |
|---|---|---|---|---|
completion | Generate the autocompletion script for a shell help-sbx sbx | its help page is not in the vendored tree | both | none |
help | Help about any command help-sbx sbx | its help page is not in the vendored tree | both | none |
login | Sign in to Docker help-sbx sbx login | --username, --password-stdin | both | §1.1 |
logout | Stop running local sandboxes and sign out help-sbx sbx logout | -y, --yes | local only | §1.1 |
version | Print the CLI version, and with --json the server and runtime component versions help-sbx sbx version | --json | both | §2.4 |
Names that appear in text but not in the tree
The help of secret set names sbx mount, and the docs name sbx ssh proxy, sbx policy approval, --model, --provider, and --usb (conflict C16). None of them has a page in the vendored tree. The manual documents only what --help prints, and §2.4 names the hidden commands once. Every sbx ... --help call also tried to open a TLS connection to login.docker.com:443 during the capture. The capture sandbox refused the connection, and the printed text did not change (research/sources/help-sbx-cloud.md).
Sources:research/sources/help-sbx.md (every sbx help page, 115 pages, sbx v0.47.0 0411f50ee4700fe7bd37e6e7e3aced563e850ca9); research/sources/help-sbx-cloud.md (sbx --cloud --help and the network note); research/sources/probes/sbx-kit-ls.txt, sbx-version.txt; research/conflicts-register.md rows C9, C16, C44, C46, C47; manual.json (section files and ids)
docker-agent commands
Every command of docker-agent 1.149.0 in the four groups of its root help, every global flag, and every serve subcommand with its default listen address.
The root help of docker-agent 1.149.0 lists 19 commands in four groups: Core, Diagnose, Advanced, and Additional help-agent docker-agent. With subcommands, the vendored tree holds 47 pages. The binary is the Homebrew build. The Desktop plugin docker agent runs the same commands with a space in the name (conflict C48). The docs page features/cli is the only other reference, because /reference/cli/docker/agent/ does not exist (conflict C68).
Global flags
Every command accepts these flags help-agent docker-agent.
| Flag | Default | Meaning |
|---|---|---|
--cache-dir | ~/Library/Caches/cagent on macOS | Override the cache directory |
--config-dir | ~/.config/cagent | Override the config directory. The docs add that DOCKER_AGENT_CONFIG_DIR and the legacy CAGENT_CONFIG_DIR set the same path docs-agent features/cli |
--data-dir | ~/.cagent, or DOCKER_AGENT_DATA_DIR | Override the data directory, which holds session.db, worktrees, and plans docs-agent features/cli |
-d, --debug | off | Enable debug logging |
--log-file | ~/.cagent/cagent.debug.log | Path of the debug log, used only with --debug |
-o, --otel | off | Enable OpenTelemetry tracing |
-h, --help | none | Print the help of the command |
Core commands
| Command | Purpose | Flags that matter | Needs | Section |
|---|---|---|---|---|
getting-started | Learn docker agent with a hands-on interactive tour help-agent docker-agent | none. The tree has no page for it (conflict C75). The docs describe a two minute skippable tour, also reachable as docker agent tour docs-agent features/cli | an interactive terminal | §5.1 |
run [<agent-file>|<registry-ref>] [message]... | Run an agent from a file, a registry reference, an alias, or the built-in default help-agent docker-agent run | see the table below | a model with a credential, Docker Model Runner, or --fake | §5.1 |
setup | Set up a model through one of four paths: a built-in provider, Docker Model Runner, a custom OpenAI-compatible endpoint, or the Claude Code harness help-agent docker-agent setup | none | a terminal. The provider path writes ~/.config/cagent/.env | §5.1 |
share push <agent-file> <registry-ref> | Push an agent file to an OCI registry, in clear, with an optional signature or MAC in the annotations help-agent docker-agent share push | --key (PEM, OpenSSH, or a 16 byte secret, or DOCKER_AGENT_ENCRYPT_KEY), --encrypt | registry access | §6.4 |
share pull <registry-ref> | Pull an agent file and, with a key, verify or decrypt it help-agent docker-agent share pull | --key, --force | registry access | §6.4 |
The flags of run
The run page lists 54 flags help-agent docker-agent run. This table groups the ones the manual uses, with their defaults.
| Group | Flags |
|---|---|
| Agent and input | -a, --agent (the team's first agent), --agent-picker [refs], --prompt-file (repeatable), --attach <image>, - reads the message from stdin |
| Output | --exec (no TUI), --last (only the final answer, needs --exec), --json (NDJSON events), --hide-tool-calls, --hide-tool-results, --on-event <type>=<cmd> |
| Approval | --safety strict|balanced|restricted|autonomous, --yolo (same as --safety autonomous, conflict C59) |
| Model | --model [agent=]provider/model (repeatable), --models-gateway, --dry-run |
| Session | --session <id or -1>, -s, --session-db (default <data-dir>/session.db), --session-read-only, --working-dir |
| Record and replay | --record [path] (a cassette plus a TUI e2e test), --fake <path>, --fake-stream [15] ms between chunks |
| Sandbox | --sandbox, --sbx (default true, and --sbx=false forces the removed docker sandbox plugin, conflict C60), --template (default docker/docker-agent-sbx-templates:latest), --sandbox-kit, --kit (repeatable), --kit-arg, --no-kit, --cloud (implies --sandbox), --sandbox-ttl (default 1h0m0s) |
| Hooks | --hook-pre-tool-use, --hook-post-tool-use, --hook-session-start, --hook-session-end, --hook-on-user-input, --hook-stop, all repeatable |
| Config | --flavor (repeatable, applied in order), --env-from-file, --code-mode-tools, --mcp-oauth-redirect-uri, --remote <addr> |
| Worktree | -w, --worktree [name] (default auto), --worktree-base <ref>, --worktree-pr <number or URL> (needs the GitHub CLI) |
| TUI | --lean, --sidebar (default true), --theme, --app-name, --disable-commands |
Diagnose commands
| Command | Purpose | Flags that matter | Needs | Section |
|---|---|---|---|---|
doctor [agent-file] | Report which providers have credentials, whether Docker Model Runner answers, and which model auto picks, and exit non-zero on an issue help-agent docker-agent doctor | --json, --env-from-file, --models-gateway | nothing. It runs docker model status --json to check the runner | §5.1 |
models [list|ls] | List the models --model accepts, from the gateway /v1/models first, then from the providers with credentials help-agent docker-agent models | -a, --all, -p, --provider, --format table|json, --models-gateway | nothing | §5.1 |
toolsets | List the 27 built-in toolset types help-agent docker-agent toolsets | --format table|json | nothing | §5.3 and R.5 |
Advanced commands
| Command | Purpose | Flags that matter | Needs | Section |
|---|---|---|---|---|
alias add <alias-name> <agent-path> | Save a name for an agent file or registry reference with run options help-agent docker-agent alias add | --yolo, --safety (wins over --yolo), --model, --hide-tool-results, --sandbox | nothing | §5.1 |
alias list|ls, alias remove|rm <alias-name> | List or remove aliases help-agent docker-agent alias list | --json | nothing | §5.1 |
board | Open a Kanban TUI where each card runs an agent in a tmux session on its own git worktree help-agent docker-agent board | none | tmux, git, and projects in ~/.config/cagent/config.yaml | none |
debug auth | Print the Docker token in use and where it came from help-agent docker-agent debug auth | --json | nothing | R.6 |
debug config <agent-file> [flavor...] | Print the canonical form of an agent file with the flavors applied help-agent docker-agent debug config | the debug group flags | nothing | §5.2 |
debug oauth list|login|remove | List stored OAuth tokens, log in to a remote MCP server, or remove a token help-agent docker-agent debug oauth | --json (list) | a browser for login | §5.3 |
debug skills <agent-file> | Show the skills an agent discovers help-agent docker-agent debug skills | --json | nothing | §4.5 |
debug title <agent-file> <question> | Generate a session title from a question help-agent docker-agent debug title | --model | a model | §5.6 |
debug tool <agent-file> <tool-name> [parameters-json] | Call one tool directly, with real side effects and no model turn help-agent docker-agent debug tool | -a, --agent, --json, --no-hook | whatever the tool needs | §5.3 |
debug toolsets <agent-file> | List the tools and parameter schemas of an agent help-agent docker-agent debug toolsets | --json | nothing | §5.3 |
eval <agent-file> [<eval-dir>] | Replay saved sessions in containers and score them help-agent docker-agent eval | -c, --concurrency (default 10), --judge-model (default openai/gpt-5.6-terra), --judge-type llm|evaluator (default llm), --agent-image (default the pinned docker/docker-agent:<version>), --base-image, --container-runtime (default docker), -e, --env, --keep-containers, --only, --output (default <eval-dir>/results), --repeat (default 1), --baseline, --regression-tolerance (conflict C65) | a container runtime and a judge model | §5.6 |
new [description] | Write a new agent file from a description help-agent docker-agent new | --model (anthropic, openai, google, dmr, or a custom provider, conflict C66), --max-iterations (default 20 for DMR, unlimited elsewhere) | a model | §5.1 |
plans create|delete|export|get|list|status|update | Edit the shared plans of the plan toolset from the host, with an optimistic lock help-agent docker-agent plans | --file (- is stdin), --title, --author, --status, --expected-version (exit code 3 on conflict), --force, --output, --json | nothing | sub_agents, transfer_task, and background_agents |
sandbox allow <host>... | Add hosts to the persistent allowlist of every later --sandbox run help-agent docker-agent sandbox allow | none | nothing | §1.2 |
sandbox deny|remove|rm <host>, sandbox list|ls | Remove a host, or list the allowlist help-agent docker-agent sandbox | none | nothing | §1.2 |
serve a2a|acp|api|chat|mcp | Start an agent as a server help-agent docker-agent serve | see the table below | a model | §6.1 to §6.3 |
sessions diff <session-a> <session-b> | Report the first tool call where two recorded sessions differ help-agent docker-agent sessions diff | --json, --fail-on-divergence, -s, --session-db | nothing | §5.6 |
The debug group shares one flag set: --code-mode-tools, --env-from-file, --flavor, the six --hook-* flags, --mcp-oauth-redirect-uri, --models-gateway, and --working-dir help-agent docker-agent debug. The same flags appear on eval, new, and every serve subcommand.
The serve subcommands
| Subcommand | Transport | Default listen address | Session database | Auth and safety flags | Section |
|---|---|---|---|---|---|
serve a2a <agent-file> | HTTP, the Agent-to-Agent protocol help-agent docker-agent serve a2a | 127.0.0.1:8082 | <data-dir>/session.db | --auth-token, --insecure-no-auth, --cors-origin, --safety, -a (the team's first agent, conflict C64) | §6.3 |
serve acp <agent-file> | stdio, the Agent Client Protocol help-agent docker-agent serve acp | none | <data-dir>/session.db | none | §6.2 |
serve api <agent-file>|<agents-dir> | HTTP, sessions and SSE help-agent docker-agent serve api | 127.0.0.1:8080 | session.db in the current directory (conflict C63) | --auth-token (empty disables auth), --max-request-size (default 1048576), --session-workingdir-root, --pull-interval (0 disables), --fake, --record | §6.1 |
serve chat <agent-file> | HTTP, /v1/chat/completions and /v1/models help-agent docker-agent serve chat | 127.0.0.1:8083 | no flag | --api-key, --api-key-env, --insecure-no-auth, --cors-origin, --safety, --conversations-max, --conversation-ttl (default 30m0s), --request-timeout (default 5m0s), --max-idle-runtimes (default 4), --max-request-size, -a (all agents if not set) | §6.1 |
serve mcp <agent-file> | stdio by default, streaming HTTP with --http help-agent docker-agent serve mcp | 127.0.0.1:8081 with --http | no flag | --auth-token, --insecure-no-auth, and --safety only with --http, --attach [latest] to expose a running TUI, --tool-name, --mcp-keepalive (stdio only), -a (all agents if not set) | §6.2 |
The help text names the four safety modes on serve a2a, serve chat, and serve mcp --http, and states no default for them. The default inside each server is therefore not documented in the help.
Additional commands
| Command | Purpose | Flags that matter | Needs | Section |
|---|---|---|---|---|
completion | Generate the autocompletion script for a shell help-agent docker-agent | its page is not in the vendored tree | nothing | none |
help | Help about any command help-agent docker-agent | its page is not in the vendored tree | nothing | none |
version | Print the version and the commit hash help-agent docker-agent version | none | nothing | §1.1 |
On the capture machine docker-agent version printed v1.149.0 and Commit: Homebrew (research/sources/probes/docker-agent-version.txt).
Sources:research/sources/help-docker-agent.md (every docker-agent help page, docker-agent v1.149.0, Commit: Homebrew); research/sources/docs-docker-agent.md (page features/cli for getting-started, --config-dir, and --data-dir); research/sources/probes/docker-agent-version.txt, docker-agent-toolsets.txt; research/conflicts-register.md rows C48, C59, C60, C63, C64, C65, C66, C68, C75; manual.json (section files and ids)
Settings keys, environment variables, and paths
Every key that sbx settings list printed on 2026-10-08, every environment variable both tools read, and every path they use on macOS, each with its source.
sbx settings list printed 31 keys on the capture machine, every one with SOURCE default (research/sources/probes/sbx-settings-list.txt). A value comes, in order, from the environment variable of the key, then a user override written with sbx settings set, then the built-in default docs-sbx Settings. RESTART yes means that daemon-side consumers that already exist need sbx daemon restart, while new sandboxes and supported CLI clients use the new value at once help-sbx sbx settings list. The CLI table cuts long descriptions. The Meaning column completes them from the docs settings page where the page has an entry. Where the page has none, the column says so and gives the full text of sbx settings list --json (capture/out/02-settings.json).
The settings keys
| Key | Type | Default | Restart | Meaning |
|---|---|---|---|---|
claude.remoteControl | bool | false | no | Let the /remote-control channel of Claude Code authenticate with its own session token instead of the proxy swapping in the host credential (docs) |
clipboard.imagePaste | bool | false | no | Let sandboxed agents read host clipboard images, so a screenshot pastes with Ctrl+V (docs) |
diagnostics.autoUpload | string | empty | no | Consent for automatic diagnostics uploads after daemon errors: yes, no, or empty for no decision yet (docs) |
env.rememberHostCommands | bool | false | no | Ask about the host commands of an environment file only when they change, after a first approval (docs) |
kit.allowExtractedAgents | bool | true | no | Admit the pinned kit references of agents that moved out of sbx into kits, even outside kit.allowedSources, and exempt them from kit.requireSignature (docs) |
kit.allowLocalKits | bool | true | no | Allow kits from local directories and ZIP files (docs) |
kit.allowedSources | json | ["docker.io/"] | no | JSON array of allowed remote kit source prefixes, matched on a path segment boundary. ["*"] allows any source (docs) |
kit.ignoreTransparencyLog | bool | false | no | Verify keyless kit signatures without a Rekor entry, for kits signed with --tlog-upload=false (docs) |
kit.requireSignature | bool | false | no | Reject unsigned kits and kits signed by an untrusted signer. ZIP kits cannot carry a signature and are rejected (docs) |
kit.trustedSigners | json | identities ending in @docker.com through the Google issuer | no | JSON array of signer policies: a keyless identity with its issuer, or a public key file (docs) |
mcp.forceLocalGateway | bool | false | yes | Use the local MCP gateway when the account would otherwise use the hosted one (docs) |
model.providers | json | {} | no | JSON object of inference endpoints for sbx run --model, each with url, wire, and apiKeyEnv (docs) |
no_proxy | string | empty | yes | Shared proxy exception list for sandbox, daemon, and supported CLI traffic (docs) |
no_proxy.daemon | string | empty | yes | Exception list for daemon and CLI requests only. A non-empty value replaces the shared list for that scope (docs) |
no_proxy.sandbox | string | empty | yes | Exception list for sandbox egress only. A non-empty value replaces the shared list for sandboxes (docs) |
platform.allowExperimentalFeatures | bool | true | no | Allow experimental features. The CLI text is complete. The docs say the default is false (conflict C18) |
platform.images.registryMirror | string | empty | no | Mirror host for template and kit references that resolve to Docker Hub, without a URL scheme (docs) |
platform.images.useDHI | bool | false | no | Use the Docker Hardened Image variant dhi/sbx-templates:<tag> for the default agent templates (docs) |
proxy | string | empty | yes | Upstream proxy for sandbox, daemon, and CLI traffic: an HTTP, HTTPS, or SOCKS5 URL, a PAC source, system, or direct (docs) |
proxy.daemon | string | empty | yes | Upstream proxy for daemon requests and supported CLI requests only (docs) |
proxy.integratedAuth | bool | false | yes | NTLM or Kerberos authentication with the Windows sign-in identity. No effect on macOS or Linux (docs) |
proxy.sandbox | string | empty | yes | Upstream proxy for sandbox egress only. It overrides proxy for that scope (docs) |
sandbox.disk.dockerVolume | string | 10g | no | Size of the /var/lib/docker volume of a new sandbox, at least 512 MiB. Existing volumes keep their size (docs) |
skills.defaultMode | string | readonly | no | Mode of the shared skills store when --skills is omitted: readonly, readwrite, or off (docs) |
ssh.agentForwardingEnabled | bool | true | yes | Let clients forward an SSH agent into sandboxes. Private keys stay on the host (docs) |
ssh.agentSocketPath | string | empty | yes | Fixed host SSH agent socket path. Empty uses the socket each client supplies (docs) |
ssh.autoCreate | bool | false | yes | Create a sandbox on SSH connect when it does not exist. The docs have no entry (conflict C20) |
ssh.defaultAgent | string | shell | yes | Built-in agent used for SSH auto-created sandboxes. The docs have no entry |
ssh.defaultTemplate | string | empty | yes | Template image override for SSH auto-created sandboxes (agent default if empty). The docs have no entry |
ssh.workspaceRoot | string | empty | yes | Host directory holding SSH auto-created sandbox workspaces (empty = mount-less, container-internal). The docs have no entry |
tls.allowNegativeSerial | bool | false | yes | Accept server certificates with a negative serial number, as some TLS-inspecting proxies issue (docs) |
Value types are bool, int, float, string, and json, and an override that conflicts with an administrator constraint is rejected help-sbx sbx settings set. Most changes apply within about five seconds help-sbx sbx settings.
Keys the docs name that the CLI did not print
The docs settings page, three guides, and one release note name keys that sbx settings list did not return (conflict C19). The manual prints them here and nowhere else.
| Key | Where the docs use it | What the docs say |
|---|---|---|
feature.model | the local model guide | sbx settings set feature.model true after platform.allowExperimentalFeatures true turns on sbx run --model docs-sbx Run a local model |
feature.sandbox-gpu | the GPU guide | the same pair of commands reveals the hidden --gpu flag docs-sbx Turn on the feature |
feature.udp-egress | the network policy page | the same pair of commands allows sbx policy allow network --protocol udp docs-sbx Allow outbound UDP |
diagnostics.autoUploadErrorCooldownInDays | the settings page | integer, default 1, the number of days between automatic uploads, applied only when diagnostics.autoUpload is yes docs-sbx Settings |
feature.ssh | the v0.34.0 release note | sbx settings set feature.ssh true enables the experimental native SSH endpoint rel-sbx v0.34.0 |
Environment variables that sbx reads
These variables configure the host side. They never set a variable inside a sandbox, and the daemon reads them only when it starts docs-sbx Settings.
| Variable | Sets | Source |
|---|---|---|
DOCKER_SANDBOXES_CLIPBOARD_IMAGE_PASTE | clipboard.imagePaste | docs-sbx Settings |
DOCKER_SANDBOXES_CLAUDE_REMOTE_CONTROL | claude.remoteControl | docs-sbx Settings |
DOCKER_SANDBOXES_USE_DHI | platform.images.useDHI | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_ALLOWED_SOURCES | kit.allowedSources | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_ALLOW_LOCAL | kit.allowLocalKits | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_ALLOW_EXTRACTED_AGENTS | kit.allowExtractedAgents | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_REQUIRE_SIGNATURE | kit.requireSignature | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_TRUSTED_SIGNERS | kit.trustedSigners | docs-sbx Settings |
DOCKER_SANDBOXES_KIT_IGNORE_TLOG | kit.ignoreTransparencyLog | docs-sbx Settings |
DOCKER_SANDBOXES_PROXY | proxy.sandbox, sandbox traffic only | docs-sbx Settings |
DOCKER_SANDBOXES_NO_PROXY | no_proxy.sandbox, sandbox traffic only | docs-sbx Settings |
DOCKER_SANDBOXES_SSH_AUTO_CREATE | ssh.autoCreate | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_SSH_DEFAULT_AGENT | ssh.defaultAgent | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_SSH_DEFAULT_TEMPLATE | ssh.defaultTemplate | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_SSH_WORKSPACE_ROOT | ssh.workspaceRoot | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_MODEL_PROVIDERS | model.providers | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_ALLOW_EXPERIMENTAL_FEATURES | platform.allowExperimentalFeatures | capture/out/02-settings.json, no docs entry |
DOCKER_SANDBOXES_TLS_ALLOW_NEGATIVE_SERIAL | tls.allowNegativeSerial | docs-sbx Settings |
HTTP_PROXY, HTTPS_PROXY, NO_PROXY, and their lowercase forms | the upstream proxy when no proxy or no_proxy setting is set | docs-sbx Upstream proxies |
DOCKER_SANDBOXES_DOCKER_SIZE | the Docker data disk size of one creation, such as 30g | docs-sbx Settings |
DOCKER_SANDBOXES_ROOT_SIZE | the root filesystem size of one creation, such as 40g | docs-sbx Troubleshooting |
DOCKER_SANDBOXES_CLONED_WORKSPACE_SIZE | the size of the private clone of a --clone sandbox | docs-sbx Troubleshooting |
DOCKER_SANDBOXES_ENABLE_VIRTIOFS_CACHE | 0 turns off the virtiofs cache of the workspace mount | docs-sbx Troubleshooting |
SBX_NO_TELEMETRY | 1 turns off CLI usage analytics | docs-sbx Troubleshooting |
DOCKER_ACCESS_TOKEN | the Docker identity of sbx env commands, with a state scope per token | help-sbx sbx env |
SBX_MCP_URL | none selects the local MCP data plane instead of the hosted gateway | help-sbx sbx mcp add |
SANDBOXES_STORAGE_ROOT | not documented in the vendored sources. The research notes name it from the v0.29.0 release note, whose body the vendored release list does not keep | none |
Host lifecycle commands of an environment file receive SBX_LIFECYCLE_PHASE, SBX_ENV_FILE, SBX_ENV_FILES, SBX_ENV_DIR, SBX_SANDBOX_NAME, SBX_AGENT, and SBX_WORKSPACE docs-sbx Environment files. After a cloud creation they also receive SBX_SANDBOX_ID help-sbx sbx env.
Environment variables that docker-agent reads
The rename of v1.30.0 changed the prefix from CAGENT_ to DOCKER_AGENT_, and the old names stay accepted (conflict C53) rel-agent v1.30.0.
| Variable | Legacy name | Meaning | Source |
|---|---|---|---|
DOCKER_AGENT_DATA_DIR | none | The data directory, the same as --data-dir | help-agent docker-agent, added in v1.147.0 rel-agent v1.147.0 |
DOCKER_AGENT_CONFIG_DIR | CAGENT_CONFIG_DIR | The config directory, the same as --config-dir | docs-agent Hooks, added in v1.100.0 rel-agent v1.100.0 |
DOCKER_AGENT_MODELS_GATEWAY | CAGENT_MODELS_GATEWAY | Route model traffic through a gateway, the same as --models-gateway | docs-agent User settings |
DOCKER_AGENT_DEFAULT_MODEL | CAGENT_DEFAULT_MODEL | The model used when none is given, as provider/model | docs-agent User settings |
DOCKER_AGENT_HIDE_TELEMETRY_BANNER | CAGENT_HIDE_TELEMETRY_BANNER | 1 hides the first-run telemetry notice only | docs-agent User settings |
TELEMETRY_ENABLED | none, no prefix | false turns telemetry off (conflict C67) | docs-agent Telemetry |
DOCKER_AGENT_AUTO_UPDATE | none | 1, true, yes, or on lets a standalone release binary update itself | docs-agent User settings, added in v1.74.0 rel-agent v1.74.0 |
DOCKER_AGENT_NO_TOKEN_EXCHANGE | none | 1 stops the exchange of the docker login access token for a Docker token | docs-agent Secrets |
DOCKER_AGENT_HUB_LOGIN_URL | none | Point the token exchange at a Docker staging environment, HTTPS docker.com URLs only | docs-agent User settings |
DOCKER_AGENT_AUTO_INSTALL | none | false turns off automatic tool installation from the aqua registry | docs-agent Tools |
DOCKER_AGENT_TOOLS_DIR | none | The directory of installed tools, default ~/.cagent/tools/ | docs-agent Tools |
DOCKER_AGENT_NO_SETUP | none | 1 stops the setup wizard from being offered when no model is usable | docs-agent features/cli |
DOCKER_AGENT_BOARD_EDITOR | BOARD_EDITOR, kept for one release | The editor the board opens a worktree in, default code | docs-agent Board, renamed in v1.102.0 rel-agent v1.102.0 |
DOCKER_AGENT_PPROF_ADDR | CAGENT_PPROF_ADDR | A loopback address for a Go pprof server. The docs still print the legacy name | rel-agent v1.139.0, docs-agent features/cli |
DOCKER_AGENT_ENCRYPT_KEY | none | The key of share push --key and share pull --key | help-agent docker-agent share push |
GITHUB_TOKEN | none | Raises the GitHub API rate limit of the auto-installer | docs-agent Tools |
CAGENT_ASKPASS_SOCKET, CAGENT_ASKPASS_TOKEN | none, still prefixed CAGENT_ | The sudo bridge of the shell toolset, set only on commands that call sudo | docs-agent Shell |
CAGENT_EXP_DEBUG_LAYOUT, CAGENT_HIDE_TELEMETRY | renamed in v1.30.0 | The new names are not documented in the vendored sources | rel-agent v1.30.0 |
Provider credential variables such as ANTHROPIC_API_KEY are listed in R.6.
Paths on macOS
The capture machine runs macOS, so the table gives macOS paths first and the Linux and Windows forms where the docs state them. Replace rohitghumare with your user name.
| Path | Holds | Tool | Source |
|---|---|---|---|
~/Library/Application Support/com.docker.sandboxes/ | the sbx state directory, removed as a last resort after sbx reset | sbx | docs-sbx Removing all state |
.../com.docker.sandboxes/sandboxes/sandboxd/sandboxd.sock | the Unix socket of sandboxd | sbx | research/sources/probes/sbx-daemon-status.txt |
.../com.docker.sandboxes/sandboxes/sandboxd/daemon.log | the daemon log | sbx | research/sources/probes/sbx-daemon-status.txt |
.../com.docker.sandboxes/sandboxes/agent-skills | the shared skills store. Linux ~/.local/state/sandboxes/sandboxes/agent-skills, Windows %LOCALAPPDATA%\DockerSandboxes\sandboxes\state\agent-skills | sbx | docs-sbx Share agent skills |
~/Library/Logs/com.docker.sandboxes/sandboxes/auditkit/ | audit records written by the daemon. Linux ${XDG_STATE_HOME:-~/.local/state}/sandboxes/sandboxes/auditkit/, Windows %LOCALAPPDATA%\DockerSandboxes\sandboxes\logs\auditkit\ | sbx | docs-sbx Where records are stored |
~/.config/sbx/credentials.yaml | the credential bindings file. Windows %APPDATA%\sbx\credentials.yaml | sbx | docs-sbx Credential bindings |
| the macOS Keychain | the secret store behind sbx secret set. Windows uses the Credential Manager, Linux the Secret Service, or the file ~/.config/com.docker.sandboxes without one | sbx | docs-sbx Credentials |
~/.sbx/run/d/containerd/containerd.sock.ttrpc | the containerd ttrpc socket the daemon binds at start, named in the error the capture recorded | sbx | help-sbx sbx kit builder history ls |
~/.sbxenv.yaml | a base environment file merged under every project file | sbx | help-sbx sbx env create |
/opt/homebrew/Caskroom/sbx/0.47.0/Sbx.app/Contents/MacOS/sbx | the CLI binary of the Homebrew cask, with mkfs.erofs under Contents/libexec | sbx | research/sources/probes/sbx-diagnose.txt |
~/.local/state/sandboxes/, ~/.cache/sandboxes/, ~/.config/sandboxes/ | the three Linux state directories, under XDG_* when set. Windows uses %LOCALAPPDATA%\DockerSandboxes | sbx | docs-sbx Removing all state |
~/.config/cagent/config.yaml | user settings, aliases, global permissions and hooks, board projects | docker-agent | help-agent docker-agent board |
~/.config/cagent/.env | the env file that docker agent setup writes provider keys into | docker-agent | help-agent docker-agent setup |
~/.config/cagent/hooks.d/ | hook drop-in files, loaded in lexicographic order | docker-agent | docs-agent Hooks |
~/.cagent/ | the data directory, which holds session.db, worktrees, and plans | docker-agent | help-agent docker-agent, docs-agent features/cli |
~/.cagent/session.db | every session, as SQLite | docker-agent | docs-agent Sessions |
~/.cagent/cagent.debug.log | the debug log, with --debug | docker-agent | help-agent docker-agent |
~/.cagent/tools/bin/ | binaries the aqua auto-installer downloads | docker-agent | docs-agent Tools |
~/.cagent/plans/, ~/.cagent/session_plans/ | shared plans and per-session plans | docker-agent | docs-agent features/cli |
~/.cagent/memory/<config-name>/memory.db | the default database of the memory toolset | docker-agent | docs-agent Memory |
~/.cagent/themes/<name>.yaml | custom TUI themes | docker-agent | docs-agent User settings |
<data-dir>/runs/<pid>.json | the discovery record of a run started with --listen | docker-agent | docs-agent features/cli |
~/Library/Caches/cagent/ | the cache directory on macOS | docker-agent | help-agent docker-agent |
~/Library/Caches/cagent/sandbox-kits/<hash> | the kit that run --sandbox stages, keyed by the agent reference | docker-agent | docs-agent Sandbox |
| a private file under the cache directory | the cached Docker bearer token. Its name is not documented | docker-agent | docs-agent Secrets |
~/.codex/skills/, ~/.claude/skills/, ~/.agents/skills/, and the project .claude/skills/, .github/skills/, .agents/skills/ | the SKILL.md directories docker-agent discovers | docker-agent | docs-agent Sandbox |
The directories of docker-agent keep the name cagent after the rename (conflict C54). Nothing migrates state from the plugin-era directories ~/.docker/sandboxes/ and ~/.sandboxd/ to the sbx state directory (conflict C17).
Sources:research/sources/probes/sbx-settings-list.txt, sbx-daemon-status.txt, sbx-diagnose.txt; capture/out/02-settings.json (the full descriptions of the four ssh.* keys, and the environment variables of the ssh.*, model.providers, and platform.allowExperimentalFeatures keys); research/sources/help-sbx.md (sbx settings, sbx settings list, sbx settings set, sbx env, sbx env create, sbx mcp add, sbx kit builder history ls); research/sources/docs-sandboxes.md (pages configuration/settings, configuration/environment-files, troubleshooting, governance/audit, workflows/agent-skills, the credentials and local model and GPU and network policy guides); research/sources/sbx-releases.md (v0.34.0); research/sources/help-docker-agent.md (root, board, setup, share push); research/sources/docs-docker-agent.md (pages features/cli, configuration/user-settings, configuration/hooks, configuration/tools, configuration/sandbox, features/sessions, guides/secrets, tools/memory, tools/shell); research/sources/docker-agent-CHANGELOG.md (v1.30.0, v1.74.0, v1.100.0, v1.102.0, v1.139.0, v1.147.0); research/conflicts-register.md rows C17, C18, C19, C20, C53, C54, C67
Agent file keys
Every key of the agent file at schema version 16, grouped by block, with its type, default, requirement, and meaning from agent-schema.json at v1.149.0.
The schema is titled Docker Agent Configuration, and the pinned copy is the file at tag v1.149.0, commit bf4169cdd31229d52385410c52c3dcc59b497858 schema Docker Agent Configuration. Its version enum runs from "0" to "16", so "16" is the current form schema version. The top level and every block below set additionalProperties to false, so an unknown key fails the load. The only required top-level key is agents schema agents. Required says whether the schema lists the key under required. Default is the schema default value, and none means the schema states none.
Top-level keys
All 16 keys come from the root properties object schema properties. The docs page configuration/overview explains the reusable blocks mcps, rag, commands, skills, and toolsets.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
version | string, "0" to "16" | none | no | Configuration version |
providers | map of ProviderConfig | none | no | Reusable provider defaults: base_url, token_key, api_type |
agents | map of AgentConfig | none | yes | The agents. At least one is required |
models | map of ModelConfig | none | no | Named model configurations |
mcps | map of MCPToolset | none | no | Reusable MCP server definitions, referenced by name from a toolset |
rag | map of RAGToolset | none | no | Reusable RAG source definitions |
commands | map of Commands | none | no | Named command groups, merged with use_commands |
skills | map of SkillsConfig | none | no | Named skill groups, merged with use_skills |
toolsets | map of Toolset | none | no | Named toolset definitions, appended with use_toolsets |
metadata | Metadata | none | no | Author, license, readme, description, version, tags |
permissions | PermissionsConfig | none | no | Tool approval patterns for the whole file |
runtime | RuntimeDefaults | none | no | Execution defaults the author wants. CLI flags and user settings win |
budget | BudgetConfig | none | no | Ceilings for one run, shared by every sub-session in it |
budgets | map of BudgetConfig | none | no | Named budgets. Agents that share a name share one ceiling |
flavors | map of object or null | none | no | Named YAML patches applied with --flavor, with JSON Merge Patch rules, key+ appends, key- removes |
evaluators | map of EvaluatorConfig | none | no | Named provider-backed assessments for tool_guard and routing hooks |
Keys under agents
Every key of one agent comes from AgentConfig schema AgentConfig. The docs page configuration/agents explains them, configuration/structured-output covers structured_output, and features/harnesses covers harness.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
model | string | none | no | A model name from models or provider/model |
fallback | FallbackConfig: models (array), retries (default 2), cooldown (default 1m) | none | no | Models tried in order when the primary fails, with retry and cool-down rules |
description | string | none | no | Description of the agent |
welcome_message | string | none | no | Message shown when the agent starts |
toolsets | array of Toolset | none | no | The toolsets of the agent |
instruction | string or array of strings | none | no | The system prompt. A list is joined with blank lines |
instruction_file | string or array of strings | none | no | Files, relative to the config file, whose content is the instruction. Exclusive with instruction |
harness | HarnessConfig: type (required, claude-code, codex, pi, opencode), model, effort, agent, thinking | none | no | An external coding CLI that runs the agent instead of a model provider |
code_mode_tools | boolean | none | no | Expose one tool that calls the others through JavaScript |
sub_agents | array of strings | none | no | Agents the parent delegates to with transfer_task: local names, OCI references, or name:reference |
handoffs | array of strings | none | no | Agents that can receive the whole conversation |
force_handoff | string | none | no | The agent that always receives the conversation after a final response |
routing | AgentRouting: allowed_agents (required), default_agent | none | no | The agents a routing hook can select, and the one used when an evaluator is uncertain |
add_date | boolean | none | no | Add the date to the context |
add_environment_info | boolean | none | no | Add cwd, git, OS, and arch to the context |
readonly | boolean | none | no | Keep only tools with a read-only annotation in every toolset |
safety | strict, balanced, restricted, autonomous | none | no | Default safety mode of new sessions on this agent, below any user choice |
redact_secrets | boolean | true | no | Install the redact_secrets builtin on tool input, model input, and tool output |
max_iterations | integer | none | no | Maximum loop iterations |
budgets | array of strings | none | no | Names of top-level budgets this agent spends against |
max_consecutive_tool_calls | integer | 0, which means 5 | no | Identical tool calls in a row before the agent stops |
max_old_tool_call_tokens | integer | none, truncation off | no | Tokens kept from old tool arguments and results. -1 turns truncation off |
max_tool_result_tokens | integer | none | no | Tokens kept from each tool result, cut middle-out |
num_history_items | integer | none | no | History items to keep |
session_compaction | boolean | true | no | Compact the session at the threshold and after a context overflow |
compaction_threshold | number | 0.9 | no | Fraction of the context window that starts compaction. The model value wins |
compaction_model | string | none | no | Model that writes the summary. Highest priority of the three levels |
add_prompt_files | array of strings | none | no | Prompt files, such as AGENTS.md, added to the context |
add_prompt_files_depth | integer | 0 | no | Levels below the working directory in which the same file names are listed by path |
commands | object or array | none | no | Named prompts for slash commands |
structured_output | object: name, description, schema, strict | none | no | A JSON schema that constrains the response, native on OpenAI and Gemini |
add_description_parameter | boolean | none | no | Add a description parameter to every tool call |
hooks | HooksConfig | none | no | Lifecycle hooks of this agent |
cache | CacheConfig: enabled (default false), case_sensitive (false), trim_spaces (false), path | none | no | Replay the previous answer to the same question |
skills | boolean or array | none | no | true loads every discovered skill. A list mixes sources, names, and inline skills |
use_commands | array of strings | none | no | Top-level command groups to merge in |
use_skills | array of strings | none | no | Top-level skill groups to merge in |
use_toolsets | array of strings | none | no | Top-level toolsets to append after the inline ones |
An inline skill under skills has name, description, and instructions as required keys, plus context: fork, model, allowed_tools, and toolsets schema InlineSkill. A command under commands is a string, or an object with description, instruction, agent, and url schema CommandConfig.
Keys under models
Every key comes from ModelConfig schema ModelConfig. The docs page configuration/models explains them, and configuration/routing covers routing.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
provider | string | none | no in the schema, yes in the docs unless first_available is set | The provider id, such as openai, anthropic, dmr |
model | string | none | no in the schema, yes in the docs | The model name |
description | string | none | no | A human summary, not sent to the model |
temperature | number | none | no | Sampling temperature |
max_tokens | integer | none | no | Maximum output tokens per response, not the context window |
top_p | number | none | no | Top-p sampling |
frequency_penalty | number | none | no | Frequency penalty |
presence_penalty | number | none | no | Presence penalty |
base_url | string | none | no | The API base URL, with ${env.VAR} substitution |
parallel_tool_calls | boolean | none | no | Allow parallel tool calls |
token_key | string | none | no | Environment variable that holds the token |
bypass_models_gateway | boolean | none | no | Connect to the provider directly even when a models gateway is set |
provider_opts | object | none | no | Provider options. For dmr: runtime_flags, context_size, keep_alive, and more in R.6 |
track_usage | boolean | none | no | Track usage |
thinking_budget | string or integer | none | no | Reasoning effort or token budget, in the forms R.6 lists |
task_budget | integer or object | none | no | Total token budget of a task, sent to Anthropic as output_config.task_budget |
routing | array of RoutingRule: model, examples (both required) | none | no | Rules that pick a model from example phrases. This model becomes the router |
auth | AuthConfig: type (required, workload_identity_federation), workload_identity_federation | none | no | A non-API-key scheme that wins over the provider path |
first_available | array of strings | none | no | Candidates in priority order. The first with credentials is used. Exclusive with the other keys |
title_model | string | none | no | Model that writes session titles |
compaction_model | string | none | no | Model that writes compaction summaries |
compaction_threshold | number | 0.9 | no | Compaction threshold for agents on this model |
capabilities | CapabilitiesConfig: image, pdf, audio, video | none | no | Attachment capabilities, when the models.dev catalogue is wrong or silent |
output_capabilities | OutputCapabilitiesConfig: image | none | no | Whether the model can generate images |
cost | CostConfig: input, output, cache_read, cache_write, USD per million tokens | none | no | Prices that override the catalogue |
Keys under providers
Every key comes from ProviderConfig schema ProviderConfig. The docs page providers/custom explains them.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
provider | string | openai when unset | no | The underlying type: openai, anthropic, google, amazon-bedrock, dmr, or a built-in alias |
api_type | openai_chatcompletions, openai_responses | openai_chatcompletions in the schema, model-dependent in the docs | no | The API schema of an OpenAI-compatible provider |
base_url | string | none | no | The endpoint, required for OpenAI-compatible providers, with ${env.VAR} substitution |
token_key | string | none | no | Environment variable that holds the token |
unload_api | string | none | no | Path or URL of the model unload endpoint, used by the unload builtin |
temperature | number | none | no | Default temperature |
max_tokens | integer | none | no | Default output tokens |
top_p | number | none | no | Default top-p |
frequency_penalty | number | none | no | Default frequency penalty |
presence_penalty | number | none | no | Default presence penalty |
parallel_tool_calls | boolean | none | no | Default for parallel tool calls |
provider_opts | object | none | no | Provider options passed to the client |
track_usage | boolean | none | no | Default usage tracking |
thinking_budget | integer or string | none | no | Default reasoning budget |
task_budget | integer or object | none | no | Default task budget |
auth | AuthConfig | none | no | A non-API-key scheme |
compaction_model | string | none | no | Default compaction model, lowest of the three levels |
The federation block under auth requires federation_rule_id (prefix fdrl_), organization_id, and identity_token, and accepts service_account_id schema FederationAuthConfig. The token source has file, env, command, url, headers, and response_field schema IdentityTokenSourceConfig.
Keys under toolsets
Every key of one toolset entry comes from Toolset schema Toolset. The docs page configuration/tools explains the shared keys, and R.5 says which type uses which.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
type | one of 27 names | none | no in the schema | The toolset type |
instruction | string | none | no | Replaces the built-in instructions, or extends them with {ORIGINAL_INSTRUCTIONS} |
toon | string | none | no | Comma-separated regular expressions of tools whose JSON output is re-encoded as TOON |
readonly | boolean | none | no | Keep only tools with a read-only annotation |
model | string | none | no | Model for the turn that processes results from this toolset |
ref | string | none | no | docker:<name> or a name from mcps |
config | any | none | no | Tool-specific configuration |
command | string | none | no | Command of a stdio MCP or LSP server |
remote | Remote: url (required), transport_type, headers, oauth | none | no | A remote MCP server |
args | array of strings | none | no | Arguments of the command |
tools | array of strings | none | no | Allow-list of tool names |
env | map of strings | none | no | Environment variables |
shared | boolean | none | no | Share the tool state across agents, for think and todo |
path | string | none | no | Storage path of memory or tasks |
shell | object | none | no | Script definitions of script: cmd, description, args, required, env, working_dir |
post_edit | array of PostEditConfig: path, cmd (both required) | none | no | Commands after an edit of filesystem or file |
api_config | ApiConfig: name, endpoint, method (required), instruction, headers, args, required, output_schema | none | no | The HTTP tool of api |
webhook_config | WebhookConfig: url (required), provider, headers, chat_id | none | no | The destination of webhook |
rag_config | RAGConfig: strategies (required), tool, docs, respect_vcs (default true), indexing_timeout, results | none | no | The sources and strategies of rag |
ignore_vcs | boolean | true | no | Exclude .git and .gitignore patterns from filesystem operations |
allow_list | array of strings | none | no | Directories the file tools can reach |
deny_list | array of strings | none | no | Directories the file tools cannot reach. Wins over allow_list |
defer | boolean or array | none | no | Load tools on demand through search_tool and add_tool |
timeout | integer | 30 when omitted | no | HTTP timeout in seconds for fetch, api, openapi |
max_output_bytes | integer | 30000 | no | Text cutoff of openapi output. 0 turns it off |
escape_html | boolean | false | no | Legacy HTML escaping in multi-URL fetch results |
allowed_domains | array of strings | none | no | Hosts fetch can reach |
blocked_domains | array of strings | none | no | Hosts fetch cannot reach. Exclusive with allowed_domains |
allow_private_ips | boolean | none | no | Permit non-public addresses for fetch, api, openapi, a2a, and remote mcp |
sudo_askpass | boolean | none | no | Prompt for a sudo password through the host UI, for shell |
recall | boolean | none | no | Expose a recall parameter on run_background_job |
url | string | none | no | URL of a2a, openapi, or open_url |
headers | map of strings | none | no | HTTP headers for openapi, a2a, and fetch |
name | string | none | no | Tool name of a2a |
file_types | array of strings | none | no | Extensions an lsp server handles |
allowed_servers | array of strings | none | no | Catalog server ids mcp_catalog offers |
blocked_servers | array of strings | none | no | Catalog server ids removed from the offer |
models | array of strings | none | no | Models model_picker can choose |
version | string | none | no | owner/repo@version for auto-install, or false to turn it off |
working_dir | string | none | no | Working directory of an mcp or lsp subprocess |
lifecycle | Lifecycle: profile (resilient, strict, best-effort), required, startup_timeout, call_timeout, restart, max_restarts, backoff | none | no | The supervisor rules of an mcp or lsp toolset |
A reusable entry under mcps has command, args, ref, remote, config, version, env, tools, instruction, name, defer, working_dir, and lifecycle schema MCPToolset. An OAuth block under remote has clientId, clientSecret, callbackPort, scopes, and callbackRedirectURL schema RemoteOAuthConfig.
Keys under permissions
The three keys come from PermissionsConfig schema PermissionsConfig, and the docs page configuration/permissions gives the pattern grammar.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
allow | array of strings | none | no | Patterns approved without confirmation, such as read_* or shell:cmd=ls* |
ask | array of strings | none | no | Patterns that always ask, even for read-only tools |
deny | array of strings | none | no | Patterns always rejected. Wins over allow |
Keys under hooks
The events come from HooksConfig schema HooksConfig, and the docs page configuration/hooks explains each one. Every event holds an array. Matcher events hold HookMatcherConfig entries with matcher, hooks (required), and preempt_yolo schema HookMatcherConfig. The other events hold hook definitions directly.
| Event | Entries | When it runs |
|---|---|---|
prompt_file_guard | definitions | Before a loaded prompt file is stored or used. Every hook must approve |
skill_content_guard | definitions | Before raw skill text is expanded. Every hook must approve |
pre_tool_use | matchers | Before a tool runs. Can allow, deny, or modify |
post_tool_use | matchers | After a tool completes, with its response |
permission_request | matchers | Before the user is asked to approve a call |
session_start | definitions | When a session begins |
user_prompt_submit | definitions | Once per user message, before the first model call |
user_steering_messages_submit | definitions | When queued mid-turn messages are appended |
user_followup_submit | definitions | When a follow-up message starts a fresh turn |
turn_start | definitions | At the start of every model call, with transient context |
turn_end | definitions | When a turn ends, for any reason |
before_llm_call | definitions | Just before each model call |
after_llm_call | definitions | After each successful model call |
session_end | definitions | When a session ends |
pre_compact | definitions | Before the transcript is compacted |
subagent_stop | definitions | When a sub-agent finishes |
on_user_input | definitions | When the agent needs user input |
stop | definitions | When the model finishes responding |
notification | definitions | When the agent sends an error or warning |
on_error | definitions | When a turn hits an error |
on_max_iterations | definitions | When max_iterations is reached |
on_agent_switch | definitions | When the active agent changes |
on_session_resume | definitions | When the user lets the run continue past max_iterations |
on_tool_approval_decision | definitions | After the approval chain decides, before the call runs or the denial is recorded |
before_compaction | definitions | Immediately before a compaction. Can veto it |
after_compaction | definitions | After a successful compaction |
tool_response_transform | matchers | Between a tool run and the record of its response. Can rewrite the output |
tool_input_transform | matchers | Before every tool call, ahead of approval. Can patch the arguments |
tool_guard | matchers | After the input transform and before approval. No safety mode bypasses it |
before_agent_run | routing definitions | Once per agent activation. Can route to another agent |
after_agent_complete | routing definitions | After an agent completes. Can route to another agent |
worktree_create | definitions | Once, after --worktree creates a worktree |
A hook definition comes from HookDefinition schema HookDefinition.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
type | command, builtin, model, evaluator | none | yes | What runs: a shell command, a named in-process function, a model, or an evaluator |
command | string | none | no | The shell command or the builtin name |
args | array of strings | none | no | Arguments for the handler |
name | string | none | no | A name for logs and events |
timeout | integer | 60 | no | Seconds before the hook is cut off |
env | map of strings | none | no | Environment for this hook only |
working_dir | string | none | no | Working directory of this hook |
on_error | warn, ignore, block | warn | no | What an error, timeout, or bad output does. pre_tool_use and tool_guard always fail closed |
strict_output | boolean | false | no | Require one JSON object and reject unknown fields |
model | string | none | no | The provider/model of a model hook |
prompt | string | none | no | The Go template a model hook renders |
schema | string | none | no | pre_tool_use_decision turns the model reply into a verdict |
system_prompt | string | none | no | A literal system message for a model hook |
evaluator | string | none | no | The top-level evaluator of an evaluator hook |
evaluator_policy | EvaluatorPolicy: decisions, min_probability, fallback (all required) | none | no | How an evaluator verdict maps to a guard decision |
routing_policy | RoutingPolicy: routes, min_probability (both required) | none | no | How an evaluator choice maps to an agent |
The type description names 17 builtins schema HookDefinition. They are add_context, add_date, add_environment_info, add_prompt_files, add_git_status, add_git_diff, add_directory_listing, add_user_info, add_recent_commits, max_iterations, redact_secrets, transform_json, limit_large_tool_results, safer_shell, http_post, snapshot, and unload.
Keys under runtime, budget, metadata, and evaluators
| Block | Key | Type | Default | Required | Meaning |
|---|---|---|---|---|---|
runtime | sandbox | boolean | none | no | Run in a sandbox by default, as --sandbox does schema RuntimeDefaults |
runtime | network_allowlist | array of strings | none | no | Hosts added to the sandbox allowlist, host or host:port |
runtime | safety | one of four modes | none | no | Default safety mode of the file, below a per-agent safety |
budget | max_cost | number | none | no | Maximum USD per run, counting only priced responses schema BudgetConfig |
budget | max_tokens | integer | none | no | Maximum input plus output tokens over the run |
budget | max_time | string | none | no | Maximum sum of turn durations, as a Go duration |
metadata | author, license, readme, description, version, tags | strings, tags an array | none | no | Descriptive fields. version is used for OCI publishing schema Metadata |
evaluators | provider, model, type, instructions | strings, type one of boolean, choice, score | none | yes | A named assessment on the typesafe or openai Decisions backend schema EvaluatorConfig |
evaluators | base_url, endpoint, token_key, bypass_models_gateway, choices, levels, timeout, cost | mixed | timeout 10s in the description | no | Endpoint, credential, and output shape of the evaluator |
The docs pages configuration/budget and configuration/flavors explain budget and flavors with examples.
Sources:research/sources/agent-schema.json at tag v1.149.0 (root properties, and the definitions AgentConfig, ModelConfig, ProviderConfig, AuthConfig, FederationAuthConfig, IdentityTokenSourceConfig, Toolset, MCPToolset, Remote, RemoteOAuthConfig, PostEditConfig, ApiConfig, WebhookConfig, RAGConfig, Lifecycle, PermissionsConfig, HooksConfig, HookMatcherConfig, HookDefinition, RuntimeDefaults, BudgetConfig, Metadata, EvaluatorConfig, EvaluatorPolicy, AgentRouting, RoutingPolicy, RoutingRule, FallbackConfig, CacheConfig, CapabilitiesConfig, OutputCapabilitiesConfig, CostConfig, HarnessConfig, InlineSkill, CommandConfig); research/sources/docs-docker-agent.md (pages configuration/overview, configuration/agents, configuration/models, providers/custom, configuration/tools, configuration/permissions, configuration/hooks, configuration/budget, configuration/flavors, configuration/routing); research/sources/README.md (the pin of the schema copy)
Toolsets
The 27 toolset types of docker-agent 1.149.0, each with the tools it exposes, its options, and what it needs outside the agent file.
docker-agent toolsets printed 27 types on 2026-10-08 (research/sources/probes/docker-agent-toolsets.txt), and the type enum of the schema holds the same 27 names schema Toolset. The Tools column comes from the docs page of each type under tools/, and not documented marks a type with no page. Options are the keys of R.4 that the type reads, with the defaults the docs state. Needs says what the toolset reaches beyond the agent file: nothing, the host shell, the network, a credential, a binary, or Docker. The shared keys instruction, tools, readonly, model, defer, and toon apply to every type and are not repeated docs-agent Tool Configuration.
The 27 types
| Type | Tools | Options | Needs |
|---|---|---|---|
a2a | one tool per remote agent, named by name or from the agent card docs-agent A2A Tool | url (required), name, headers, allow_private_ips | the network and the remote A2A server. A token in headers when the server was started with --auth-token |
api | one tool per api_config, named by api_config.name docs-agent API Tool | api_config with name, endpoint, method (GET or POST in the docs, five verbs in the schema), instruction, args, required, headers, output_schema. timeout (30), allow_private_ips | the network. Tokens in headers through ${env.VAR} |
background_agents | run_background_agent, list_background_agents, view_background_agent, stop_background_agent docs-agent Background Agents Tool | none | agents listed under sub_agents |
background_jobs | run_background_job, list_background_jobs, view_background_job, stop_background_job, wait_background_job (timeout default 60) docs-agent Background Jobs Tool | env, recall (false) | the host shell |
environment | not documented. The probe summary says it reports the OS and the resolved shell, read-only, with no arguments | none documented | nothing |
fetch | fetch with urls, format (text, markdown, html), timeout (1 to 300) docs-agent Fetch Tool | timeout (30), allowed_domains, blocked_domains, allow_private_ips (false), headers, escape_html (false) | the network. GET only. Non-public addresses are refused by default |
file | not documented. The probe summary says it reads, writes, and edits individual files | post_edit, allow_list, deny_list, which the schema names for filesystem and file | nothing |
filesystem | read_file, read_multiple_files, write_file, edit_file, list_directory, directory_tree, create_directory, remove_directory, search_files_content docs-agent Filesystem Tool | ignore_vcs (true), post_edit (path, cmd), allow_list, deny_list | nothing |
git | git_status, git_log (limit default 20, path), git_branches, git_show (ref), git_blame (path required, rev) docs-agent Git Tool | none | nothing. It uses go-git and needs no git binary |
lsp | lsp_workspace, lsp_hover, lsp_definition, lsp_references, lsp_document_symbols, lsp_workspace_symbols, lsp_diagnostics, lsp_code_actions, lsp_rename, lsp_format, lsp_call_hierarchy, lsp_type_hierarchy, lsp_implementations, lsp_signature_help, lsp_inlay_hints docs-agent LSP Tool | command (required), args, env, file_types, working_dir, version, lifecycle | a language server binary, installed from the aqua registry when absent |
mcp | the tools the server lists, filtered by tools docs-agent MCP Tool | ref (docker:<name> or an mcps name), or command, args, env, version, working_dir, or remote with url, transport_type (streamable or sse), headers, oauth. config, lifecycle, allow_private_ips | ref: docker: needs the MCP Gateway and Docker. command needs the binary. remote needs the network and an OAuth login when the server asks for one |
mcp_catalog | search_remote_mcp_servers, enable_remote_mcp_server, list_remote_mcp_servers, disable_remote_mcp_server, reset_remote_mcp_server_auth docs-agent MCP Catalog Tool | allowed_servers, blocked_servers | the network. OAuth for servers that require it. No gateway, because the catalog subset is streamable HTTP only |
memory | add_memory, get_memories, delete_memory, search_memories, update_memory docs-agent Memory Tool | path (~/.cagent/memory/<config-name>/memory.db) | a SQLite file on disk |
model_picker | change_model (model required), revert_model docs-agent Model Picker Tool | models (required) | credentials for the listed models |
open_url | one tool, open_url by default docs-agent Open URL Tool | url (required), name | a browser on the host, through open, xdg-open, or rundll32 |
openapi | one tool per operation of the document docs-agent OpenAPI Tool | url (required), headers, timeout (30), max_output_bytes (30000), allow_private_ips | the network, for the document and every call |
plan | write_plan, read_plan, list_plans, delete_plan, update_plan_from_file, export_plan_to_file, set_plan_status, get_plan_status docs-agent Plan Tool | none | the store under ~/.cagent/plans/, shared by every agent with the type |
rag | one search tool named in rag_config.tool docs-agent RAG Tool | rag_config with docs, strategies (chunked-embeddings, semantic-embeddings, bm25), results, respect_vcs (true), indexing_timeout | an embedding model, so a provider credential or Docker Model Runner, and a SQLite database per strategy |
scheduler | create_schedule (prompt, when required, name), list_schedules, cancel_schedule (id required) docs-agent Scheduler Tool | none | a running session that supports recall. Schedules are not persisted |
script | one tool per entry under shell, named by the key docs-agent Script Tool | shell.<name>.cmd, description, args, required, env, working_dir | the host shell |
session_context | list_sessions, read_session docs-agent Session Context Tool | none | the session database |
shell | shell with cmd (required), cwd (.), timeout (30) docs-agent Shell Tool | env, sudo_askpass (false), safer (deprecated and ignored) | the host shell. Every command is classified safe, destructive, or unknown before approval |
tasks | create_task, get_task, update_task, delete_task, list_tasks, next_task, add_dependency, remove_dependency docs-agent Tasks Tool | path (tasks.json) | a JSON file on disk |
think | one reasoning tool. The page names no tool docs-agent Think Tool | shared in the schema | nothing. No side effects |
todo | create_todo, create_todos, update_todos, list_todos docs-agent Todo Tool | shared (false) | nothing |
user_prompt | user_prompt with message (required), title, schema docs-agent User Prompt Tool | none | an elicitation handler, which the TUI and CLI provide and some MCP clients do not |
webhook | send_webhook with message docs-agent Webhook Tool | webhook_config with url (required), provider (generic by default, or slack, discord, ifttt, telegram, mattermost, rocketchat, googlechat, teams), headers, chat_id | the network. The URL is itself the credential on Slack and Mattermost |
Names that are not types
Two tools arrive without a toolsets entry. sub_agents injects transfer_task, and handoffs injects the tool of the same name (conflict C56). The docs built-in table lists both as types, and the schema enum and the probe exclude both.
| Name | Tool | Injected by | Needs |
|---|---|---|---|
transfer_task | transfer_task with agent, task, expected_output (all required), always auto-approved docs-agent Transfer Task Tool | sub_agents on the caller | the named sub-agent |
the tool of handoffs | one tool with agent (required) that moves the conversation to a local agent and opens no network connection docs-agent Handoff Tool | handoffs on the caller | the named agent in the same file |
session_plan | write_session_plan, read_session_plan, exit_plan_mode docs-agent Session Plan Tool | the docs list it as a built-in type. The v1.149.0 enum and the probe do not name it, so the pinned binary does not accept it as a type | a per-session file under ~/.cagent/session_plans/ |
The docs name session_plan as a type, and the schema names 27 types without it. The schema ranks above the docs in this manual, so the manual treats session_plan as not available at the pin.
Sources:research/sources/probes/docker-agent-toolsets.txt; research/sources/agent-schema.json (Toolset, MCPToolset, Remote, ApiConfig, WebhookConfig, RAGConfig, ScriptShellToolConfig, PostEditConfig, Lifecycle); research/sources/docs-docker-agent.md (page configuration/tools and the 28 pages under tools/: a2a, api, background-agents, background-jobs, fetch, filesystem, git, handoff, lsp, mcp-catalog, mcp, memory, model-picker, open-url, openapi, plan, rag, scheduler, script, session_context, session_plan, shell, tasks, think, todo, transfer-task, user-prompt, webhook); research/conflicts-register.md row C56
Providers and models
Every provider id docker-agent 1.149.0 knows, the credential it reads, the provider/model reference form, the DMR endpoint order, the auto rule, and the keys that tune a model.
docker-agent doctor printed 20 providers with their credential variables on the capture machine, every one not set (research/sources/probes/docker-agent-doctor.txt). The docs name 31 ids, and the schema description of provider lists the built-in aliases schema ProviderConfig. The table below is the union, in the order of the doctor output and then of the docs. The doctor prints canonical ids such as fireworks-ai, and the legacy ids stay accepted as aliases (conflict C58).
The provider ids
| Provider id | Legacy id | Credential | Source |
|---|---|---|---|
anthropic | none | ANTHROPIC_API_KEY, or auth.type: workload_identity_federation | doctor, docs-agent Models |
openai | none | OPENAI_API_KEY | doctor |
chatgpt | none | none. A browser sign-in through docker agent setup, shown as CHATGPT_OAUTH_TOKEN | doctor, docs-agent Model Providers |
github-copilot | none | GITHUB_TOKEN or GH_TOKEN, a PAT with the copilot scope | doctor, docs-agent Model Providers |
google | none | GOOGLE_API_KEY or GEMINI_API_KEY. Vertex AI uses GOOGLE_GENAI_USE_VERTEXAI, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION | doctor, docs-agent Models |
mistral | none | MISTRAL_API_KEY | doctor |
openrouter | none | OPENROUTER_API_KEY | doctor |
baseten | none | BASETEN_API_KEY | doctor |
ovhcloud | none | OVH_AI_ENDPOINTS_ACCESS_TOKEN | doctor |
groq | none | GROQ_API_KEY | doctor |
fireworks-ai | fireworks | FIREWORKS_API_KEY | doctor, docs-agent Models |
deepseek | none | DEEPSEEK_API_KEY | doctor |
cerebras | none | CEREBRAS_API_KEY | doctor |
togetherai | together | TOGETHER_API_KEY | doctor, docs-agent Models |
huggingface | none | HF_TOKEN | doctor |
moonshotai | moonshot | MOONSHOT_API_KEY | doctor, docs-agent Models |
vercel | none | AI_GATEWAY_API_KEY | doctor |
amazon-bedrock | none | AWS_BEARER_TOKEN_BEDROCK, or the AWS credential chain, which the doctor counts as three more variables | doctor, docs-agent Models |
opencode | opencode-zen | OPENCODE_API_KEY | doctor, docs-agent Model Providers |
opencode-go | none | OPENCODE_API_KEY | doctor |
dmr | none | none. A local Docker Model Runner | docs-agent Models |
ollama | none | none. An optional base_url | docs-agent Models |
xai | none | XAI_API_KEY | docs-agent Models |
nebius | none | NEBIUS_API_KEY | docs-agent Models |
nvidia | none | NVIDIA_API_KEY | docs-agent Models |
minimax | none | MINIMAX_API_KEY | docs-agent Models |
cloudflare-workers-ai | none | CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID | docs-agent Models |
cloudflare-ai-gateway | none | CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID, and CLOUDFLARE_GATEWAY_ID | docs-agent Models |
requesty | none | REQUESTY_API_KEY | docs-agent Models |
azure | none | AZURE_API_KEY and a base_url | docs-agent Models |
a name under providers | none | the variable in token_key | schema ProviderConfig |
The doctor lists eleven ids that the docs tables do not count as providers, and the docs list ten that the doctor did not print. A credential can also come from ~/.config/cagent/.env, from a credential_helper command in the user config, from Docker Desktop, or from a 1Password op:// reference docs-agent Secrets.
The model reference
| Form | Example | Meaning | Source |
|---|---|---|---|
provider/model | dmr/ai/qwen3, anthropic/claude-sonnet-4-5 | An inline model on a known provider | docs-agent Models |
a name from models | model: local | A named model with its own keys | docs-agent Models |
name/model | my_gateway/gpt-4o | A model on a provider defined under providers | docs-agent Provider Definitions |
a,b | anthropic/claude-sonnet-4-5,openai/gpt-5 | An alloy. The runtime alternates between the models in one conversation | docs-agent Models |
auto | model: auto | The first cloud provider with a credential, else a pulled DMR model | docs-agent Set up a model |
first_available: [...] | a list of references | The first candidate whose credentials are configured, resolved at load time | docs-agent Models |
--model [agent=]provider/model | --model root=dmr/ai/qwen3 | A CLI override, repeatable per agent | help-agent docker-agent run |
The auto rule
auto picks the first cloud provider with a configured credential and then a locally pulled Docker Model Runner model docs-agent Set up a model. On the DMR side it prefers the model named in model: when that model is already pulled. Otherwise it takes the first available non-embedding model, instead of asking to pull ai/qwen3:latest docs-agent Docker Model Runner. It never picks a provider defined under providers docs-agent Provider Definitions. DOCKER_AGENT_DEFAULT_MODEL sets the model used when none is given docs-agent User settings. On the capture machine, with no credential and the runner unreachable, the doctor still printed auto -> dmr/ai/qwen3:latest and reported one issue (research/sources/probes/docker-agent-doctor.txt).
The DMR endpoint
| Step | Endpoint | Source |
|---|---|---|
| 1 | base_url on the model or the provider, when set | docs-agent Docker Model Runner |
| 2 | the endpoint the docker model plugin reports. The doctor runs docker model status --json through the Desktop context | research/sources/probes/docker-agent-doctor.txt |
| 3 | the default http://127.0.0.1:12434/engines/llama.cpp/v1 when the plugin is not found | docs-agent Docker Model Runner, 03-stack.md note 14 |
| in a container | http://model-runner.docker.internal/engines/v1, with unload_api: /engines/_unload on the provider | docs-agent Docker Model Runner |
| unload | the base_url with the trailing /v1 replaced by _unload, called by the unload builtin on on_agent_switch | docs-agent Docker Model Runner |
The runner needs no API key, and the API is not authenticated (conflict C84). The docs state only that the default URL is used when the plugin is not found. The order above puts that default third, after base_url and the plugin, as the research file reads it.
The models gateway
--models-gateway and DOCKER_AGENT_MODELS_GATEWAY route model traffic through one address, and bypass_models_gateway: true or a custom base_url sends a model directly to its provider docs-agent Model Configuration. docker agent models asks the gateway /v1/models first and uses the providers with credentials when the gateway answers nothing usable docs-agent features/cli. The Docker gateway needs a Docker token. Desktop hands out one that lasts 15 minutes. When Desktop has none, docker-agent exchanges the docker login access token for a fresh one over HTTPS docs-agent Secrets. The exchange is cached under the cache directory, DOCKER_AGENT_NO_TOKEN_EXCHANGE=1 turns it off, and docker agent debug auth shows the token in use.
Keys that tune a model
| Key | Applies to | Values | Source |
|---|---|---|---|
temperature, top_p, frequency_penalty, presence_penalty | every provider | sampling parameters, sent per request | schema ModelConfig |
max_tokens | every provider | output tokens per response, not the context window | schema ModelConfig |
thinking_budget | OpenAI | none, minimal, low, medium, high, xhigh, max. xhigh needs gpt-5.2 or later, none and max need gpt-5.6 or later | schema ModelConfig |
thinking_budget | Anthropic | an integer from 1024 to 32768, adaptive, adaptive/<effort>, or an effort level | schema ModelConfig |
thinking_budget | Amazon Bedrock with Claude | an integer, or low, medium, high | schema ModelConfig |
thinking_budget | Gemini 2.5 | an integer, -1 dynamic, 0 off, 24576 at most | schema ModelConfig |
thinking_budget | Gemini 3 | minimal on Flash only, low, medium, high | schema ModelConfig |
thinking_budget | DMR on llama.cpp | sent as llamacpp.reasoning-budget through _configure. On vLLM as thinking_token_budget per request. Ignored on MLX and SGLang | docs-agent Docker Model Runner |
task_budget | Anthropic | an integer or {type: tokens, total: N}, sent as output_config.task_budget | schema ModelConfig |
provider_opts.context_size | DMR | the context window, sent through _configure. max_tokens is never the window | docs-agent Docker Model Runner |
provider_opts.runtime_flags, raw_runtime_flags | DMR | flags for the inference runtime, as a list or one string. Exclusive with each other | docs-agent Docker Model Runner |
provider_opts.keep_alive | DMR | a Go duration, 0 unloads at once, -1 never unloads | docs-agent Docker Model Runner |
provider_opts.mode | DMR | completion, embedding, reranking, image-generation | docs-agent Docker Model Runner |
provider_opts.speculative_draft_model, speculative_num_tokens, speculative_acceptance_rate | DMR | speculative decoding with a draft model | docs-agent Docker Model Runner |
provider_opts.gpu_memory_utilization, hf_overrides | DMR on vLLM | engine settings sent through _configure | docs-agent Docker Model Runner |
provider_opts.supports_images, supports_pdf | DMR | declare attachment types, because DMR models are not in the models.dev catalogue | schema ModelConfig |
provider_opts.http_headers | OpenAI-compatible providers | headers on every request, such as the Copilot integration id | schema ModelConfig |
fallback.models, retries (2), cooldown (1m) | every provider | models tried after a failure, with backoff | schema FallbackConfig |
capabilities, output_capabilities, cost | every provider | attachment flags, image output, and USD prices that override the catalogue | schema ModelConfig |
title_model, compaction_model, compaction_threshold (0.9) | every provider | cheaper models for titles and summaries, and the compaction point | schema ModelConfig |
providers.<name> defaults | every provider | temperature, max_tokens, thinking_budget, task_budget, and the other defaults a model inherits | docs-agent Provider Definitions |
The docs give a default reasoning effort per provider docs-agent Models. It is medium on OpenAI always-reasoning models, off on Anthropic, -1 on Gemini 2.5, and model-dependent on Gemini 3.
Sources:research/sources/probes/docker-agent-doctor.txt, docker-agent-models-list.txt; research/sources/agent-schema.json (ModelConfig, ProviderConfig, FallbackConfig); research/sources/docs-docker-agent.md (pages concepts/models, configuration/models, providers/overview, providers/custom, providers/dmr, getting-started/set-up-a-model, guides/secrets, configuration/user-settings, features/cli); research/sources/help-docker-agent.md (run, models, doctor); /Users/rohitghumare/.cache/aiefs-manuals-wip/docker-research/03-stack.md note 14; research/conflicts-register.md rows C58, C84
Sources and the conflicts register
Two help trees, one schema, one kit specification, the vendored docs, the release notes, and one capture kit, with a ruling for every conflict a section touches.
The pin is sbx 0.47.0 and docker-agent 1.149.0 (manual.json). The agent file schema is at version 16, and the Sandbox Kit Spec at milestone v3.0.0-m.8. The pin date is 2026-10-07, and the facts were verified on 2026-10-08. Every file named below lives under research/sources/, and research/sources/README.md records its origin, commit, line count, and license.
The pin
| Item | Value |
|---|---|
| sbx | v0.47.0, commit 0411f50ee4700fe7bd37e6e7e3aced563e850ca9, Homebrew cask, released 2026-10-05 |
| docker-agent | v1.149.0, Homebrew build, tag commit bf4169cdd31229d52385410c52c3dcc59b497858, released 2026-10-07 |
| agent file schema | version 16, the agent-schema.json at the v1.149.0 tag |
| Sandbox Kit Spec | v3.0.0-m.8, tag commit 129be2ff45e8f9463450eb3cf04ddcb52c2b76e5, 2026-10-02, a pre-release |
| docs.docker.com | 252 pages served 2026-10-08, source repository docker/docs at 2c8a358489b56cd24069cc9e3a3d9a7b376dcfe7 |
| capture machine | one Mac, macOS, with the probes recorded between 06:55Z and 06:58Z on 2026-10-08 |
The ranked sources
When two sources disagree, the higher one wins and the section says so. The table repeats the ranking of manual.json with the citation keys of quoteSources.
| Rank | Source | Cited as | Used for |
|---|---|---|---|
| 1 | the help text that sbx 0.47.0 and docker-agent 1.149.0 print | help-sbx, help-agent with the command | every command, flag, default, and path |
| 2 | the agent file schema at v1.149.0 | schema with the definition | every key, type, and default of the agent file |
| 3 | Sandbox Kit Spec v3 at tag v3.0.0-m.8 | kitspec with the section, kitspec-main for unreleased changes, kitcap for a capability page | the kit descriptor and each capability |
| 4 | the Docker documentation, fetched 2026-10-08 | docs-sbx, docs-agent, docs-dmr, docs-mcp, docs-sbx-api, docs-desktop with the page | rules and limits the help text does not state |
| 5 | the release notes | rel-sbx, rel-agent with the version | when a behaviour appeared, changed, or was removed |
| 6 | Docker blog posts and talks, through the research files | blog, talk with a date, declared null | history and positioning only, never a rule |
| 7 | the capture kit | the file under capture/out/ | every command output, file, and record shown |
The vendored files
| File | Origin | Commit, tag, or version |
|---|---|---|
help-sbx.md, help-sbx-cloud.md | the full sbx --help tree and sbx --cloud --help, program output | sbx v0.47.0 0411f50ee4700fe7bd37e6e7e3aced563e850ca9 |
help-docker-agent.md | the full docker-agent --help tree | docker-agent v1.149.0, Commit: Homebrew |
help-legacy-docker-sandbox.md, help-legacy-docker-agent.md | the docker sandbox and docker agent plugin trees, kept for comparison | plugin v0.12.0 f13b3c1a96a8be40b06473bb3db0c26dbfe1878c, plugin v1.32.4 bd55840ec12b55874dd9fccf88912f9b6bb3e3f3 |
agent-schema.json | docker/docker-agent agent-schema.json | tag v1.149.0 bf4169cdd31229d52385410c52c3dcc59b497858 |
docker-agent-CHANGELOG.md | docker/docker-agent CHANGELOG.md, newest entry v1.149.0, no v1.146.0 entry | tag v1.149.0 bf4169cdd31229d52385410c52c3dcc59b497858 |
docker-agent-README.md | docker/docker-agent README.md | main 7a69c316f03c635d8d1951bc4677c59b47beedc7 |
SPEC-v3-at-v3.0.0-m.8.md | docker/sandbox-kit-spec docs/spec/SPEC-v3.md at the tag | tag v3.0.0-m.8 129be2ff45e8f9463450eb3cf04ddcb52c2b76e5 |
SPEC-v3.md, kit-capabilities.md, kit-spec-extras.md, kit.schema.json | the same specification on main, 20 capability pages, the README and governance files, and the kit schema | main 4be7f4dff4d647f10c51dcdb392ce3fa18dc3e96, 2026-10-07 |
docs-sandboxes.md | 79 pages under /ai/sandboxes/ | docker/docs main 2c8a358489b56cd24069cc9e3a3d9a7b376dcfe7, served 2026-10-08 |
docs-docker-agent.md | 108 pages under /ai/docker-agent/ | the same |
docs-model-runner.md, docs-mcp.md, docs-compose-models.md, docs-sandboxes-api.md, docs-desktop-release-notes.md | 8, 10, 3, 44 pages, and the Desktop release notes to 4.94.0 | the same |
sbx-releases.md | the GitHub releases of docker/sbx-releases, stable tags v0.21.0 to v0.47.0 only, bodies not kept | API read 2026-10-08, repository main 2329d12106fee653c0890152947fdd827e00cfd0 |
mcp-gateway-README.md, model-runner-README.md, compose-for-agents-README.md | the README of each repository | main a34df45d4ec0e941a9853ad768c4f6cd818966b3, ed3e67a8205b8d068b9c65b30d3708132a231bba, bfd4fe952591495af757a1a737c7eacc78c75c15 |
probes/ | 21 read-only command runs with timestamps and exit codes | sbx v0.47.0, docker-agent v1.149.0, 2026-10-08 |
The sbx help text and release notes are proprietary program output of Docker Inc., quoted as short quotations. The repositories docker/docker-agent, docker/sandbox-kit-spec, docker/docs, and docker/model-runner are Apache-2.0. The repository docker/mcp-gateway is MIT, and docker/compose-for-agents is Apache-2.0 or MIT (research/sources/LICENSES.md).
The conflicts register
research/conflicts-register.md holds 122 conflicts, C1 to C122, in nine groups, and 30 pieces of stale advice, S1 to S30. The table below keeps every C row that a section entry of the plan cites or whose ruling names a section. That is 94 rows, each with the ruling the manual follows. The 28 rows that no section uses are cut to keep the table short, and the next table names them by group. Their substance appears in the S rows below or in the register itself. Kind is the register's own label.
| Group | Rows cut |
|---|---|
| A, sbx | C1, C2, C3, C5, C6, C8, C10, C11, C14, C22, C23, C24, C30 |
| B, docker-agent | C55, C57, C74 |
| C, Docker Model Runner | C78, C80, C81, C82, C85 |
| D, MCP gateway and Toolkit | C86, C87 |
| E, Compose | C94, C95, C97 |
| F, Offload | C100 |
| G, Cloud Sandboxes | C102 |
| Id | Kind | What disagrees | Ruling |
|---|---|---|---|
| C4 | talk vs docs | libkrun or Firecracker as the VMM, against a VMM Docker wrote | A VMM Docker wrote, on Hypervisor.framework, WHP, and KVM. libkrun stays unverified |
| C7 | removed | API keys exported in the shell, against the secret store since v0.35.0 | Secrets come from the secret store. -e KEY is a plain variable, not a secret |
| C9 | docs vs CLI | five to nine agent names in blogs, against 11 in sbx run --help | The 11 names of the help, the 8 create subcommands, and cagent as an alias |
| C12 | docs vs repo | a 512 MB default kit volume, against 20 GiB in the spec docs | Both numbers with their sources, no single default |
| C13 | docs vs CLI | v1, v2, and v3 kits, against sbx kit commands that name only v1 and v2 | The sbx kit artifact commands are v1 and v2 tooling. A v3 kit is an OCI image built with Docker tooling |
| C15 | community vs docs | no Docker Desktop needed, against a required sign-in | Both true. No Desktop, but a Docker account sign-in |
| C16 | undocumented | sbx mount, sbx ssh proxy, sbx policy approval, --model, --provider, --usb in text, absent from the tree | Only what --help prints, and one note on the hidden names |
| C17 | undocumented | plugin-era state directories, against the sbx state directory | No state migration exists |
| C18 | docs vs CLI | platform.allowExperimentalFeatures default false in the docs, true in the CLI | The value and SOURCE column from the capture |
| C19 | docs vs CLI | feature.* keys in the docs, absent from sbx settings list | Only keys the CLI returns, plus a note on the documented ones |
| C20 | undocumented | four ssh.* keys in the CLI, absent from the docs: ssh.autoCreate, ssh.defaultAgent, ssh.defaultTemplate, and ssh.workspaceRoot | Included, with the full descriptions of sbx settings list --json |
| C21 | docs vs CLI | 11 built-in secret services in the docs, 13 in the CLI | 13 services |
| C25 | renamed | the plugin name rule, against the sbx rule | The sbx rule: 2 to 63 characters, letters, digits, hyphens, periods, default reserved |
| C26 | undocumented | host.docker.internal:3128 or gateway.docker.internal:3128, against docs with no address | The address and variables the capture shows |
| C27 | blog stale | UDP and ICMP blocked for good, against UDP behind feature.udp-egress | UDP rules exist behind the experimental flag. ICMP is blocked. DNS is policy-controlled |
| C28 | experimental | filesystem policies in ls and audit, against policy log that does not support them | Filesystem rules list but do not log in v0.47.0 |
| C29 | talk vs docs | near-instant start, against seconds | The manual's own timing on one machine, no vendor number |
| C31 | removed | docker sandbox exec -d, against sbx exec -d not supported | Said so in 2.2 |
| C32 | renamed | docker sandbox save into the host daemon, against the runtime image store | Said so in 2.3 and the migration table |
| C33 | undocumented | Gordon runs on the host, against sbx reset clearing Gordon sessions | The help line, nothing more |
| C34 | community vs docs | header matching rules for custom secrets, against --header and --format cloud only | The captured request on the receiver. --header and --format cloud only |
| C35 | docs vs CLI | proxy-managed sentinels, against GHO_SBX_PROXY_MANAGED and docker-placeholder-value | The captured values |
| C36 | docs vs docs | cloud needs 0.45.0, against 0.45.1 | 0.45.1 |
| C37 | undocumented | audit JSONL needs 0.39.0, against no word on subscriptions | What the directory holds after the runs |
| C38 | removed | relative --command helpers, against a fresh temporary directory since v0.46.0 | Absolute helper paths outside writable mounts |
| C39 | removed | -p binds both loopbacks, against tcp4 since v0.42.0 | The tcp4 rule |
| C40 | removed | kits from any registry, against kit.allowedSources since v0.34.0 | The setting is changed for a local registry and the error without it is shown |
| C41 | removed | global rules only, against per-sandbox policies since v0.29.0 | The two scopes global and local |
| C42 | docs vs CLI | ssh <name>.sbx through a managed block, against sbx setup ssh details | The help text facts |
| C43 | docs vs CLI | templates and kits as different things, against a kit as the positional agent | Three words defined once in 2.1 |
| C44 | undocumented | the TUI as the dashboard, against sbx with no command | Both entry points |
| C45 | community vs docs | 16K guest pages, against the sbx diagnose line | The diagnose line |
| C46 | docs vs CLI | sbx kit ls, against unknown command | The ten sbx kit subcommands the help lists |
| C47 | docs vs CLI | the same verbs with --cloud, against hidden and cloud-only verbs | Each verb marked local, cloud, or both |
| C48 | docs vs CLI | plugin v1.32.4, against the Homebrew v1.149.0 | v1.149.0 from Homebrew, both versions captured |
| C49 | docs vs docs | Docker Agent in Desktop 4.63, against the first release note at 4.64.0 | 4.63 per the docs, first note 4.64.0 |
| C50 | undocumented | a stated bundle per Desktop release, against releases that state none | A footnote in 7.3 |
| C51 | renamed | cagent commands and docs, against docker agent and /ai/docker-agent/ | Migration rows S5 to S9 |
| C52 | removed | cagent config, feedback, build, catalog, exec, against the v1.23.4 restructure | Migration rows S6 and S7 |
| C53 | renamed | CAGENT_* variables, against DOCKER_AGENT_* since v1.30.0 | The new names with the legacy aliases |
| C54 | renamed | a complete rename, against directories still named cagent | The paths as they are |
| C56 | docs vs repo | handoff and transfer_task as types, against 27 types without them | Not a type. Separate rows marked implicit |
| C58 | renamed | fireworks, together, moonshot, opencode-zen, against canonical ids | Canonical ids with an alias column |
| C59 | renamed | --yolo as the way, against --safety autonomous | --safety autonomous, with --yolo as its alias |
| C60 | docs vs CLI | --sandbox needs Desktop, against --sbx default true | --sbx=false has no working target on Desktop 4.80.0 or later |
| C61 | docs vs docs | docker/docker-agent-sbx-templates:latest, against docker/sandbox-templates:docker-agent | Two images for two launch paths |
| C62 | undocumented | --yolo inside --sandbox, against a cut default | The mode stays unknown in this edition. The run stopped at "attempt to write a readonly database (1032)" before it printed one (27-sandbox-run.txt) |
| C63 | docs vs CLI | <data-dir>/session.db, against session.db in the current directory for serve api | Both defaults |
| C64 | renamed | serve a2a -a defaults to root, against the team's first agent | The v1.149.0 text |
| C65 | docs vs CLI | eval -c as the CPU count and an Anthropic judge, against 10 and openai/gpt-5.6-terra | 10 and openai/gpt-5.6-terra |
| C66 | docs vs CLI | new --model with 30 providers, against four named in the help | new auto-selects among those, run --model takes any provider |
| C67 | docs vs repo | telemetry off through DOCKER_AGENT_*, against TELEMETRY_ENABLED=false | Both printed |
| C68 | undocumented | a CLI reference at /reference/cli/docker/agent/, against a 404 | The features page and the help tree |
| C69 | undocumented | N hook built-ins, against eight confirmed names | Only the confirmed built-ins |
| C70 | docs vs docs | serve a2a as full A2A, against listed limitations | The captured card and each limitation the capture shows |
| C71 | undocumented | the card at /.well-known/agent-card.json, against no stated path | The path that answered and the protocolVersion field |
| C72 | undocumented | Gordon as separate, against docker ai calling Docker Agent | One sentence in 1.1 |
| C73 | undocumented | ref: docker:<name> through the gateway, against no word on the Toolkit | Both outcomes |
| C75 | undocumented | docker agent with no arguments runs run, against getting-started listed first | What its help says |
| C76 | docs vs docs | served agents forward budgets, against no forwarding | Budgets do not apply to served agents |
| C77 | docs vs CLI | docker model configure, against a missing command | Whichever exists. Docker Agent sets context_size through _configure |
| C79 | docs vs docs | /anthropic/v1/messages, against /v1/messages | The one that answered 200 |
| C83 | undocumented | context_size applied, against issue #4522 | The request body the runner received |
| C84 | docs vs docs | no key needed, against an unauthenticated API | The same fact, said in 6.5 |
| C88 | docs vs repo | an invite-only gateway, against an MIT gateway | Two things share a name, separated in 6.5 |
| C89 | docs vs repo | no default for --verify-signatures, against true | The default from the help |
| C90 | docs vs repo | registry references as supported, against partly implemented | Marked partly implemented in the Toolkit |
| C91 | docs vs docs | one gateway for everything, against a separate sandbox gateway | Three gateways named |
| C92 | undocumented | a known bundled gateway version, against 0.42.2 last stated | Both outputs |
| C93 | docs stale | the Action installs v0.22.0, against v0.44.1 | One sentence in 7.3 |
| C96 | community vs docs | Compose starts sandboxes, against no such feature | One sentence in 6.5 |
| C98 | blog stale | 300 free GPU minutes, against no public price | Offload is out of scope |
| C99 | docs vs docs | Offload as the cloud path, against sbx --cloud | Said so in 7.1 |
| C101 | docs vs docs | prices documented, against prices in a blog and template load --help | Shapes from the help, prices from the blog with its date |
| C103 | docs vs docs | pay-as-you-go on a Personal account, against Personal and Pro | Personal or Pro, the rest unverified |
| C104 | docs vs docs | docker exec and healthchecks work in the cloud, against the VM filesystem being reached instead | A documented limitation |
| C105 | talk vs docs | Warp Oz on Docker cloud sandboxes, against no confirmation | Omitted |
| C106 | blog stale | the sandbox primitive inside Kubernetes, against no other mention | Omitted |
| C107 | removed | the plugin still in Desktop, against removal in 4.80.0 | The error text the capture shows |
| C108 | removed | --mount-docker-socket, --load-local-template, --pull-template, against none in sbx | Migration rows S1 to S4 |
| C109 | renamed | network proxy flags, against policy subcommands with no bypass | Migration rows S10 and S11, bypass with no replacement found |
| C110 | removed | the legacy default allowed hosts, against three presets | The legacy list only in 7.3 |
| C111 | removed | cagent in Desktop, against removal in 4.81.0 | Migration row S8 |
| C112 | removed | keys from the daemon environment and state under ~/.docker/sandboxes/, against the sbx store | One table in 7.3 |
| C113 | docs vs CLI | index annotations that the frontend promotes (kitspec §9.3), against an index with an attestation manifest and no annotations | 4.1 prints both. The annotations sit on the platform manifest (28-index.json, 28-manifest.json), which a consumer reads when the index has none |
| C114 | docs vs CLI | a floating docker/sandbox-kit:3 that never moves for a milestone, against a resolve to 3.0.0-m.8 | 4.1 prints both (28-buildx.txt). A build that needs the same frontend every time names the exact version tag |
| C115 | docs vs CLI | docker buildx build -f kit.yaml --push as the publish command, against a schema 2 manifest with no annotations from the default docker driver | 4.1 says to build with a docker-container builder (28-manifest-docker-driver.json, 28-manifest.json). An image without the annotation is not a kit |
| C116 | docs vs CLI | a local source directory passed to sbx during development, against sbx kit inspect failing on Desktop 4.94.0 | 4.1 prints both (13-kit-v3-inspect.txt, 28-kit-inspect-source.txt). No capture ran sbx run on a directory |
| C117 | docs vs CLI | source builds in the sbx-kit-builder sandbox, against a build through the host Docker daemon and a builder "not created" | 4.1 and 4.3 print the help line and the status (13-kit-builder-status.txt). Where a successful source build runs stays unstated |
| C118 | docs vs repo | agent file config version 15 as current, against a schema enum to "16" and captured files that load | "16" from the schema (5.2, R.4). The docs page still names 15 |
| C119 | docs vs CLI | session titles made from the first message, against five sessions titled Running agent | 5.6 prints the five rows (21-session-db.txt) |
| C120 | docs vs CLI | a skill name and mode lists that A2A 1.0.1 requires, against a card with an empty skill name and two empty mode lists | 6.3 prints the card against the rules (23-agent-card.json). A client accepts the empty values |
| C121 | docs vs CLI | "A2A artifact support not yet integrated", against a task whose artifacts hold the answer | 6.3 says the capture contradicts the limitation (23-a2a.http) |
| C122 | docs vs CLI | a release note that points to docker sbx, against a notice that points to the product page and a plugin list that still shows sandbox v0.13.0 | 7.3 prints the notice (29-docker-sandbox.txt, 29-docker-plugins.txt). The plugin entry stays, and the command only prints the notice |
Stale advice
The migration section prints every row. The last column is what the register says to do now.
| Id | The advice as printed | Stopped being true | Do this instead |
|---|---|---|---|
| S1 | docker sandbox run <agent> | Desktop 4.80.0, 2026-06-29 | sbx run <agent> |
| S2 | docker sandbox run --mount-docker-socket kiro | Desktop 4.58.0, 2026-01-26 | Every sandbox has a private Docker Engine |
| S3 | --load-local-template | Desktop 4.61, 2026-02-18 | sbx template load FILE, then --pull never -t TAG |
| S4 | --pull-template missing | v0.21.0, 2026-03-31 | --pull always, missing, or never, default always |
| S5 | docker sandbox create cagent . | Desktop 4.80.0 | sbx create docker-agent ., with cagent kept as an alias |
| S6 | cagent run agent.yaml, cagent new, cagent exec | v1.23.4, 2026-02-19, and Desktop 4.81.0 | docker agent run, docker agent new, docker agent run --exec |
| S7 | cagent push, pull, acp, api, mcp, a2a | v1.23.4 | docker agent share push or pull, docker agent serve acp, api, mcp, a2a |
| S8 | cagent version, brew install cagent | v1.30.0, 2026-03-09, and Desktop 4.81.0 | docker agent version, brew install docker-agent, winget install Docker.Agent |
| S9 | cagent config, feedback, build, catalog | v1.23.4 | Removed. The catalog is the Hub namespace agentcatalog/* |
| S10 | docker sandbox network proxy S --allow-host api.example.com | Desktop 4.80.0 | sbx policy allow network api.example.com [--sandbox S] |
| S11 | docker sandbox network log --json | Desktop 4.80.0 | sbx policy log [S] --json |
| S12 | docker sandbox save S TAG into host Docker | Desktop 4.80.0 | sbx template save S TAG [-o FILE] into the runtime store |
| S13 | a proxy set by hand at host.docker.internal:3128 | sbx, where the daemon configures the proxy | Nothing. The capture prints what the sandbox sees |
| S14 | API keys in ~/.zshrc and a Desktop restart | v0.35.0, 2026-07-10 | sbx secret set SERVICE or sbx secret import |
| S15 | sbx run claude --branch | v0.31.0, 2026-05-28 | sbx run claude --clone |
| S16 | the kit v1 grammar with schemaVersion: "1" | v2 recommended 2026-09-09, v3 published 2026-09-24 | v2 spec.yaml for sbx kit artifacts, v3 kit.yaml with # syntax=docker/sandbox-kit:3 for OCI kits |
| S17 | sbx mcp catalog | v0.45.0, 2026-09-21 | sbx mcp add --url <registry or manifest URL> |
| S18 | sbx mcp enable github-official --sandbox my-project | never in a help tree | sbx mcp add, then --static-mcp or sbx mcp load |
| S19 | Docker Agent v2.x renamed the CLI | never true | The latest is v1.149.0 |
| S20 | docs.docker.com/ai/cagent/ | the v1.30.0 rename | docs.docker.com/ai/docker-agent/ |
| S21 | CAGENT_MODELS_GATEWAY, CAGENT_CONFIG_DIR, CAGENT_PPROF_ADDR | v1.30.0, still accepted | DOCKER_AGENT_MODELS_GATEWAY, DOCKER_AGENT_CONFIG_DIR, DOCKER_AGENT_PPROF_ADDR |
| S22 | docker/sandbox-templates:cagent | the rename | docker/sandbox-templates:docker-agent |
| S23 | a --command helper written as ./helper or cat token | v0.46.0, 2026-09-28 | An absolute path outside writable sandbox mounts |
| S24 | -p 3000:8080 binds IPv4 and IPv6 | v0.42.0, 2026-09-07 | Default tcp4. Write 3000:8080/tcp for both |
| S25 | kits from any registry | v0.34.0, 2026-06-26 | Add the prefix to kit.allowedSources |
| S26 | docker cp of the agent binary and agent run dev-team.yaml inside | sbx and the docker-agent template | sbx run docker-agent . or docker agent run --sandbox agent.yaml |
| S27 | one sandbox per workspace | names default to <agent>-<workdir> | Name sandboxes with --name |
| S28 | Windows 10 is supported | v0.35.0, 2026-07-10 | Windows 11 with the Windows Hypervisor Platform |
| S29 | the Toolkit gateway at host.docker.internal:8811 in five steps | v0.38.0, 2026-08-06 | sbx mcp add and sbx mcp load. The Toolkit gateway is separate |
| S30 | a paid Anthropic judge model for eval | v1.147.0, 2026-10-05 | Default openai/gpt-5.6-terra, any provider/model works |
Sources:manual.json (pin, sources, quoteSources); research/sources/README.md (origins, commits, tags, versions, and the trim of 2026-10-08); research/sources/LICENSES.md; research/conflicts-register.md (rows C1 to C122 and S1 to S30, with the plan's citations in research/plan.md used to pick the rows); research/sources/probes/ (the probe timestamps); capture/out/13-kit-builder-status.txt, 13-kit-v3-inspect.txt, 21-session-db.txt, 23-a2a.http, 23-agent-card.json, 27-sandbox-run.txt, 28-buildx.txt, 28-index.json, 28-kit-inspect-source.txt, 28-manifest.json, 28-manifest-docker-driver.json, 29-docker-plugins.txt, 29-docker-sandbox.txt (the files that rows C62 and C113 to C122 name)
Glossary and index of figures
Each term below has one meaning across the manual and names the section that defines it, and the index lists every figure by the claim it makes.
Every term is defined once, in the section the last column names, with the sources that section cites. The meanings below come from the help text, the schema, the kit specification, and the docs named in the Sources line.
Glossary
| Term | Meaning in this manual | Defined in |
|---|---|---|
| Agent | One entry under agents, with a model, an instruction, toolsets, and optional sub-agents | 5.2 |
| Agent file | The YAML or HCL file docker-agent run loads, titled Docker Agent Configuration in the schema: agents, models, providers, toolsets, and the rules between them | 5.2 |
| Alloy | A model reference of two models with a comma between them, a,b, that the runtime alternates between in one conversation | 5.2 |
| Attestation | A signed statement attached to an artifact: SLSA provenance on a kit, or the DSSE publication statement on a shared agent | 4.3 and 6.4 |
| Audit record | One JSONL line the daemon writes for each decision under the auditkit directory, rotated by time, count, and size | 7.2 |
| Background agent | A sub-agent started by run_background_agent that runs while the parent continues | sub_agents, transfer_task, and background_agents |
| Binding | A record in credentials.yaml of the credential mechanism and the domains approved for one service | 3.5 |
| Capability | A typed, versioned request a kit makes of the host, com.docker.sandbox/<name>@N, answered as granted, refused, or prompted | 4.2 |
| Cassette | The file --record writes with the model API interactions of a run, and --fake replays | 5.1 |
| Clone mode | The --clone workspace mode: the host repository is mounted read-only at /run/sandbox/source and the agent works on a private clone, whose commits return through the sandbox-<name> remote | 3.2 |
| Cloud sandbox | A sandbox run through sbx --cloud on Docker's compute, with no host workspace, a shape, and a TTL | 7.1 |
| Delegation | The transfer_task tool: the parent sends a task to a sub-agent, which runs in a sub-session and returns a result | sub_agents, transfer_task, and background_agents |
| Environment file | sbxenv.yaml, which declares the agent, kits, workspace, secrets, ports, MCP servers, and host commands of one sandbox | 4.4 |
| Environment plan | The list of everything applying an environment file would set up, printed by sbx env plan and approved before create or run acts | 4.4 |
| Eval | A saved session that docker-agent eval replays in a container and scores | 5.6 |
| Flavor | A named YAML patch under flavors, applied with --flavor before the file is parsed | 5.2 |
| Forward proxy | The host proxy that HTTP and HTTPS requests from a sandbox pass through. It enforces policy and injects credentials | 3.4 |
| Gateway | The one MCP endpoint a sandbox sees, served on the host, through which every registered server is reached. The Toolkit gateway and the hosted gateway are separate things with the same name | 3.6 |
| Governance profile | A named profile from remote governance policies, assigned to a sandbox with --profile | 7.2 |
handoff | The tool that handoffs: injects. It moves the whole conversation to another agent in the same session, and the previous agent leaves the loop | sub_agents, transfer_task, and background_agents |
| Harness | An external coding CLI, claude-code, codex, pi, or opencode, that runs an agent instead of a model provider | 5.1 |
| Hook | A command, builtin, model, or evaluator that runs at one of the 33 lifecycle events of an agent | 5.5 |
| Kit | One OCI image whose manifest annotation vnd.docker.sandbox.kit.descriptor carries its declarations. For the sbx kit commands, a v1 or v2 artifact with a spec.yaml | 4.1 |
| Kit argument | A value for an argument a kit declares, given as --kit-arg name=value | 4.1 |
| Kit set | A kind: set descriptor that lists kits and is never published | 4.1 |
| Mixin | A kit of kind: mixin: an overlay on the filesystem of a workload, zero or more per composition | 4.1 |
| Model reference | provider/model, a name from models, auto, an alloy, or a first_available list | 5.2 |
| Models gateway | An address set with --models-gateway that model traffic routes through, with a Docker token | R.6 |
| Permission rule | An allow, ask, or deny pattern on a tool name and its arguments, evaluated deny, then allow, then ask | 5.5 |
| Policy | A set of rules that controls what sandboxes can reach. Local rules apply to all sandboxes or to one | 3.3 |
| Policy scope | global, all sandboxes, or local, one sandbox named with --sandbox | 3.3 |
| Preset | allow-all, balanced, or deny-all, chosen once with sbx policy init | 3.3 |
| Provider | A model API docker-agent knows by id, such as anthropic or dmr, or an entry under providers with its own base URL | 5.2 |
| Proxy type | The PROXY column of sbx policy log: forward, forward-bypass, transparent, network, or browser-open | 3.4 |
| Proxy-managed | A credential whose real value stays on the host while the proxy swaps its sentinel into outbound requests. Also the literal sentinel value proxy-managed | 3.5 |
| Rule | One allow or deny entry of a policy, with a RULE_ID, a resource pattern, and a protocol | 3.3 |
| Safety mode | strict, balanced, restricted, or autonomous: what the runtime does with a tool call that no permission rule matched | 5.5 |
| Sandbox | A microVM with its own kernel and a private Docker Engine, in which the agent runs as a container. It has a name, a workspace, a template, and a lifecycle | 2.1 |
sandboxd | The host daemon that owns every local sandbox, reached over a Unix socket | 2.4 |
| Secret | A value sbx secret set stores in the host keychain for one of 13 services, a custom host, or a registry. It never enters the sandbox | 3.5 |
| Sentinel | The placeholder value the agent sees in place of a secret, such as proxy-managed | 3.5 |
| Session | The record of one conversation in session.db: every message, tool call, sub-agent run, and cost | 5.6 |
| Shape | A billable cloud size, micro, small, medium, large, or xl, from 1 vCPU and 2048 MiB to 16 vCPU and 32768 MiB | 7.1 |
| Skills store | The shared directory of SKILL.md skills that sbx links into every sandbox, read-only by default | 4.5 |
| Static set | The MCP servers fixed at creation with --static-mcp, as opposed to servers attached later with sbx mcp load | 3.6 |
| Sub-agent | An agent listed in sub_agents that the parent delegates to with transfer_task | sub_agents, transfer_task, and background_agents |
| Template | A container image a sandbox starts from: the default docker/sandbox-templates:<agent> image, a saved snapshot, or a loaded tar | 2.3 |
| Toolset | One entry under toolsets with a type from the 27 built-in types, which gives the agent a set of tools | 5.3 |
| Transparent proxy | The host proxy that intercepts TCP traffic other than HTTP and HTTPS. It enforces policy and injects nothing | 3.4 |
| TTL | The time-to-live of a cloud sandbox, 1 hour by default, under a 24 hour ceiling from creation | 7.1 |
| Workload | A kit of kind: workload: the root filesystem with entrypoint, cmd, env, user, and workdir, exactly one per composition | 4.1 |
| Workspace | The host directory a sandbox mounts at the same absolute path, read-write by default, with extra paths marked :ro | 2.1 |
Index of figures
Each number links to its figure, and the claim is the bold sentence that opens the caption.
| Fig. | The claim it makes |
|---|---|
| 0.1 | Each of the eight arrow styles marks one kind of exchange, and every label is a real command, rule, status, or event name from the capture. |
| 1.1 | sbx owns every layer from the daemon to the guest kernel, and docker-agent owns its binary, its files, and the loop inside the guest. |
| 1.2 | Both paths end with docker-agent in a microVM, but sbx run starts from a template image and docker-agent run --sandbox starts from your agent file. |
| 1.3 | docker-agent stages a kit, has sbx create the VM, and opens two hosts on the proxy, and the model call from inside then leaves through that proxy. |
| 2.1 | A sandbox is absent, running, stopped, or removed, and every sbx verb in this section moves it along exactly one edge. |
| 2.2 | A template moves from a stopped sandbox into the runtime image store, out as a tar, and into a new sandbox with --pull never -t. |
| 2.3 | Everything sandboxd owns on macOS sits under one Application Support directory, with the audit log under Logs and a symlink at ~/.sbx/run. |
| 3.1 | The agent sits on a private Docker Engine inside a guest kernel, and five doors cross the hypervisor line: mount, network, secret, MCP, and SSH agent. |
| 3.2 | The host repository enters the VM read-only at /run/sandbox/source, the agent commits to a private clone, and git-daemon serves that clone back as a host remote. |
| 3.3 | A matching deny ends the walk at once, an allow from any active scope admits the host, and a request that matches nothing is denied as implicit. |
| 3.4 | The first CONNECT is denied inside TLS by the proxy itself, one allow rule later the same CONNECT becomes a forward-bypass tunnel, and both leave a log row. |
| 3.5 | A host allow passes any request to github.com, body included, while a network-policy@2 entry denies one method on one host and leaves the rest open. |
| 3.6 | The sandbox only ever holds sbx-cs-<rand>, the forward proxy swaps it for the stored value on the bound host, and a direct connection gets no swap. |
| 3.7 | The host store holds m101-deepwiki once, each sandbox gets its own gateway at one URL, and a static set is fixed while load changes a dynamic one live. |
| 4.1 | The frontend copies the hello-kit descriptor into one manifest annotation, derives six more annotations from its fields, and records its own release in built-by. |
| 4.2 | One workload and two mixins merge into one grant set where allows union, deny wins, and a removed deny on the next version stops for approval. |
| 4.3 | A v2 kit is validated and packed on the host, signed and pushed with two referrers, verified from the registry, and consumed by create or kit add. |
| 4.4 | sbx env plan turns the file into margin-marked rows, approval on create records them, and a later plan prints only the rows that moved. |
| 5.1 | run loads the agent file, alternates model calls and tool calls in one loop, and writes the answer, the events, a session, and with --record a cassette. |
| 5.2 | files.yaml resolves local to the dmr provider and finds the endpoint itself, while dmr.yaml resolves qwen through providers.runner to a fixed URL. |
| 5.3 | Each toolsets entry becomes named tools in the model request: built-in types run in process, and mcp goes through a gateway, a child process, or a URL. |
| 5.4 | transfer_task sends writer only the task and returns its sentence to root, while the session move makes reviewer answer every later message. |
| 5.5 | The preempting hook allows every call only as advice, the patterns decide rm and echo, and the safety mode alone decides pwd. |
| 5.6 | The two replays of files.yaml match on all five turns, the guarded run differs at turn 0, and eval scores the same agent per tool call. |
| 6.1 | A session is created first and stored, and one POST to the agent path returns the whole turn as SSE frames from team_info to stream_stopped. |
| 6.2 | The same pong.yaml answers an editor through stdin and stdout and an MCP client through HTTP POSTs to port 8081, with different method names on each path. |
| 6.3 | The client reads the card, sends one blocking SendMessage that makes one model call, and gets back a completed task that GetTask returns again. |
| 6.4 | One share push stores files.yaml as an OCI manifest with four annotations, and share pull and run both read it back by the same reference. |
| 6.5 | Five callers reach the same runner through five base URLs, and every path after them is open to any client that reaches the port. |
| 6.6 | Compose pulls the model, injects two variables into printer, and runs the gateway with the API socket, which refuses a request without its bearer token. |
| 7.1 | sbx move copies the sandbox filesystem as one image and nothing else, so secrets, mounts, local rules, and processes stay behind while the destination starts a TTL clock. |
| 7.2 | One policy decision becomes one JSONL record on the developer machine, and only an enforced organization policy makes the daemon write it. |
| 7.3 | Each old name survived its successor for months, and Docker Desktop removed the two of them one week apart, on 2026-06-29 and 2026-07-06. |
| 7.4 | Eight claims were confirmed by a recorded file, five are reported without a test, and the community disputes one more. |
Sources:the definition paragraphs of sections 2.1 to 7.2 as research/plan.md plans them; research/sources/help-sbx.md (sbx create claude, sbx template, sbx kit, sbx policy, sbx policy init, sbx policy allow network, sbx policy log, sbx secret, sbx secret set-custom, sbx mcp ls, sbx run, sbx template load, sbx ttl, sbx policy profile, sbx daemon status, sbx skills); research/sources/help-docker-agent.md (run, eval, serve); research/sources/agent-schema.json (AgentConfig, HooksConfig, HarnessConfig, Toolset); research/sources/SPEC-v3-at-v3.0.0-m.8.md §1, §3.4, §7; research/sources/docs-sandboxes.md (pages architecture, configuration/credentials, governance/audit); research/sources/docs-docker-agent.md (pages concepts/models, concepts/agents, configuration/flavors, configuration/permissions, features/sessions, features/cli); the figure blocks of front.md and every file under sections/
Docker Sandboxes and Docker Agent 101, edition 2026.10, by Rohit Ghumare. It is part of AI Engineering from Scratch, an open source course at aiengineeringfromscratch.com. The newest edition is always at aiengineeringfromscratch.com/manual-docker-sandboxes-101.html.
Docker Sandboxes (sbx) and Docker Agent (docker-agent) sbx 0.47.0, docker-agent 1.149.0 is the subject, at commit bf4169cdd31229d52385410c52c3dcc59b497858 of 2026-10-07, checked on 2026-10-08. Every listing comes from a recorded run of the capture kit in the manual's source directory, and every quoted rule is checked word for word against the vendored sources.
© 2026 Rohit Ghumare · MIT license. You can copy and share this manual. Keep this page and the copyright line with it. Report an error at github.com/rohitg00/ai-engineering-from-scratch/issues.