Agent Substrate: How Suspended Agents Share a Warm Pool
If you run agents in containers today, most of the compute you reserve sits waiting on a model reply, a tool, or a person. Agent Substrate changes the unit you schedule. An agent becomes an actor that lives in a snapshot, and it borrows a warm worker pod only while it does work. This part covers how that swap works, what a snapshot holds, and how to size the pool.
Agent Substrate is an open-source runtime under Apache-2.0, published under its own GitHub organization, agent-substrate/substrate. Its README carries the line "This is not an officially supported Google product." Community meetings, the mailing list, and meeting notes run on Google tooling. The network egress doc says the certificate extension for actor identity "will change after the CNCF donation is complete."
Google Cloud announced it on May 20, 2026, in the same post that introduced Agent Executor. On September 15, Google made it available on GKE for non-production workloads. Production support is "available via allowlist," per that post. The project is in early development, and "the APIs are almost guaranteed to change," the repository says. Treat everything below as a reading of the design and docs as of today. Nothing here was run on a cluster. Every manifest and command comes from the project's docs and demos.
Where it sits in the stack
This series reads one vertical slice of the agent stack, bottom to top. Kubernetes owns nodes and pods. Agent Substrate runs on top of Kubernetes and owns the fast part: which sandboxed actor is live on which pod right now. Google's AX, short for Agent Executor, sits above Substrate and gives you tasks, workspaces, gateways, and models. Part 2 covers AX.
Above AX sits the harness that runs the agent loop. The Substrate README names Claude Code, Codex, and Antigravity as workloads it can host. Above the harness sits the framework that authored the agent, such as ADK or LangChain. MCP, A2A, and AP2 do not live in one layer. Any layer uses them to reach tools, other agents, and payments.
"AX" also has a second meaning. Netlify co-founder Mathias Biilmann uses it for agent experience, a design discipline in the spirit of UX and DX. That usage is unrelated to Google. In this series, AX always means the Agent Executor runtime.
"It is not an SDK for building agents, but rather a system for running them at scale," the Substrate README says of its scope. It does not know what a model is, and it does not check whether the workload is an agent at all.
Agents mostly wait
"Agents and 'agent-like' workloads are generally very bursty, spending most of their time waiting for input or events," the architecture doc says. They also run untrusted code, so each one gets its own sandbox. A deployment also runs a great many of them. The usual result is thousands of single-tenant pods. Each pod holds CPU and memory for an agent that is idle almost all the time.
Substrate separates two things that a pod normally fuses. An actor is one instance of a workload, identified by an atespace and a name. A worker is a pre-started pod that waits to receive an actor's state. A worker hosts at most one actor at a time. The control plane maps a large population of mostly suspended actors onto a small pool of workers. The README demo juggles about 250 stateful actors across 8 pods, which the project calls "30x+ oversubscription."
Two sets of numbers circulate for this. They measure different things, and neither set was measured here. The architecture doc lists north star targets. They are 100 ms activation latency at the 95th percentile, 1 billion actors per cluster, and 1000 wakeup events per second. The README and the GKE post report two figures that Google claims for the shipped code. The first is "sub-500ms resume operations at over 500 suspend/resume activations per second," and the second is "10x higher density than standard container runtimes."
Each square is an actor. Blue ones are running right now. The rest are suspended in object storage and cost no pod. Set the number of actors, the share of time each one is busy, and the headroom you want. The figure shows how many workers the pool needs.
Actor population
Worker pool
Illustrative arithmetic, not a Substrate tool. Running at once is the population times the busy share. Workers are that number times one plus headroom, rounded up. Real traffic is bursty and correlated, so size the pool for measured peaks. The default of 250 actors on 8 pods mirrors the README demo.
Two kinds of object, two kinds of storage
The worker side is ordinary Kubernetes. A WorkerPool is a custom resource in the ate.dev/v1alpha1 group that the atecontroller turns into a Deployment of worker pods. The counter demo's pool, from the repo:
apiVersion: ate.dev/v1alpha1
kind: WorkerPool
metadata:
name: counter
namespace: ate-demo-counter
labels:
workload: counter
spec:
replicas: 3
workerImage: ko://github.com/agent-substrate/substrate/cmd/ateom-gvisor
template:
nodeSelector:
ate.dev/substrate-version: "${SUBSTRATE_VERSION}"
resources:
limits:
cpu: "1"
memory: 1Gi
requests:
cpu: 250m
memory: 1Gi
From the substrate docs (demos/counter), not run here.
Two details in that file matter when you write your own pool. The comments explain that capacity is read from the limits, while the Kubernetes scheduler packs nodes by the requests. So the CPU request sits below the limit to keep the pool placeable. The nodeSelector on ate.dev/substrate-version matters too. The README warns that "a node added later hosts no workers until you label it with the installed version."
The actor side stays out of Kubernetes on purpose. Actors and workers are records in a PostgreSQL-backed store. The architecture doc explains: "The Kubernetes API server is not designed to handle millions of resources." Actor state also changes many times a second.
The docs say an ActorTemplate is "not a CRD." It lives in the Substrate control plane under an atespace. You create it through the ate API with the kubectl ate plugin. Below is the counter demo's template, trimmed to the fields that this part discusses:
metadata:
atespace: ate-demo-counter
name: counter
workerSelector:
matchLabels:
workload: counter
containers:
- name: counter
image: ko://github.com/agent-substrate/substrate/demos/counter
wakeupProbe:
httpGet:
path: /readyz
port: 80
volumeMounts:
- name: data
mountPath: /home/counter
resources:
limits:
- name: cpu
quantity: "1"
- name: memory
quantity: 512Mi
snapshotConfig:
onPause: SNAPSHOT_CONTENT_SCOPE_FULL
onCommit: SNAPSHOT_CONTENT_SCOPE_FULL
storageLocation: gs://${BUCKET_NAME}/ate-demo-counter/
sandboxConfig:
sandboxClass: SANDBOX_CLASS_GVISOR
configName: gvisor-default
volumes:
- name: data
durableDir: {}
From the substrate docs (demos/counter/counter-template.yaml.tmpl), trimmed, not run here.
The workerSelector matches the pool's metadata labels, and that is how a template claims its pool. The template's own resources size the sandbox. The pool's resources size the pod that hosts it.
That split gives the practical sizing rule. You scale the worker pool on how many actors run at the same moment, not on how many exist. A pool sized to the population pays for idle pods, and Substrate exists to remove idle pods. A pool sized to the mean with no headroom turns every burst into waiting. The repo ships an autoscaled WorkerPool demo that drives a HorizontalPodAutoscaler from the ate_workerpool_workers metric. That metric counts assigned workers, so the pool can follow real concurrency instead of a guess.
One actor's life
The architecture doc gives the lifecycle as a state machine. An actor is created suspended, and a resume moves it to running on some worker. From running it can be suspended, which uploads a checkpoint and frees the worker, or paused, which keeps a short-term checkpoint on the node. A paused actor can resume on the node that holds its snapshot. It can also be suspended, which uploads the node-local snapshot and ends the pin.
[*] --> SUSPENDED : CreateActor
SUSPENDED --> RESUMING : ResumeActor
RESUMING --> RUNNING : restore / boot complete
RUNNING --> SUSPENDING : SuspendActor
SUSPENDING --> SUSPENDED : checkpoint complete
RUNNING --> PAUSING : PauseActor
PAUSING --> PAUSED : node-local checkpoint complete
PAUSED --> RESUMING : ResumeActor (pinned to the snapshot's node)
PAUSED --> SUSPENDING : SuspendActor (uploads the node-local snapshot)
SUSPENDED --> [*] : DeleteActor
From docs/architecture.md, verbatim.
RevertActor takes a running, paused, or crashed actor back to suspended. It discards any local pause checkpoint and keeps the last completed snapshot. Eviction has a deadline, per the API guide. An actor whose worker pod is evicted gets SIGTERM and 30 minutes to be suspended. After that, the actor is killed and marked crashed, and it loses everything since its last snapshot.
Drive one actor, my-counter-1, through the documented transitions. Watch where its state lives: on a worker, on a node's disk, or in object storage. Also watch which worker it lands on next.
Worker pool
Node disk
Object storage
Transitions follow the state machine in docs/architecture.md. The docs state that a resume after a pause is pinned to the snapshot's node. They do not say whether the worker is released during the pause, so the figure marks that worker as not documented. The scheduler chooses which free worker a suspended actor lands on. The figure rotates through free workers to show that the actor can move.
The resume path is the core of the design. In the counter demo, you create an actor from its template and reach it through the router with one header:
# create the actor from the template in its atespace
kubectl ate create actor my-counter-1 -a ate-demo-counter --template counter
# expose the router locally
kubectl port-forward -n ate-system svc/atenet-router 8000:80
# any request with this header resumes the actor if it is suspended
curl -X POST \
-H "ate-target-actor: ate-demo-counter/my-counter-1" \
http://localhost:8000
# inspect, hibernate, and remove it
kubectl ate get actor my-counter-1 -a ate-demo-counter
kubectl ate suspend actor my-counter-1 -a ate-demo-counter
kubectl ate delete actor my-counter-1 -a ate-demo-counter
From the substrate docs (demos/counter/README.md), not run here.
Behind that curl, atenet-router, which is Envoy with an ext_proc external processor, reads ate-target-actor. It asks ate-api-server to resume the actor, then waits. The control plane claims a free worker from the actor's pool. The node agent atelet tells ateom inside the worker pod to restore the snapshot. The router then opens an authenticated tunnel to the worker's atunnel listener on port 443 and forwards the original request. The client sees one slow request instead of an error.
Substrate never suspends an idle actor on its own. Idle detection belongs to the layer above, and AX takes on that job in part 2. The architecture doc describes the handoff. When the user or a higher-level system is done with an actor, "it can request the Agent Substrate to suspend the actor."
When the pool is full for a moment, the router holds the request instead of failing fast. Its request parking feature retries the resume with backoff until a worker frees up or the park budget runs out. The --parked-request-budget flag sets that budget, and the default is 5 seconds. When the budget runs out, the router does not cancel a resume that is already in flight, so a slow restore is served late. Without a free worker, the client gets a 503 reading "no free workers available."
What a snapshot holds
Resume speed and correctness both depend on what the snapshot captured. Substrate makes that a per-template choice between two scopes. A Full snapshot holds process memory, the root filesystem delta on top of the image, and any DurableDir volumes. That is everything an actor needs to come back hot. A Data snapshot holds only the DurableDir contents and discards memory and the rest of the filesystem. It is cheaper to write and store.
You set the scope per trigger. The onPause trigger picks what a pause captures on the node, and the onCommit trigger picks what a suspend uploads. The onCommit scope must be a subset of the onPause scope. Every template also gets a Golden Snapshot, a full capture taken once from a temporary boot when the template is created. A new actor first resumes from the golden snapshot. After that, each actor resumes from its own last snapshot.
A Data snapshot needs a boot source for everything it dropped. The onResume.fromData field picks that source. ColdBoot, the default, starts the containers fresh and restores the volume contents. Golden restores the golden snapshot's warm memory and serves the actor's own data on top. Golden is currently micro-VM only.
The counter demo uses Full for both triggers so that an in-memory counter survives a suspend. Its README describes the alternative: with Data scope, "only the durable-volume counter would survive."
If you write a workload for Substrate, follow these rules:
- If a value must survive a Data-scope suspend, write it under a
DurableDirmount, never only in memory. - Do not read an actor's identity from an environment variable. A variable baked into the golden snapshot stays frozen for every actor restored from it. The API guide projects identity fields as files on a per-actor mount. It does this "precisely so they carry the correct values after a resume from a shared snapshot."
- Give every container a
wakeupProbe.ResumeActorreturns only after each probed container answers 200. When every container declares one, the template controller also skips its default wait of about 20 seconds before it takes the golden snapshot.
The number of durable volumes depends on the isolation class. A micro-VM template can declare several, because they are subdirectories of one shared virtio-fs mount. A gVisor template gets one until gVisor accepts more than a single durable mount.
Two isolation classes
Each WorkerPool picks a sandbox class, and each class has its own ateom image that drives checkpoint and restore inside the worker pod.
- gVisor (
ateom-gvisor, the default) runs the workload underrunsc. It uses gVisor's own checkpoint and restore of the sandboxed process tree. The architecture doc notes it "currently requires arunscversion with the--allow-connected-on-saveflag to work around a bug in networking resumption during checkpointing." - Micro-VM (
ateom-microvm) runs the workload in a Kata Containers guest on the Cloud Hypervisor VMM. It captures a memory-only VM snapshot. It restores the snapshot on demand withuserfaultfddemand paging, so pages load as the guest touches them.
The sandbox binaries are not baked into worker images. They come from a cluster-scoped SandboxConfig, which the template names through sandboxConfig.configName. The gvisor-default reference in the template above is one example.
The docs show the trade between the two classes. gVisor is the default, and the demos use it. Micro-VMs get multiple durable volumes and the Golden resume source today, and they need /dev/kvm on the node. The repo's local guide covers Linux hosts with KVM and Apple Silicon Macs through Lima.
What each component owns
| Component | Runs as | Owns |
|---|---|---|
ate-api-server | control plane service | actor lifecycle, scheduling onto workers, and snapshot coordination, with state in PostgreSQL |
atecontroller | Kubernetes controller | reconciles CRDs, for example a WorkerPool into a Deployment |
atelet | node DaemonSet | pulls images, drives sandbox lifecycle through ateom, streams snapshots to and from GCS or S3 |
ateom | inside each worker pod | runs, checkpoints, and restores the workload with runsc or Kata and Cloud Hypervisor, and hosts atunnel |
atenet-router | Envoy with ext_proc | ingress: resolves the target actor, triggers resume, parks requests |
kubectl-ate | kubectl plugin | create, get, suspend, and delete actors and templates |
Summarized from docs/architecture.md and docs/glossary.md.
The split follows the premise. Kubernetes handles the slow, low-frequency work of creating and scaling pods, which it does well. Substrate handles the high-frequency work of deciding which actor occupies which pod right now. It keeps that data in a store built for that write rate. Outbound traffic goes through atunnel and an egress policy point that trusts nothing the actor claims. Part 3 covers that contract and the parts that are not built yet.
Worked example: a pool for coding agents
Take a team that runs 400 coding agents, each of which wakes for model calls and tool runs. Suppose your own traces show that a typical agent does work 5% of the time. At the busiest minute of the day, that share rises to 9%. These numbers are illustrative, so measure your own.
- Mean concurrency is 400 times 0.05, or 20 running actors. Peak concurrency is 400 times 0.09, or 36.
- Sizing to the mean with 25% headroom gives 25 workers. That covers ordinary load, and peak bursts wait in the park.
- Sizing to the peak with 10% headroom gives 40 workers. That is ten times fewer than 400 always-on pods.
- Each worker pod in the demo pool requests 1 GiB of memory with a 1 GiB limit. So 40 workers hold 40 GiB, where 400 pods would hold 400 GiB.
Check four things before you trust the plan:
- Each template's memory limit fits inside the worker's limit, because the pool advertises its limits as the per-actor ceiling.
- The snapshot bucket is in the same region as the cluster, because every resume after a suspend streams the snapshot back from it.
- Every new node carries the
ate.dev/substrate-versionlabel. - Your clients tolerate a request that can wait up to the park budget before it is served.
When it goes wrong
| Symptom | Likely cause | Fix from the docs |
|---|---|---|
| Requests return 503 "no free workers available" | pool smaller than peak concurrency for longer than the park budget | raise replicas, autoscale on ate_workerpool_workers, or raise --parked-request-budget |
| New nodes never host workers | node lacks the version label | kubectl label node <node> ate.dev/substrate-version=<build version>, or put the label on the node pool |
| In-memory state is gone after resume | Data-scope commit, so memory cold-boots | use Full scope on commit, or keep the state in a DurableDir |
| Every actor reports the same identity | identity read from an environment variable frozen in the golden snapshot | read the per-actor metadata files instead |
| Actor marked crashed after a node drain | eviction took longer than the 30 minute suspend window | suspend actors before draining, and treat the last snapshot as the recovery point |
| Resume works for memory but connections break | gVisor networking resume bug | use a runsc build with --allow-connected-on-save |
Adoption steps
- Before you touch the cluster, measure from traces how much of the time your agents are busy. The whole case rests on that number.
- Decide what state must survive a suspend. Then pick Full or Data scope per trigger to match.
- Move anything that must survive a Data-scope suspend onto a
DurableDir. Read identity from the projected files. - Pick the isolation class. Use gVisor unless you need several durable volumes, the Golden resume source, or VM-level isolation.
- Add a
wakeupProbeto every container so resumes block until the workload is ready. - Size the pool for measured peak concurrency. Then put an autoscaler on assigned workers.
- Decide who suspends idle actors, because Substrate does not do it. Part 2 shows how AX takes that job.
Substrate has a narrow scope. It does not build agents, pick models, or decide when an agent is idle. Its design assumes that agents spend most of their time waiting, and it releases the pods they would hold while they wait. Part 2 reads the executor built on top of it, and part 3 reads its egress contract and the gaps it leaves.
Keep reading