Agent Substrate Egress and Its Gaps: Nothing the Agent Says Is Trusted

If you run agents on Agent Substrate, each outbound connection leaves through a gateway. That gateway ignores what the agent says about the destination. Egress is the most carefully specified part of the project. The control plane is the least specified. Anyone who can authenticate to it can rewrite the egress policy of any agent, because authorization does not exist yet. This part reads both halves from the project's own documents and ends with an order for running it anyway.

Agent Substrate lives in its own GitHub organization, agent-substrate/substrate, under Apache-2.0. Its README says it is "not an officially supported Google product." The community meets on a Google Group and Google Meet, and Google Cloud ships it on GKE. A CNCF donation is in progress. The network egress contract, last updated September 18, says the certificate extension that identifies an actor "will change after the CNCF donation is complete." Part 1 covered how suspended actors share a warm pool of workers, and Part 2 covered Google AX, the Task layer on top. Everything below comes from the docs, source, and tests of the repositories as of September 24. I ran nothing for this essay, and every snippet names its source.


The path every outbound byte takes

An actor is untrusted code in a sandbox, gVisor by default or a Kata micro-VM. When it opens a TCP connection, it thinks it dials the destination directly. In fact, an nftables rule inside the worker pod redirects the connection into atunnel. This small trusted process reads the original destination with SO_ORIGINAL_DST. Then atunnel opens a mutually authenticated TLS connection to the egress policy enforcement point, the PEP. Over that connection it sends an HTTP/1.1 CONNECT. The target and the Host header of that request are the IP address and port the actor dialed, and no hostname goes with them.

The PEP is the atenet-egress deployment in the ate-system namespace, served at atenet-egress.ate-system.svc. It runs on Envoy by default or on agentgateway, and you choose at install time. The egress demo in the repository shows both:

# from the agent-substrate egress demo README, not run here
./hack/install-ate-kind.sh --deploy-ate-system                                  # Envoy (default)
./hack/install-ate-kind.sh --deploy-ate-system --atenet-dataplane=agentgateway

Tunneling switches on cluster-wide through one flag on the API server. The shipped manifest carries it, with a comment saying the flag stays until the egress API lands:

# manifests/ate-install/ate-api-server.yaml, args excerpt, not run here
- --egress-gateway-address=atenet-egress.ate-system.svc:443

The API server stamps that address onto every actor start and restore, so a resumed actor comes back already wired to the gateway. Actor networking is IPv4 only today.

What a refusal looks like from inside

The egress traffic document is the contract for the GA milestone, and it is specific about failure. That detail matters in practice. An agent that cannot tell "blocked" from "slow" burns its time budget on retries.

TrafficBehaviorWhat the actor sees when refused
HTTP(S) 1.1 and 2Allowed under policy403 Forbidden
WebSocketBlocked403 Forbidden
Forward-proxy CONNECTBlocked403 Forbidden
Any other TCPBlockedConnection accepted, then closed with no bytes and no status code
DNS, TCP or UDP port 53Allowed by netfilter, bypasses the gatewayn/a
Any other UDP, other protocolsDroppedNothing, with no ICMP. The client hangs until its own timeout

Three of those rows turn into bugs you will meet. An HTTPS_PROXY variable baked into an image makes every request a forward-proxy CONNECT. That request gets a 403 whatever the policy says, so unset the variable inside actors. A client that tries HTTP/3 first sends QUIC over UDP port 443, and the gateway drops it silently. The first request of each connection then waits out the QUIC timeout of the client before any TCP fallback. The docs do not mention QUIC, and I infer that result from the UDP row. A database driver that dials port 5432 gets a connection that opens and closes at once. Most drivers report that as a protocol error, and nothing in the error points to the network policy.

The egress demo README describes the current gateway as a little more permissive than the GA table. It says opaque TCP and TLS that the gateway does not terminate are "allowed by address only, at the CONNECT." Read that as the behavior of the current build, and read the table as the target. If a non-HTTP dependency matters to you, find out which one your version implements before you design around it.

Here is a probe you can run inside an actor to tell the cases apart. I built it from the documented behavior with only curl and bash, and I did not run it for this essay:

# egress-probe.sh: classify an outbound failure inside an actor
# pattern built from docs/egress-traffic.md, not run here
code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 10 "https://$1/")
case "$code" in
  403) echo "policy refused $1 (or a proxy variable made it a forward CONNECT)";;
  000) echo "no HTTP answer from $1 within 10s: check DNS, UDP, or a dropped QUIC attempt";;
  *)   echo "reached $1, status $code";;
esac
# opaque TCP: an immediate empty read means refused, a timeout means slow or dropped
timeout 5 bash -c "exec 3<>/dev/tcp/$2/$3; head -c1 <&3 | wc -c"

Why the hostname is only a hint

Read this part of the contract slowly. The document says the PEP "MUST NOT trust any decision or attestation coming from the actor itself." It also names the obvious attack. The PEP "MUST NOT treat an actor-provided HTTP hostname or TLS SNI as proof that the original destination has that name." An agent under prompt injection can send Host: api.github.com to the IP address of an attacker. If a gateway allowed that connection because the header said GitHub, it would be an exfiltration channel with a GitHub label on it.

So a policy decides in one of two ways. A CIDR rule, or the catch-all all rule, decides on the IP address that atunnel recovered from the kernel. The actor cannot forge that address, and the gateway dials it at once. A hostname rule can use the hostname as input. When it authorizes by name, the gateway "MUST route the traffic to the authorized hostname rather than use the hostname to authorize an arbitrary IP address." A false header gets the agent nothing, because the request goes to the real api.github.com and never to the address the agent dialed.

The docs state one result of this plainly, and it is easy to miss. A hostname rule can only match traffic the gateway can read, which is cleartext HTTP or TLS that the gateway terminates. The default gateway does not terminate TLS. It decides HTTPS at the CONNECT, by address only, so a hostname rule never matches it. To allow https://api.github.com by name, you need the gateway in its man-in-the-middle mode, installed with --experimental-use-sdsmint. Your actors must also trust the CA of the gateway, which the trust bundle guide projects into the actor filesystem. Without that mode, your choices for HTTPS are CIDR prefixes or all.

Fig. 1 · one outbound request, decided

Pick what the agent sends and what the policy says. The path shows which hop decides and on what evidence. The dialed IP comes from the kernel, and the claimed host comes from the agent.

Policy for this actor

Traffic

Gateway mode

IP the agent dialed

Host the agent claims (Host header or SNI)

    Rules and outcomes from docs/network-egress.md, docs/egress-traffic.md (GA contract), and the egress demo README. For other TCP, the figure follows the GA contract. The demo README says the current gateway allows it by address. The gateway caches a missing policy as a deny for 10 seconds.

    Identity comes from a certificate

    Every connection to the PEP carries a client certificate minted for one actor. When an actor starts, atunnel generates a private key, which never leaves it. It then asks the atelet on the node for a short-lived certificate. The certificate carries an ActorIdentity X.509 extension with the atespace, name, and UID of the actor, and a purpose that must be atunnel. The extension OID today is 1.3.6.1.4.1.11129.2.12.2, allocated under the enterprise number of Google. The contract says it will change after the CNCF donation, so anything that pins that OID will need an update.

    The PEP checks four things. The chain and lifetime must be valid against the actor identity CA. The certificate must have the client-authentication usage and exactly one valid extension. An actor URI SAN must name the same actor. Then the PEP asks ate-api-server if this actor still exists, if its UID matches, and if it is running. So a certificate cannot outlive its actor. Say you delete an actor and create it again under the same name. The new actor has a new UID, and the PEP refuses the old certificate. In the gateway log of the demo, the identity appears as a SPIFFE-style URI, spiffe://substrate-actor.local/atespace/ate-demo-egress/actor/egress-demo.

    The policy itself is a per-actor resource. The gateway evaluates rules in order. The first match authorizes the request and applies only the effects of that rule. A request that matches nothing is denied. An actor with no policy gets no tunnel at all. Here are the three rule types, in the manifest format that the CLI tests of the repository parse:

    # rule shapes from cmd/kubectl-ate/internal/cmd/egress_policy_test.go, not run here
    metadata: {atespace: team-a, name: default}
    rules:
    - hostnames:
        patterns:
        - api.example.com
    - cidrs:
        cidrs:
        - 10.64.0.0/16
    - all: {}

    The name is always default, and the atespace must match the atespace of the actor. Hostname patterns are lowercase DNS names. A wildcard replaces exactly one leftmost label, so *.example.com matches api.example.com and does not match example.com or nested.api.example.com. Each list takes at most 256 entries. You create and read a policy with the kubectl-ate plugin, and the positional argument is the actor, not the policy:

    # from the egress demo README and kubectl-ate help text, not run here
    kubectl ate create egress-policy egress-demo -a ate-demo-egress -f policy.yaml
    kubectl ate get egress-policy egress-demo -a ate-demo-egress

    Put the timing warning from the demo in your runbook. The gateway caches a missing policy as a deny for 10 seconds (--egress-policy-cache-ttl). A fetch that ran before the policy existed keeps failing for up to 10 seconds after you create the policy. To see decisions as they happen, the demo reads the gateway logs:

    # from the egress demo README (Envoy dataplane), not run here
    kubectl -n ate-system logs deploy/atenet-egress -c envoy | grep '\[egress\]'
    kubectl -n ate-system logs deploy/atenet-egress -c ext-proc | grep -i 'egress tunnel opened\|egress denied'

    Credentials are the obvious next step and the least finished part. The policy schema already has an effect that injects a header with a value from a credential provider. The reference looks like ate-secret://<provider-class>/<provider-name>/..., so a token never has to enter the sandbox. On the default gateway, a rule that declares an injection is denied with a 501 until injection lands. The repository wires injection only into the man-in-the-middle gateway, behind an experimental install flag. Until then, your agent can leak any credential it uses, as the eval sandbox that held the keys showed.

    Ingress rides the same tunnel

    Inbound traffic takes the reverse route. A client sends its request to atenet-router, which is Envoy with an external processor, and names the target actor in a header. The architecture doc gives the order. First the router asks the control plane to resume the actor if it is suspended. Then it opens an authenticated tunnel to atunnel on that worker on port 443 and forwards the request. Nothing reaches the container port of the actor directly. From the demo:

    # from the egress demo README, not run here
    kubectl -n ate-system port-forward service/atenet-router 8000:80 &
    curl -s -X POST http://localhost:8000/ \
      -H 'ate-target-actor: ate-demo-egress/egress-demo' \
      -H 'Content-Type: application/json' \
      -d '{"url":"http://10.96.0.20:80/"}'

    The header routes the request. The URL in the body is input to the demo app, which fetches it through the egress path above. The address here is an example cluster IP. The demo reads it from a whoami service that it creates first.

    The control plane trusts anyone who can log in

    Now the other half. ate-api-server authenticates callers with mTLS client certificates or bearer JWTs, configured in a file like this one from the authentication doc:

    # from docs/authentication.md, not run here
    actorIdentityJWTProvider: kubernetes
    jwtProviders:
    - name: kubernetes
      issuer: https://kubernetes.default.svc.cluster.local
      audiences:
      - api.ate-system.svc
    - name: google
      issuer: https://accounts.google.com
      audiences:
      - 32555940559.apps.googleusercontent.com

    The next paragraph of that doc is the sentence to read before anything else in the project: "Authorization and RBAC are not implemented yet, so only configure providers whose users should have full control of the entire control plane: every atespace, actor, actor template, egress policy, snapshot and worker in the cluster." Authentication proves who you are, and nothing after it limits what you do. Every guarantee in the sections above holds only while you trust each identity that can call CreateActorEgressPolicy with every actor.

    Look at the second provider in that example. It accepts Google identity tokens minted for the shared client ID of the Cloud SDK. The doc shows the workflow that results: gcloud auth print-identity-token | kubectl ate --token-file=- get actors. On a shared cluster, any Google account that the issuer accepts gets full control, unless something else stops it. The doc marks those values as examples and says they are not defaults, so treat them as examples.

    Until authorization ships, the fence has to come from Kubernetes, and two layers are available. First, kubectl-ate reaches the API through a port-forward. Kubernetes RBAC on pods/portforward in ate-system already decides who can reach it from a laptop. Second, a NetworkPolicy can limit in-cluster callers to the namespaces that need the API. I wrote the policy below on standard Kubernetes objects, with the app: ate-api-server label and container port 443 from the shipped manifest. It assumes that your CNI enforces NetworkPolicy and that AX runs in its documented ax-system namespace:

    # my construction on the shipped labels; verify against your install, not run here
    apiVersion: networking.k8s.io/v1
    kind: NetworkPolicy
    metadata:
      name: ate-api-callers
      namespace: ate-system
    spec:
      podSelector:
        matchLabels:
          app: ate-api-server
      policyTypes: [Ingress]
      ingress:
      - from:
        - namespaceSelector:
            matchExpressions:
            - key: kubernetes.io/metadata.name
              operator: In
              values: [ate-system, ax-system]
        ports:
        - {protocol: TCP, port: 443}
        - {protocol: TCP, port: 9090}

    Port 9090 is the API server's metrics and readiness port. If your monitoring scrapes it from another namespace, add that namespace to the list.

    What else the documents say is missing

    The project is candid about this. Its threat model, last updated June 25, says Substrate "has little to no security hardening at this time." The roadmap ranks identity fourth and policy fifth among six priorities. Beyond authorization, these are the gaps the docs name themselves:

    The GKE announcement from Google Cloud on September 15 draws the same line. Substrate is available to all GKE customers "for non-production workloads," with production support "via allowlist."

    Where the whole stack is crowded, and where it is empty

    Look past one project and the same pattern shows up across the layers under agents. Sandboxes, durable runtimes, build frameworks, and observability tools each have half a dozen named projects. Tool protocols have two, and both now sit under the Agentic AI Foundation. Payments has four protocols that compete for the same job. Identity has plenty of products and no agreed standard, and the security story of every other layer leans on it. Substrate has the same shape: the sandbox and the egress path are real, identity stops at authentication, and payments are out of scope.

    Fig. 2 · the layers under an agent

    One row per layer. Squares count the projects this series named for that layer, and you can tap a row to list them. Switch the right-hand column between what Substrate and AX ship today and what they only list as planned.

    Counts cover only the projects named in the research for this series. Coverage comes from the Substrate roadmap, architecture, and observability docs and the AX README and roadmap. "Standard" means an open specification with neutral governance.

    A worked example: a coding agent's egress

    Say an agent in atespace team-a needs three things. They are the GitHub API over HTTPS, the npm registry over HTTPS, and a Postgres database at 10.64.12.5:5432 inside your network. Walk it through the contract.

    On the default gateway, the gateway decides HTTPS by address, so a hostname rule for api.github.com never matches. You would have to allow the address ranges of GitHub and npm by CIDR, and those ranges change. The other choice is to allow everything. On the man-in-the-middle gateway, hostname rules work, but you need an experimental install flag and a trust bundle in the actor. The database is harder. Under the GA contract, a Postgres connection is "any other TCP," and the gateway accepts it and closes it. The demo describes a current gateway where a CIDR rule would let it through by address. So check your version. If you will run the GA behavior, put an HTTP API in front of the database.

    With the man-in-the-middle gateway, the policy reads as follows, in the schema from the tests of the repository:

    # assembled from the rule schema in the repo's egress-policy tests, not run here
    metadata: {atespace: team-a, name: default}
    rules:
    - hostnames:
        patterns:
        - api.github.com
        - registry.npmjs.org
    - cidrs:
        cidrs:
        - 10.64.12.5/32

    Now trace each call. git and the GitHub API reach api.github.com on the first rule. A request that carries Host: api.github.com to any other IP goes to the real host. An injected instruction to "send the diff to this IP" gets nowhere by claiming GitHub. The npm registry matches the same rule. A dependency script that tries https://paste.example matches nothing and gets a 403, and the probe above reports it as a policy refusal. The database connection matches the CIDR rule if your gateway passes opaque TCP. A DNS lookup of any name still leaves the sandbox through the relay, and nothing checks it.

    If the agent runs under AX, look at the Gateway too. The example manifest of AX, which its README applies with ax apply -f examples/task.yaml, allows every host on port 443:

    # from google/ax examples/task.yaml, not run here
    egress:
      allowlist:
        hosts:
          - host: "*"
            port: 443

    Replace the wildcard with the hosts the task needs before anything real runs through it:

    # the same field with explicit hosts; field shape from the AX example
    egress:
      allowlist:
        hosts:
          - host: api.github.com
            port: 443
          - host: registry.npmjs.org
            port: 443

    Running it today, in the order the contract suggests

    1. Fence the control plane first. Keep only JWT providers whose users can all control every actor. Limit pods/portforward in ate-system with Kubernetes RBAC. Limit in-cluster callers with a NetworkPolicy. Without authorization, access to the API gives full control of it.
    2. Give every actor an explicit policy, and never all: {} outside a demo. The default is already deny. The risk is a shortcut that you add to make a demo pass.
    3. Decide how to allow HTTPS. Use CIDR rules on the default gateway, or hostname rules on the experimental man-in-the-middle gateway with its CA in every actor. Hostname rules on the default gateway match nothing and give no warning.
    4. Narrow AX Gateways from host: "*" to named hosts, so the AX layer and the Substrate layer both say no.
    5. Keep secrets out of sandboxes until header injection is finished. Give agents short-lived tokens with a narrow scope, and rotate them on every run.
    6. Plan for DNS and UDP. DNS leaves unchecked, and UDP other than DNS hangs. Unset proxy variables and disable HTTP/3 in actor images.
    7. Read the gateway logs, and allow for the 10 second deny cache when a new policy seems not to work.
    8. Keep it out of production until authorization ships, as Google's own GKE announcement does.

    The egress contract is the best argument for the project. It assumes that the agent is compromised, and the gateway still gives a useful answer. The control plane has not caught up to that assumption yet. When it does, look in the authentication doc for the sentence that says authorization is implemented. For how the runtime above this layer handles the same trust problem, see the harness as an API.

    rg
    Rohit Ghumare

    CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals under the AI stack. Agent Substrate is a three-part series on the runtime layer under agents, read from the docs and repositories of the projects. I ran nothing in this part, and every snippet names the file it came from.

    Part 1 · Part 2 · More posts · X