The Harness Is an API Now. Here Is What You Still Own.

If your agents run long enough to crash, you can now rent the loop that keeps them alive. Three things stay yours. You own the computer the agent works on and the keys that computer can read. You also own the recovery code around an event stream that does not replay. This essay shows each one as requests, config, and failure handling.

On September 10, OpenAI put the agent harness behind a URL. The Agents API, in public beta, takes a POST to /v1/agents/sessions with a model, a tool list, and an environment. From then on, OpenAI's servers run the loop that Codex runs on your laptop. The loop calls the model and executes tools. It summarizes old context when the window fills, sends work to subagents, and continues after a disconnect. The same morning, Baseten announced it had bought Blaxel, a company that builds sandboxes for agents. Together, the two posts split the agent stack in a new place. You now rent the loop from the model vendor. The computer that the loop drives is a separate market.


What moved across the wire

A harness is everything around the model call. I have written a lot about that layer on this blog. Two examples are why it is where the leverage is and how Codex treats compaction as an API primitive. Until now, the harness lived in your process, even when the model ran on someone else's GPU. Your code held the transcript and decided when to summarize. It dispatched each tool call and retried the failed ones. When your process died, the harness died with it.

The Agents API moves that code to OpenAI. The overview page gives the split in one sentence:

OpenAI manages sessions, orchestration, context compaction, and recovery while your application provides tools and chooses its execution environment.

The harness it runs is the open-source Codex harness. You can read the logic that coordinates model calls, tools, and context. You do not run it.

The pricing makes the move cheap to try. OpenAI says the API itself has no extra fees. You pay the token rates of the model you select and standard rates for OpenAI's built-in tools. If OpenAI hosts the sandbox, you also pay container rates.

Four nouns and a turn

The API has four concepts. An agent is the model, instructions, tools, and MCP servers. An environment is an optional computer where the agent runs commands and edits files. A session is a durable instance of an agent working on tasks. Events and items are the data that flows in and the data that gets saved. Events stream live progress. Items are the stored messages and tool calls that you can list later.

The unit of work inside a session is a turn. The sessions guide defines a turn by the session state that your message arrives in. A message sent to an idle session starts a new turn. A message sent while a turn is active steers that turn. Turns run asynchronously. Your application either keeps a stream open or registers a webhook that reports each change of session state. To stop the agent and keep the conversation, send an agent.session.input.cancel event.

Two capabilities that used to take weeks of your own code are now fields. Context compaction happens on the server. The harness "automatically compacts earlier context as a session approaches its context limit," and you write no compaction logic. Subagents are a config block, multi_agent: { enabled: true, max_concurrent_subagents: 3 }. Each subagent keeps its own context, and the main agent merges their results. Tool search loads tool definitions only when the agent needs them. Programmatic tool calling lets the agent run calls in parallel and filter the results in code before they enter context.

The request below creates a whole session. It defines an inline agent with delegation on and points it at a machine that you connect later. The response carries session.id, session.environment.id, and session.environment.remote_url. Keep all three.

curl https://api.openai.com/v1/agents/sessions \
  -H "OpenAI-Beta: agents=v1" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "agent": {
      "model": "gpt-6-astra",
      "instructions": "Prepare release notes from the repository. Have one subagent identify customer-visible changes and another check migration guides and examples, then combine their findings.",
      "multi_agent": {
        "enabled": true,
        "max_concurrent_subagents": 3
      }
    },
    "environment": {
      "type": "self_hosted",
      "workspace_directory": "/workspace"
    }
  }'

From the Agents API multi-agent guide, as of Sep 25, 2026. Every request needs the OpenAI-Beta: agents=v1 header, and the SDKs add it. The coordinator and its subagents share one environment. A new subagent does not get another machine.

Steering and stopping use the same call with different event types. A message that arrives during an active turn redirects that turn. A message that arrives at an idle session starts the next turn.

def send_message(client, session_id, text):
    client.beta.agents.sessions.events.create(
        session_id,
        events=[{
            "type": "agent.session.input.message",
            "input": [{"role": "user", "content": [{"type": "input_text", "text": text}]}],
        }],
    )


def cancel_turn(client, session_id):
    client.beta.agents.sessions.events.create(
        session_id, events=[{"type": "agent.session.input.cancel"}]
    )

From the sessions guide, as of Sep 25, 2026, trimmed to the two helpers. Cancel stops the turn and keeps the session and its saved items.

Three places the tools can run

The environment field separates two things that every local harness combines. One is the loop that decides what to do. The other is the machine that does it. The architecture page offers three values.

The self-hosted path adds a new component. Inside your environment you run codex exec-server, which the docs call the executor. It registers with the API using an environment ID and a restricted key, then opens an outbound WebSocket to codex-cloud-environments.chatgpt.com and waits for commands. "All connections are outbound," the self-hosted guide says, and the executor reconnects if the connection drops. No connection comes into your network. The harness in OpenAI's cloud decides to run npm test. The executor in your container runs it and sends back the output.

On your side, the executor needs two commands and two outbound firewall rules. https://api.openai.com is for registration, and wss://codex-cloud-environments.chatgpt.com is for commands. REMOTE_URL and ENVIRONMENT_ID are the values that the session returned. Use them unchanged on every reconnect.

mkdir -p /workspace
npm install -g @openai/codex@alpha

CODEX_API_KEY="$OPENAI_EXECUTOR_API_KEY" \
codex exec-server \
  --remote "$REMOTE_URL" \
  --environment-id "$ENVIRONMENT_ID"

From the self-hosted guide and the webhooks guide, as of Sep 25, 2026. OPENAI_EXECUTOR_API_KEY is the environment key, never your application key.

If you also rent the computer, change the network setting on day one. Hosted sandboxes allow all outbound traffic by default. Restricted mode takes 1 to 100 exact host names, no wildcards, and redirect targets need their own entries.

{
  "vault_ids": ["vault_123"],
  "environment": {
    "type": "openai_hosted",
    "network": {
      "access": "restricted",
      "allowed_domains": ["api.github.com"]
    }
  }
}

From the vaults guide, as of Sep 25, 2026. Merge these fields into the create-session body next to agent. allowed_domains lets the sandbox reach a host. The allowed_hosts list of a vault credential lets the proxy give that host the secret. Both lists must name the host.

OpenAI lists nine sandbox partners with first-class integrations: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel. The partners compete on things that the harness ignores, such as cold start, GPU shapes, VPC deployment, and file persistence. A customer quote in the launch post gives a result for the split. Hypha says that a separate harness and sandbox cut its failed agent responses by 86%. The vendor selected that testimonial. The mechanism is still real. A harness that does not share a process with rm -rf and an out-of-memory test suite has fewer ways to die.

Fig. 1 puts the pieces on one board, so you can see which jobs you outsource in each setup.

Fig. 1 · who runs what

Ten jobs every agent runtime does. Pick a setup to deal them into your column or the provider's, or tap a chip to move it yourself. The readout checks whether any shipping product actually offers the split you built.

your code0

OpenAI0

Assignments follow the Agents API architecture, self-hosted, lifecycle, and security docs as of September 12, 2026. "Your own loop" means your code calling the Responses API directly.

What a crash costs now

The best argument for a rented loop is durability. A local harness lives only as long as the process that runs it. If you kill the process mid-turn, you lose the in-flight model call. The tool call it waited on becomes an orphan. You also lose any part of the transcript that was not on disk. Every team that runs agents for hours writes a checkpoint layer for this problem.

In the Agents API, the session is the durable object and your process is a viewer. OpenAI keeps the session's configuration, turns, and items. If your application restarts, the turn keeps running. One limit is sharp: "Streams do not replay missed events." After a disconnect, retrieve the session and list its saved items to catch up. A reconnected stream does not deliver the backlog. The agent can also wait on one of your function tools while your process is down. In that case, the session stays in requires_action until something answers.

The environment can still die, and the lifecycle page says so directly. "An agent session can outlive its environment." A mid-turn disconnect "can fail a tool even if the turn completes." The disconnect does not restart a killed command. Later input can request a reconnection. The API then waits up to five minutes for the executor before it fails the submission. "Reusing the environment ID does not restore files in replacement compute." A deleted session "neither stops its environment nor emits a deletion webhook." You must find the bill for a forgotten sandbox yourself.

The durability that you rent stops at the stream. Two pieces of code are still yours. The first rebuilds your view after a disconnect. The docs give the procedure in five steps:

  1. Open a new stream and buffer it.
  2. While the stream stays open, fetch the session and its saved items.
  3. Restore state from the items.
  4. Drop updates for items that you already have.
  5. Resume.

The code is short.

def resume_view(client, session_id, handle_event, saved):
    with client.beta.agents.sessions.events.stream(session_id) as events:
        session = client.beta.agents.sessions.retrieve(session_id)
        for item in client.beta.agents.sessions.items.list(session_id, order="asc", limit=100):
            saved[item.id] = item
        for event in events:
            if getattr(event, "item_id", None) in saved:
                continue
            handle_event(event)
    return session

Our pattern, built from the documented recovery procedure in the events guide. OpenAI does not ship this sample. events.stream, sessions.retrieve, and items.list are the documented SDK calls. The skip rule is ours. Items saved mid-update can still need their final events.

The second is the handler that brings the machine back. Before the API waits for an offline executor, it emits agent.session.action_required. That event has required_action.type set to environment_connection. The API then holds the submission for up to five minutes. The docs split the work in two. The handler verifies and queues the event, and a worker checks the state again.

import json
import os
import queue

from fastapi import FastAPI, Request, Response
from openai import AsyncOpenAI, InvalidWebhookSignatureError

app = FastAPI()
client = AsyncOpenAI(webhook_secret=os.environ["OPENAI_WEBHOOK_SECRET"])
jobs = queue.Queue()


@app.post("/webhooks/openai")
async def handle_webhook(request: Request):
    payload = await request.body()
    try:
        client.webhooks.verify_signature(payload=payload, headers=request.headers)
    except (InvalidWebhookSignatureError, ValueError):
        return Response("Invalid signature", status_code=400)
    event = json.loads(payload)
    action = event["data"].get("required_action") or {}
    if event["type"] == "agent.session.failed" or action.get("type") == "environment_connection":
        jobs.put((event["type"], event["data"]["id"]))
    return Response(status_code=200)


async def work(session_id, start_executor, stop_compute):
    session = await client.beta.agents.sessions.retrieve(session_id)
    if session.status == "failed":
        return stop_compute(session_id)
    if any(a.type == "environment_connection" for a in session.required_actions or []):
        start_executor(session.environment.id, session.environment.remote_url)

Signature verification and the FastAPI shape come from the webhooks guide. The queue, the recheck, and the two callbacks follow the handler and worker jobs in the lifecycle guide. Your provider calls are start_executor and stop_compute. The code uses only documented session fields: status (failed, requires_action), required_actions entries with a type, and environment.id and environment.remote_url.

Fig. 2 · kill something mid-turn

One turn of a coding agent, seven steps. Choose where the harness runs and what dies while step 3, a long npm test, is still executing. The ledger reports what survives, following the documented behavior.

runtime

what dies at step 3

Behavior for the Agents API rows comes from the sessions, lifecycle, and manage-sessions guides. For hosted sandboxes, OpenAI manages recovery and publishes no contract for sandbox loss. The figure states this and does not guess.

The sandbox key and the other keys

A remote harness that drives a local executor needs a credential inside the box. OpenAI's design for that credential is good to copy, even if you never use the API. It uses two keys. Your application's OPENAI_API_KEY gets api.agents.read, api.agents.write, and api.responses.write. The docs say to keep it outside the sandbox. The executor gets a separate environment key as CODEX_API_KEY. You create that key on its own dashboard tab and set every other permission to None.

The security page says plainly that the key is visible. "Agent-generated code can read the environment key." The defense is scope: "This key only permits connecting environments. It cannot authorize any other API action." Assume that the agent can read everything in its box. Put nothing valuable in the box.

Third-party secrets follow the same rule. For OpenAI-hosted sandboxes, vaults store a credential outside the session. The sandbox code sees an environment variable that holds a placeholder. A network proxy puts in the real secret only for approved hosts. An MCP credential in a vault is bound to one server URL. OpenAI's service uses it, and code in the sandbox does not.

For self-hosted environments, the docs tell you to run that proxy yourself: "This is infrastructure you provide." You still own that proxy, and it sets the limit on the damage from a prompt injection. In the same week, an eval sandbox with real provider keys was the entry point for an attack on about 30 AI companies. I took that attack apart in the eval sandbox essay.

The table below gives the placement of each credential, so you can audit it.

CredentialScopesLives inIf the agent reads it
OPENAI_API_KEY (application)api.agents.read, api.agents.write, api.responses.write. Add api.vaults.* only if the app manages vaultsYour application, outside the sandboxThe agent must not read it. This key runs sessions and spends tokens
Environment key as CODEX_API_KEYConnects environments; every other permission set to NoneThe sandboxAssume that it does. The key cannot authorize any other API action
OPENAI_WEBHOOK_SECRETVerifies webhook signaturesThe webhook handlerIt must never reach the sandbox. Keep it apart from the executor key
Third-party token (GitHub, Stripe)Whatever the provider grantsA vault (hosted) or your proxy (self-hosted)It sees a placeholder that only works in an HTTPS header to an approved host

From the self-hosted, security, and webhooks guides, as of Sep 25, 2026.

For a hosted sandbox, the vault credential is one request. The sandbox reads $GITHUB_TOKEN and gets a placeholder. The proxy puts in the token only on HTTPS to api.github.com on port 443 or 8443.

POST /v1/vaults/{vault_id}/credentials

{
  "name": "GitHub API token",
  "auth": {
    "type": "environment_variable",
    "secret_name": "GITHUB_TOKEN",
    "secret_value": "<read from your secret manager, never a prompt or file>",
    "networking": {
      "type": "limited",
      "allowed_hosts": ["api.github.com"]
    }
  }
}

Body assembled from the field table in the vaults guide, as of Sep 25, 2026. The placeholder cannot sign anything locally. Put operations that need the real secret in a function tool that your application runs. To rotate the secret, start a new session. A running sandbox keeps the credential that it started with.

Fig. 3 · where each secret can be read

Three secrets an Agents API deployment handles. Choose the environment, then choose where each secret lives. Each row reports whether code the agent writes can read it, and what a stolen copy can do.

Scopes and placement rules from the Agents API self-hosted, security, and vaults guides. "Your proxy" is infrastructure that you run. The docs recommend it for self-hosted environments and do not ship one.

Same morning, different layer

Baseten sells inference. On the day OpenAI moved the harness into its own API, Baseten bought the layer underneath it. Blaxel built microVM sandboxes. In Baseten's words, they "suspend and resume in 25 milliseconds, up to 5x faster than other sandbox products." Blaxel also built Agent Drive, a distributed filesystem for artifacts. Idle sandboxes run "at close to zero cost." The press release names Abridge, Clay, Cursor, Lovable, Mercor, and OpenEvidence as customers. Terms were not disclosed. Baseten's stated thesis is locality: agents should "act next to the models they're calling, not across a network boundary."

The two moves point in opposite directions, and both make sense. OpenAI bets that the loop depends on the model, so the lab should own it. The launch post promises "versioned access" to harness improvements "with each model launch." Baseten bets that agents spend most of their wall-clock time on the computer, so the computer should sit next to the GPUs. For a harness that dials out, the sandbox vendors in the middle become interchangeable. E2B had already shipped templates on September 7 that run Cursor's self-hosted agent machines on E2B infrastructure. With the executor pattern, any harness can drive any sandbox.

A worked example: one flaky-test session

A CI job fails intermittently on a pull request. The agent must reproduce the failure, find the cause, and propose a patch without a push. The steps below use everything above, and I label the assumptions.

  1. Create. Your CI hook creates a self_hosted session with workspace_directory set to /workspace. Delegation has a cap of 3. One subagent can bisect commits while another reads the history of the test. Store the session ID, environment ID, and remote URL with the pull request.
  2. Connect. The agent.session.created webhook carries the environment ID and connect.remote_url. Your worker starts a container on your own runner pool and clones the branch. It then starts codex exec-server with the environment key. Your application opens the event stream and then sends the task.
  3. Lose the machine. After six minutes, the runner is preempted. The turn keeps running on OpenAI's side. The npm test in flight fails, and nothing restarts it. When the agent next needs the environment, the API emits environment_connection. Your handler starts a replacement inside the five-minute window. The replacement starts from a fresh clone, because a reused environment ID does not restore files. Data such as a dependency cache comes from the snapshot of your provider.
  4. Lose the viewer. Your dashboard process also restarts. It runs resume_view and rebuilds the transcript from saved items. Then it continues with live events.
  5. Finish. On agent.session.turn.completed, check the tool results and the final message as well as the event. Then retire the session and the container together.

You can calculate the model bill before the run. Assume that the turn reads 250,000 input tokens in total. Of those, 30,000 are written to the cache once, 20,000 are never cached, and 200,000 are cache reads. The turn writes 12,000 output tokens. The Astra essay gives the GPT-6 Astra list prices per million tokens:

The total is about $1.38 per flaky test, or $55 a day at 40 of them. The runner is your compute, so no container rate applies. The token counts are examples. The prices are list prices.

import time

import openai


def retire(client, session_id, stop_compute, attempts=5):
    for attempt in range(attempts):
        try:
            client.beta.agents.sessions.delete(session_id)
            break
        except openai.ConflictError:
            time.sleep(2 ** attempt)
    else:
        raise RuntimeError(f"session {session_id} still busy after {attempts} tries")
    stop_compute(session_id)

A pattern built from the hosted-sandbox and lifecycle guides, as of Sep 25, 2026. Deletion can return 409 while setup or execution finishes, so retry with a limit. Then stop your compute separately, because a deleted session does not stop it. openai.ConflictError is the 409 exception in the Python SDK.

When it breaks

SymptomWhyFix
The UI shows half a turn after a reconnectStreams do not replay missed eventsOpen a new stream, fetch the session and saved items, then apply live events (resume_view)
The session is idle but the work is wrongIdle means ready for input, not successWait for turn.completed, turn.failed, or turn.cancelled, then read tool results
Follow-up input hangs, then failsThe executor is offline. The API waits up to five minutesStart compute from the environment_connection webhook. Set client and proxy timeouts above five minutes
The replacement machine has no filesReusing the environment ID does not restore themClone on start and restore caches from provider snapshots
Sandbox bills keep arrivingDeleting a session does not stop compute and emits no webhookRetire session and compute in one job
Delete returns 409Setup or execution is still finishingBounded retry with backoff
A hosted request is blockedRedirect targets and subdomains need their own entriesList every exact host in allowed_domains
A signing step fails with a vault secretThe sandbox only ever holds a placeholderMove the operation into a function tool your application runs

Each row traces to the sessions, webhooks, lifecycle, hosted-sandbox, or vaults guide, as of Sep 25, 2026.

What you give up

A rented loop has costs that the quickstart does not show.

The API is still a beta. OpenAI says it will change quickly from feedback before general availability. Some of these costs will get smaller. The data-residency cost will stay, because a server-side harness must keep the transcript somewhere.

When to rent the loop

The decision is about which failures you prefer to own. Some agents run for hours, use subagents, and die in ways that you patch with checkpoints. For those agents, the Agents API gives you a durable session and a compaction strategy that someone else maintains. The executor lets you keep your own computer. Other agents are short, or their data cannot leave your boundary. Or you must route between labs on each turn. In those cases, keep the loop in your process and use this launch as a reference design.

In both cases, copy three things from it. Split the loop from the machine it drives, and make the machine dial out. Give the machine a key that can do exactly one thing. Keep every other secret behind a proxy that the agent's code cannot read. These rules also apply to a harness that you run on your laptop.

If you adopt it, do the steps in this order:

  1. Create two keys first. Give the application key only the three session scopes. Set every other permission on the environment key to None.
  2. Start with none or with a hosted sandbox that allows only exact hosts. Move to self_hosted when you need your own image, compute, or private network.
  3. Ship the webhook handler before the executor. The handler verifies and queues. The worker checks the session again and then starts compute.
  4. Write stream recovery before you build any interface on top of the stream.
  5. Put third-party secrets in vaults or behind your own proxy. Keep all signing operations behind a function tool.
  6. Retire sessions and compute in one job with a bounded retry.
  7. Record the model and the date on every eval run. The harness can change on a model launch without a deploy on your side.
rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. API behavior here comes from OpenAI's September 10 launch post and the Agents API guides. I used the overview, architecture, sessions, self-hosted sandboxes, sandbox lifecycle, security, and vaults guides. The Blaxel details come from Baseten's announcement and press release. I read all of them on September 12, 2026. The API is in public beta, so check the current docs before you build on a specific behavior.

Related: Harness Engineering · The Harness, Not the Model · More posts · X