GPT-6 Astra Changed the Shape of a Turn

If your agent runs on GPT-6 Astra, your tool loop loses three assumptions. The model is no longer idle while a tool runs. A user correction no longer has to wait for the turn to end. A change to reasoning effort is no longer free. Each change comes with an API, and each one moves state into your harness. You now keep a registry of pending jobs, a journal of mid-turn messages, and a log of the effort that applied to each turn.

On September 3, OpenAI shipped GPT-6 Astra. The same API changelog entry added three controls that break that rhythm. The model can call tools that it does not wait for. User messages can land in the middle of a running response. You can change the reasoning setting without paying for your prompt again. The launch post leads with benchmarks. The engineering changes are in the changelog, and they change what your harness has to track.


The contract Astra loosens

A classic function call is a blocking contract. The model emits a function_call item with a call_id, and the response ends. The turn cannot continue until your application sends back a function_call_output with the same call_id. Parallel tool calls softened this a little, because the model can ask for three lookups at once. It still waits for all of them before it does anything else. If one lookup takes 14 seconds, the model sits idle for 14 seconds. That is true even when half of the remaining work never needed that result.

Every harness I have taken apart on this site has this idle time. Codex and Claude Code both treat the tool call as a hard synchronization point. Most of the engineering around them (timeouts, background shells, subagents) works around that point from the outside. The Astra API moves one of those workarounds inside the protocol.

Async tools: the model keeps working

The mechanism is one field. Set async: true on a function or custom tool definition. When the model calls that tool, the call item carries async: true as well. The async tool calling guide describes the rest. The model can "continue working after issuing that call, before your application returns the output." One response can contain both the async call and an answer to an independent part of the request. When your job finishes, you send its output in a later request, matched by the original call_id. You chain that request with previous_response_id to the latest response at that time.

Two details decide if this is useful or dangerous. First, execution stays where it was: "Your application still executes the tool. Async tools don't move execution to OpenAI or manage your background jobs." You now own a registry of pending work that outlives a single request. Second, the model needs a way to say "I need that result now." OpenAI gives a pattern for this and no built-in. You define a synchronous wait_for_tasks tool and give every async tool a task_handle argument. Then you bind each handle to its call_id and running job.

When the model calls wait, return the finished results on their original call IDs first. Then return the wait status on the ID of the wait call itself. The results are then in context when the model resumes. The same shape gives you a tool that asks the user a question while the model keeps gathering facts. The figure below runs one incident-triage turn both ways. Drag the latencies and the amount of independent work, and watch where the savings come from.

Fig. 1 · one turn, two contracts

Prompt: "Find why checkout p99 regressed and draft the incident note." The model launches query_traces and fetch_deploy_log, has some work that does not depend on them (outlining the note, reviewing the PR already in context), then must wait before it analyzes. Durations are illustrative seconds.

14 s
6 s
9 s
Blocking call 0 s
model
query_traces
fetch_deploy_log
    async: true 0 s
    model
    query_traces
    fetch_deploy_log
      model working (numbered steps)your tool runningmodel blocked

       

      The arithmetic in that verdict is the whole case for and against the feature. With a blocking call, wall-clock time is the model work plus the slowest tool. With an async call, the slow tool overlaps with the work that does not depend on it. So the saving is the smaller of two numbers: the latency of the slowest tool, or the independent work available. A request with no independent work saves nothing. A request with a lot of it saves up to the full tool latency. Async tools are a scheduling feature, and the payoff is in proportion to how parallel your task already was.

      The guide also lists where async does not apply. It works for function and custom tools that your application runs. It does not work for hosted tools like web search. Do not mark tools async when you use programmatic tool calling. In multi-agent mode, do not combine async tools with parallel tool calls. Support also depends on the model: "supported by GPT-6 Astra and later models."

      Here is the tool pair from the guide, one async lookup and the synchronous wait tool. I trimmed only the descriptions:

      [
        {
          "type": "function",
          "name": "lookup_price",
          "async": true,
          "description": "Look up a product price in the background. Choose a fresh task_handle unique within this conversation, including completed tasks.",
          "strict": true,
          "parameters": {
            "type": "object",
            "properties": {
              "sku": { "type": "string" },
              "task_handle": { "type": "string" }
            },
            "required": ["sku", "task_handle"],
            "additionalProperties": false
          }
        },
        {
          "type": "function",
          "name": "wait_for_tasks",
          "description": "Wait for selected tasks whose results you need. Pass a nonempty list of distinct task_handles from your earlier async tool calls. Results arrive on their original calls; this tool returns status only.",
          "strict": true,
          "parameters": {
            "type": "object",
            "properties": {
              "task_handles": { "type": "array", "items": { "type": "string" } }
            },
            "required": ["task_handles"],
            "additionalProperties": false
          }
        }
      ]

      Tool definitions from the async tool calling guide. wait_for_tasks is your function and is not a built-in. If you omit async, the tool stays synchronous.

      The guide leaves the registry to you. This is the smallest version that follows its rules. Bind every handle to its original call_id and job. Never reuse a handle in the conversation. When the model waits, send finished results on their own call IDs before the wait status.

      import json
      from concurrent.futures import ThreadPoolExecutor
      
      pool = ThreadPoolExecutor(max_workers=8)
      registry = {}  # task_handle -> {"call_id": ..., "job": Future, "delivered": bool}
      
      def dispatch(item, run_tool):
          args = json.loads(item.arguments)
          if item.name == "wait_for_tasks":
              return wait(item, args["task_handles"])
          handle = args.get("task_handle")
          if item.async_:
              if handle in registry:
                  raise ValueError(f"task_handle reused: {handle}")
              registry[handle] = {"call_id": item.call_id,
                                  "job": pool.submit(run_tool, item.name, args),
                                  "delivered": False}
              return []  # nothing to send yet; the model keeps working
          return [{"type": "function_call_output", "call_id": item.call_id,
                   "output": json.dumps(run_tool(item.name, args))}]
      
      def wait(item, handles):
          outputs, done = [], []
          for h in handles:
              entry = registry[h]
              result = entry["job"].result()  # block only on what the model asked for
              if not entry["delivered"]:
                  outputs.append({"type": "function_call_output", "call_id": entry["call_id"],
                                  "output": json.dumps({"task_handle": h, **result})})
                  entry["delivered"] = True
              done.append(h)
          outputs.append({"type": "function_call_output", "call_id": item.call_id,
                          "output": json.dumps({"status": "completed",
                                                "completed_task_handles": done})})
          return outputs  # send as input with previous_response_id = latest response

      This is our pattern, built from the documented contract, and it is not OpenAI sample code. The ordering rule (results on original call IDs, then wait status) and the handle rules come from the guide. The Python SDK exposes the field as item.async_ in the OpenAI sample.

      Steering: a message that lands mid-response

      The second control is for the user who types while the agent works. Before Astra, a correction had two outcomes. You could cancel the response and lose its progress. Or you could wait for it to finish and then pay for a second turn that undoes part of the first. Mid-turn steering adds a third outcome. It works only over a WebSocket connection to the Responses API. GPT-5.6 and earlier models do not support it.

      The event flow is precise, so read it slowly. After the response.created event, you send response.steer with the response ID as previous_response_id. The server replies response.steer.accepted. That reply means the input is in the queue, and the model has not acted on it yet. Then the server "finishes the current output item and any hosted tool work already running." It ends the original response with response.incomplete and incomplete_details.reason: "steered". Then it creates a continuation response that includes your update. The continuation inherits the original request settings, and token and tool-call limits apply to each response separately.

      One exception causes most of the trouble. If the response ends and needs a client tool result or an approval, the steer stays in the queue. The server sends response.steer.pending with a required_input list that names the calls it needs. It then waits for you to return them with a normal response.create. The server prepends the queued steer for you, so if you send it again, it appears twice. Slide the arrival point below to see both paths.

      Fig. 2 · the steer injector

      A running response, resp_1, drafts an incident report. The user sends "keep it to the customer-facing impact" while it streams. Pick the output item that is in flight when the steer arrives.

      item 3
      output items
        WebSocket events

           

          That event log gives three rules. Steering never rewrites output that already reached your application, never undoes earlier actions, and never cancels a tool that already started. Queued steering lives on the connection and not on the stored response, so a dropped socket can lose it. The guide says to record every steer you send and compare it against response history before a replay. A failure is final. response.steer.failed means the input "will not apply it automatically later," with error codes such as steering_not_supported and too_many_pending_steers. If your harness sends steers and forgets them, it will sometimes drop the most important sentence the user wrote.

          On the wire, a steer is one event on the open socket. The only state to keep is its ID:

          {
            "type": "response.steer",
            "previous_response_id": "resp_1",
            "input": "Keep it to the customer-facing impact."
          }

          Event shape from the mid-turn steering guide. The input text is ours. The server answers response.steer.accepted with a steer.id.

          steers = {}  # steer.id -> {"input": ..., "state": "accepted"}
          
          def on_event(ev, send):
              kind = ev["type"]
              if kind == "response.steer.accepted":
                  steers[ev["steer"]["id"]] = {"input": pending_text.pop(0), "state": "accepted"}
              elif kind == "response.steer.pending":
                  steers[ev["steer"]["id"]]["state"] = "pending"
                  outputs = [run_client_tool(req["call_id"]) for req in ev["required_input"]]
                  # do not resend the steer: the server prepends it to this request
                  send({"type": "response.create", "model": "gpt-6-astra",
                        "previous_response_id": ev["steer"]["previous_response_id"],
                        "input": outputs})
              elif kind == "response.steer.failed":
                  steers[ev["steer"]["id"]]["state"] = "failed"
                  surface_to_user(ev["steer"]["input"], ev["error"])  # never silently dropped

          This is our journal around the documented events. The response.steer.pending event carries required_input, and you must not repeat the queued steer. A failed steer "will not apply it automatically later" (guide). pending_text, run_client_tool, and surface_to_user are yours.

          The launch post describes the same behavior at the product level. In Codex, Astra "can ask asynchronously while continuing work that doesn't depend on your reply." It proceeds on sensible assumptions for small gaps and waits on consequential decisions. Async tools and steering are the two halves of that. The model can ask without a stop, and you can answer without a restart.

          Effort became a message

          The third control looks small, and it is about money. Reasoning effort used to be a request-level setting, reasoning.effort. Astra accepts low, medium, high, xhigh, and max per its model page. Some conversations need high effort for one hard step and low effort for routine follow-ups. They used to change that field between requests. The OpenAI migration guidance now says to keep the request-level value fixed "to preserve the prompt prefix for caching." To change effort, you put a configuration_update input item before the next user message.

          response = client.responses.create(
              model="gpt-6-astra",
              reasoning={"effort": "low"},
              input="Draft a database migration plan.",
              store=True,
          )
          
          response = client.responses.create(
              model="gpt-6-astra",
              previous_response_id=response.id,
              reasoning={"effort": "low"},          # unchanged, so the prefix stays cached
              input=[
                  {"type": "configuration_update", "reasoning": {"effort": "high"}},
                  {"role": "user", "content": "Analyze the failure modes and propose rollback steps."},
              ],
              store=True,
          )

          From the reasoning guide example, with our comment. The request-level value stays low on both calls. Only the input item raises effort.

          The reasoning guide sets the rules. The update applies to the next response and to every one after it, until another update overrides it. The reasoning.effort field of the response keeps reporting the request-level value. Your logs will be wrong unless you record the updates yourself. Two updates cannot sit next to each other in history. Updates do not mix with automatic compaction or truncation, and the standalone compact endpoint rejects histories that contain them. The feature works in standard, single-agent mode only.

          The price sheet shows why this matters. Astra lists $10 per million input tokens, $1 per million cached input tokens, $12.50 per million for cache writes, and $50 per million output tokens. Prompts over 272K input tokens are billed at twice the input and cache rates. A cached read costs a tenth of a fresh one. An agent conversation is mostly prefix: system prompt, tool schemas, and every earlier turn. The figure prices eight turns both ways.

          Fig. 3 · the effort strip

          Eight turns of one agent session: 24,000 tokens of system prompt and tools, plus 6,000 new tokens per turn (illustrative). Tap a turn to switch its effort. Input uses the Astra list rates. Output is the same in both rows, so the figure leaves it out.

           

          Assumption, stated: a changed request-level reasoning.effort invalidates the cached prefix, so that turn writes the whole prefix again at the cache-write rate. The OpenAI guidance warns about this worst case. A configuration_update keeps the prefix, so only the new tokens are written.

          The pattern applies beyond Astra. Some things sit in the prefix and change between turns: a tool list you rebuild, a timestamp in the system prompt, a setting encoded early. Each one turns a cheap cached read into an expensive write. I made the same argument about the Anthropic cache pricing in Fable 5.1's two doors. OpenAI has now moved one of the most common per-turn changes out of the prefix and into the message stream.

          The monitor in the loop

          You cannot turn off the fourth change. OpenAI says Astra "meets the Critical threshold in cybersecurity under our Preparedness Framework," and it shipped with misalignment monitoring. The monitor "reviews model reasoning and actions asynchronously and can stop a conversation when it identifies a potential issue." It focuses on consequential contexts, such as transfer of or access to sensitive data, or destructive changes.

          Coverage depends on how you keep conversation state:

          Request typeMonitoredCan auto-stop
          Responses API with persisted reasoning, WebSockets, or OpenAI compactionyesyes
          Responses API using none of thoseyesno, webhook alerts only
          Chat Completionsnono, other checks apply

          A stop arrives as HTTP 403 with code misalignment_policy_violation. The guide is blunt about what that means: "The API does not provide a general way to resume a conversation stopped by misalignment monitoring." The check is asynchronous, so "an action may already have completed before monitoring identifies a concern. A stopped request does not undo earlier actions." Alerts can also reach you through a safety.alert.created webhook that carries only an alert ID. You fetch the alert with a key that has the api.safety.alerts.read permission.

          Put that next to async tools and the design constraint is clear. When a 403 arrives, your application can have three jobs in flight, a steer in the queue, and a user question pending. Nothing on the OpenAI side will unwind them. The launch post adds that in ChatGPT and Codex, a paused task can ask you to review the action. But "in the API, the task will stop." The same post says Astra will refuse more advanced security work, such as proof-of-concept exploits. OpenAI plans less restrictive access through its Daybreak program. If you build defensive tooling, plan for refusals on the public model and for stops you did not trigger.

          import openai
          
          try:
              response = client.responses.create(**request)
          except openai.PermissionDeniedError as err:          # HTTP 403
              if (err.body or {}).get("code") != "misalignment_policy_violation":
                  raise
              halt_dispatch(conversation_id)                   # no more tools, no retry
              record(conversation_id, request_id=err.request_id,
                     pending=[h for h, e in registry.items() if not e["delivered"]],
                     steers=steers)
              page_operator(conversation_id, reason=err.message)

          A handler written to the three steps in the misalignment monitoring guide: stop dispatching and never auto-retry, preserve IDs and tool records, show a human. Match the error code and ignore the message text. Streaming code must also watch for the error mid-stream. The helper names are ours.

          curl "https://api.openai.com/v1/safety/alerts/$ALERT_ID" \
            -H "Authorization: Bearer $OPENAI_API_KEY"

          Retrieving an alert after a safety.alert.created webhook, per the same guide. The key needs api.safety.alerts.read. ALERT_ID is the data.id of the webhook.

          The migration list, read as a harness spec

          The model guide has the usual migration checklist, and every line of it changes code somewhere. Astra has no none effort, so use low. It does not accept custom temperature, top_p, or log probabilities. It supports Chat Completions, but its tool calling requires the Responses API. Fast mode costs 2x, and Batch and Flex cost half. The window is 1,050,000 tokens with up to 128,000 output tokens.

            {
          -   "model": "gpt-5.6",
          +   "model": "gpt-6-astra",
          -   "reasoning": { "effort": "none" },
          +   "reasoning": { "effort": "low" },
          -   "temperature": 0.2,
          -   "top_p": 0.9,
          -   "top_logprobs": 5,
          -   "prompt_cache_retention": "...",
          +   "prompt_cache_options": { "ttl": "30m" },
              "tools": [ ...function tools, now on the Responses API... ],
              "input": [ ... ]
            }

          The request changes from the Astra migration quickstart. Use low because there is no none effort. Remove temperature, top_p, and top_logprobs when effort is not none. Move tool calling to Responses. From GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options.ttl set to "30m". The old retention value is elided.

          One line in the prompting section deserves more attention. The guide says Astra "can be more sensitive to instructions contained in skills and other files, such as AGENTS.md," and recommends an audit of them. This is the other side of better instruction following. It is the same supply-chain surface I wrote about in the AGENTS.md practices nobody uses. A model that obeys your files more closely also obeys a poisoned file more closely.

          If a skill causes you to ask for permission or confirmation, pause, leave requested work unfinished, or diverge from the user's intent, name and link to the exact SKILL.md file you read, quote the relevant instruction, and briefly explain how it applies. Distinguish explicit skill requirements from your interpretation of guidelines.

          Verbatim from the Astra prompting guidance. The guidance suggests it to find silent and conflicting guidance across skills and AGENTS.md files.

          When it goes wrong

          SymptomCauseFix
          400 on the first Astra requesttemperature, top_p, or top_logprobs sent with reasoning on, or effort noneRemove the sampling fields and use low effort
          The model resumes without the result it waited forWait status sent before, or instead of, the resultsSend each result on its original call_id, then the wait status
          The wrong job's result lands in contextA task_handle was reused after its task finishedKeep the registry for the whole conversation and reject reuse
          A user correction appears twiceThe steer was resent after response.steer.pendingReturn only the required tool results. The server prepends the steer
          A correction silently disappearsThe socket dropped with the steer still queued, or it failedJournal by steer.id, compare with history before replaying, surface failures
          Compaction request rejectedHistory contains configuration_update itemsCompact with a compaction_trigger item, then add a fresh update
          403 misalignment_policy_violationThe monitor stopped the conversationHalt dispatch, keep records, hand to a person, and do not retry

          Each row maps to a rule in the migration quickstart, the async, steering, reasoning, and monitoring guides.

          A worked example: one incident session

          Take the incident agent from Fig. 1 through a twelve-turn session. Five of the turns fan out to query_traces and fetch_deploy_log with the default latencies of the figure. That is 14 and 6 seconds of tool time against 9 seconds of work that does not need them. Effort runs low for four turns and high for two while the agent forms a hypothesis. Then it runs low for four more and high for the last two while the agent writes the rollback plan. The prompt is 24,000 tokens of system text and tool schemas, and each turn adds 6,000.

          The token counts and the assumption that a changed request-level effort rewrites the whole cached prefix are the stated worst case from Fig. 3. Prices are the OpenAI list rates for gpt-6-astra. Afterwards, check three things. Every async handle in the registry ended delivered. The steer journal has no accepted entry without a later response. The effort log matches the configuration_update items you sent, because the field of the response reports only the request-level value.

          An adoption order that follows the turn

          1. Make the request valid. Apply the migration diff and move tool calls to Responses before you touch anything else. A 400 is the cheapest failure you will ever get.
          2. Add the stop path. Wire the 403 handler first. Every later feature adds in-flight state that a stop leaves behind.
          3. Move effort into the message stream. Freeze request-level effort, send configuration_update items, and log them yourself.
          4. Mark one slow, idempotent tool async. Add the registry and wait_for_tasks. Then measure the saving against the independent work that is available. If the saving rounds to zero, stop.
          5. Turn on steering last. It needs the WebSocket transport, the steer journal, and the reconnect comparison. If any of the three is missing, this feature loses the words of the user and gives no sign.
          6. Audit your instruction files. Run the skills audit prompt over every SKILL.md and AGENTS.md that the agent loads. Astra obeys them more closely.

          The vendor numbers in the launch post are strong. OpenAI reports 57.9% on Terminal-Bench 4.0, against 55.8% for Claude Fable 5.1 and 37.3% for GPT-5.6. It also says Astra often finishes tasks with fewer output tokens than the models it compares against. Treat those as OpenAI measurements until independent ones arrive. The API changes need no benchmark. Before Astra, a turn was a straight line with pauses in it. With Astra, a turn is a set of overlapping timelines. They are model work, your jobs, the corrections of the user, and a monitor that watches all three. You still have to write the harness that keeps those timelines straight.

          rg
          Rohit Ghumare

          CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals under the AI stack. API behavior here comes from the OpenAI GPT-6 Astra launch post and the API changelog entry of September 3. It also comes from the async tool calling, mid-turn steering, reasoning, model, and misalignment monitoring guides, read in September 2026. Benchmark figures are from OpenAI. Figure durations and token counts are illustrative. Prices are the OpenAI list rates for gpt-6-astra.

          Related: Inside the Codex Harness · Harness Engineering · More posts · X