Agents & tools · Glossary term
What is Durable Execution?
Running a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
Why does Durable Execution matter?
Agent tasks often span model calls, tools, approvals, and external systems. A transient process should not be the only record of progress.
Durable Execution in practice
Persist each workflow transition, use idempotency keys for external writes, and resume from the latest checkpoint after a worker restarts.
What is the common confusion about Durable Execution?
Durable execution does not make every operation safe automatically. Side effects still need idempotency and compensation rules.
Learn Durable Execution in the course
Lessons that name Durable Execution in a title or section
- Stateful Graph Orchestration — Durable Execution and Checkpoints
Agent is a state machine; nodes are functions; edges are transitions; state is checkpointed after each node. Resume from any failure at the last successful checkpoint.
- Long-Running Background Agents: Durable Execution
Production long-horizon agents do not run in while True. Every LLM call becomes an activity with checkpoint, retry, and replay. Temporal's OpenAI Agents SDK integration went GA March 2026.
- Production Runtimes: Queue, Event, Cron
Production agents run on six runtime shapes: request-response, streaming, durable execution, queue-based background, event-driven, and scheduled. Pick the shape before you pick the framework.
- Production Scaling — Queues, Checkpoints, Durability
Scaling multi-agent systems to thousands of concurrent runs requires durable execution — work queues plus checkpoints, so any worker can resume any run after any crash, provided lease handling,…
Covered in Phase 14: Agent Engineering, Phase 15: Autonomous Systems and Phase 16: Multi-Agent & Swarms.
Related terms
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
- IdempotencyThe property that repeating the same operation with the same identity does not create additional side effects beyond the first successful…
- Approval GateA control point that blocks a consequential action until an authorized person or policy grants permission.
- Compensating ActionA deliberate operation that semantically counteracts a completed side effect when the original operation cannot be rolled back atomically.
- OrchestrationThe control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
- RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
More terms in Agents & tools
- Agent
- Agent Harness
- Agent Memory
- Agent State
- Agent Skill
- Approval Gate
- Checkpoint
- Compensating Action
- Delegation
- Function Calling
- Human-in-the-Loop (HITL)
- MCP (Model Context Protocol)
- Multi Round-Trip Request (MRTR)
- Orchestration
- Planning
- ReAct
- Sandbox
- Skill Bundle
- Skill Catalog
- Skill Discovery
- Skill Invocation
- Stateless MCP
- Structured Output
- Swarm
- Termination Condition
- Tool Contract
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.