Agents & tools · Glossary term
What is Checkpoint?
A durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references. In model training, it can store parameters, optimizer state, scheduler state, and the training position.
Why does Checkpoint matter?
Long-running workflows and training runs can recover from interruption without replaying completed work or losing expensive progress.
Checkpoint in practice
Save an agent's accepted patch and test evidence after a verified step, or save a training run's weights, optimizer state, random state, and data position before shutdown.
What is the common confusion about Checkpoint?
A workflow checkpoint and a model-training checkpoint serve the same recovery goal but preserve different state. Neither is merely a transcript or a weights file with no resume metadata.
Learn Checkpoint in the course
Start with
- Checkpoint Save and Resume
Train interrupts kill runs; checkpoints let them continue. Save model, optimizer, scheduler, loss history, step counter, and RNG state, atomically, so a kill at any moment leaves a valid file on disk.
- Repo Memory and Durable State
Chat history is volatile. The repo is durable. The workbench stores agent state in versioned files so the next session, the next agent, and the next reviewer all read from the same source of truth.
Lessons that name Checkpoint in a title or section
- Agent State Machines — Graphs, Nodes, Checkpoints
A ReAct loop written by hand is a while True. The same loop written as an explicit graph is something you can checkpoint, interrupt, branch, and time-travel through. The agent hasn't changed.
- Stateful Graph Orchestration — Durable Execution and Checkpoints
Agent is a state machine; nodes are functions; edges are transitions; state is checkpointed after each node. Resume from any failure at the last successful checkpoint.
- Checkpoints and Rollback
Every graph-state transition persists. When a worker crashes, its lease expires and another worker picks up at the latest checkpoint. Cloudflare Durable Objects hold state across hours or weeks.
- Production Scaling — Queues, Checkpoints, Durability
Scaling multi-agent systems to thousands of concurrent runs requires durable execution — work queues plus checkpoints, so any worker can resume any run after any crash, provided lease handling,…
- Sharded Checkpoint and Atomic Resume
A 70B-parameter training job is paused by a node failure every few hours. The checkpoint format decides whether you lose 30 minutes or 30 hours.
- Gradient Checkpointing and Activation Recomputation
Backprop keeps every intermediate activation. At 70B parameters and 128K context that is 3 TB of activations per rank. Checkpointing trades FLOPs for memory: recompute instead of save.
- Show-o and Discrete-Diffusion Unified Models
Transfusion mixes continuous and discrete representations. Show-o (Xie et al., August 2024) goes the other way: text tokens use causal next-token prediction, image tokens use masked discrete…
- Long-Running Background Agents: Durable Execution
Production long-horizon agents do not run in while True. Every LLM call becomes an activity with checkpoint, retry, and replay. Temporal's OpenAI Agents SDK integration went GA March 2026.
Taught in Phase 14: Agent Engineering and Phase 19: Capstone Projects.
Also covered in Phase 10: LLMs from Scratch, Phase 11: LLM Engineering, Phase 12: Multimodal AI, Phase 15: Autonomous Systems and Phase 16: Multi-Agent & Swarms.
Related terms
- Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
- Durable ExecutionRunning a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- Activation CheckpointingA training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
- Agent MemoryInformation stored outside the model and selected for use in later agent steps, such as prior decisions, user preferences, task episodes,…
- Compensating ActionA deliberate operation that semantically counteracts a completed side effect when the original operation cannot be rolled back atomically.
- HandoffA structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints,…
- IdempotencyThe property that repeating the same operation with the same identity does not create additional side effects beyond the first successful…
- RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
More terms in Agents & tools
- Agent
- Agent Harness
- Agent Memory
- Agent State
- Agent Skill
- Approval Gate
- Compensating Action
- Delegation
- Durable Execution
- Function Calling
- Human-in-the-Loop (HITL)
- MCP (Model Context Protocol)
- Multi Round-Trip Request (MRTR)
- Orchestration
- Planning
- ReAct
- Sandbox
- Skill Bundle
- Skill Catalog
- Skill Discovery
- Skill Invocation
- Stateless MCP
- Structured Output
- Swarm
- Termination Condition
- Tool Contract
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.