Agents & tools · Glossary term

What is Checkpoint?

A durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references. In model training, it can store parameters, optimizer state, scheduler state, and the training position.

Why does Checkpoint matter?

Long-running workflows and training runs can recover from interruption without replaying completed work or losing expensive progress.

Checkpoint in practice

Save an agent's accepted patch and test evidence after a verified step, or save a training run's weights, optimizer state, random state, and data position before shutdown.

What is the common confusion about Checkpoint?

A workflow checkpoint and a model-training checkpoint serve the same recovery goal but preserve different state. Neither is merely a transcript or a weights file with no resume metadata.

Learn Checkpoint in the course

Start with

  • Checkpoint Save and Resume

    Train interrupts kill runs; checkpoints let them continue. Save model, optimizer, scheduler, loss history, step counter, and RNG state, atomically, so a kill at any moment leaves a valid file on disk.

    Phase 19: Capstone Projects

  • Repo Memory and Durable State

    Chat history is volatile. The repo is durable. The workbench stores agent state in versioned files so the next session, the next agent, and the next reviewer all read from the same source of truth.

    Phase 14: Agent Engineering

Lessons that name Checkpoint in a title or section

  • Agent State Machines — Graphs, Nodes, Checkpoints

    A ReAct loop written by hand is a while True. The same loop written as an explicit graph is something you can checkpoint, interrupt, branch, and time-travel through. The agent hasn't changed.

    Phase 11: LLM Engineering

  • Stateful Graph Orchestration — Durable Execution and Checkpoints

    Agent is a state machine; nodes are functions; edges are transitions; state is checkpointed after each node. Resume from any failure at the last successful checkpoint.

    Phase 14: Agent Engineering

  • Checkpoints and Rollback

    Every graph-state transition persists. When a worker crashes, its lease expires and another worker picks up at the latest checkpoint. Cloudflare Durable Objects hold state across hours or weeks.

    Phase 15: Autonomous Systems

  • Production Scaling — Queues, Checkpoints, Durability

    Scaling multi-agent systems to thousands of concurrent runs requires durable execution — work queues plus checkpoints, so any worker can resume any run after any crash, provided lease handling,…

    Phase 16: Multi-Agent & Swarms

  • Sharded Checkpoint and Atomic Resume

    A 70B-parameter training job is paused by a node failure every few hours. The checkpoint format decides whether you lose 30 minutes or 30 hours.

    Phase 19: Capstone Projects

  • Gradient Checkpointing and Activation Recomputation

    Backprop keeps every intermediate activation. At 70B parameters and 128K context that is 3 TB of activations per rank. Checkpointing trades FLOPs for memory: recompute instead of save.

    Phase 10: LLMs from Scratch

  • Show-o and Discrete-Diffusion Unified Models

    Transfusion mixes continuous and discrete representations. Show-o (Xie et al., August 2024) goes the other way: text tokens use causal next-token prediction, image tokens use masked discrete…

    Phase 12: Multimodal AI

  • Long-Running Background Agents: Durable Execution

    Production long-horizon agents do not run in while True. Every LLM call becomes an activity with checkpoint, retry, and replay. Temporal's OpenAI Agents SDK integration went GA March 2026.

    Phase 15: Autonomous Systems

Taught in Phase 14: Agent Engineering and Phase 19: Capstone Projects.

Also covered in Phase 10: LLMs from Scratch, Phase 11: LLM Engineering, Phase 12: Multimodal AI, Phase 15: Autonomous Systems and Phase 16: Multi-Agent & Swarms.

  • Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
  • Durable ExecutionRunning a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • Activation CheckpointingA training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
  • Agent MemoryInformation stored outside the model and selected for use in later agent steps, such as prior decisions, user preferences, task episodes,…
  • Compensating ActionA deliberate operation that semantically counteracts a completed side effect when the original operation cannot be rolled back atomically.
  • HandoffA structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints,…
  • IdempotencyThe property that repeating the same operation with the same identity does not create additional side effects beyond the first successful…
  • RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.

More terms in Agents & tools

Open the Agents & tools list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.