Reliability & operations · Glossary term

What is Rollback?

Restoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.

Why does Rollback matter?

Agent and model changes can fail in production despite pre-deployment evaluation, so recovery must be designed before rollout.

Rollback in practice

Retain versioned artifacts and configuration, define rollback triggers, rehearse the command and data implications, and verify service health after restoration.

What is the common confusion about Rollback?

Code rollback does not automatically reverse database migrations, external side effects, cached outputs, or data written by the bad release.

Learn Rollback in the course

Lessons that name Rollback in a title or section

  • MCP Registry Supply Chain: Admission, Drift, and Rollback

    A registry entry tells you what a publisher declared. Production admission proves what you fetched, what you observed, what you approved, and what you can safely restore.

    Phase 13: Tools & Protocols

  • Checkpoints and Rollback

    Every graph-state transition persists. When a worker crashes, its lease expires and another worker picks up at the latest checkpoint. Cloudflare Durable Objects hold state across hours or weeks.

    Phase 15: Autonomous Systems

  • Building a Complete LLM Pipeline

    Everything from Lessons 01 to 12 is one stage of one pipeline. This lesson is the scaffold that turns those stages into a single end-to-end run: tokenize, pre-train, scale, SFT, align, evaluate,…

    Phase 10: LLMs from Scratch

  • Speculative Decoding and EAGLE-3

    Phase 7 · Lesson 16 proved the math: the Leviathan rejection rule preserves the verifier's distribution exactly. This lesson is the training-stack view of 2026 production speculative decoding.

    Phase 10: LLMs from Scratch

  • Scope Contracts and Task Boundaries

    The model does not know where the work ends. A scope contract is a per-task file that says where the work begins, where it ends, and how to roll back if it spills.

    Phase 14: Agent Engineering

  • Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs

    LLM rollouts combine the hardest parts of software deployment: no unit tests, diffuse failure modes, delayed signals. The sequence is (1) shadow mode — duplicate prod requests to candidate model,…

    Phase 17: Infrastructure & Production

Covered in Phase 10: LLMs from Scratch, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems and Phase 17: Infrastructure & Production.

  • Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
  • Durable ExecutionRunning a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.

Sources

More terms in Reliability & operations

Open the Reliability & operations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.