Reliability & operations · Glossary term

What is Incident Response?

The coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens service, data, safety, or security.

Why does Incident Response matter?

During an incident, clear roles and evidence matter more than improvised heroics, especially when model behavior and distributed dependencies obscure the failing boundary.

Incident Response in practice

Define severity and command roles, preserve traces and audit records, stop harmful actions, communicate impact, verify recovery, and track corrective work to completion.

What is the common confusion about Incident Response?

Incident response manages the event and its consequences. Root-cause analysis and long-term prevention continue after immediate service is restored.

Learn Incident Response in the course

Start with

Taught in Phase 17: Infrastructure & Production.

  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
  • PostmortemA durable incident record that explains impact, detection, response, contributing conditions, recovery, and owned follow-up actions…
  • AvailabilityThe proportion of eligible service interactions or time windows in which users can obtain the defined acceptable service under a stated…
  • Error BudgetThe amount of unsuccessful service allowed by a service-level objective over its measurement window before the objective is exhausted.

Sources

More terms in Reliability & operations

Open the Reliability & operations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.