AI-native development · Glossary term

What is Observability?

The ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors, and evaluation signals.

Why does Observability matter?

AI failures often span model, retrieval, tools, and orchestration. You need correlated evidence to locate the failing boundary.

Observability in practice

Record a trace ID across retrieval, model calls, tool execution, approvals, and final scoring while applying redaction and access controls.

What is the common confusion about Observability?

Logging collects events. Observability makes those events structured and connected enough to answer operational questions.

Learn Observability in the course

Start with

  • Agent Observability: Langfuse, Phoenix, Opik

    Three open-source agent observability platforms dominate 2026. Langfuse (MIT) — 6M+ installs/month, tracing + prompt management + evals + session replay.

    Phase 14: Agent Engineering

Lessons that name Observability in a title or section

Taught in Phase 14: Agent Engineering.

Also covered in Phase 11: LLM Engineering, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.

  • TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
  • Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
  • Incident ResponseThe coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens…
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
  • PostmortemA durable incident record that explains impact, detection, response, contributing conditions, recovery, and owned follow-up actions…
  • SaturationThe degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
  • Service Level Indicator (SLI)A quantitative measure of service behavior at a defined user-relevant boundary, such as successful request ratio or latency below a…
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.

More terms in AI-native development

Open the AI-native development list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.