AI-native development · Glossary term
What is Observability?
The ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors, and evaluation signals.
Why does Observability matter?
AI failures often span model, retrieval, tools, and orchestration. You need correlated evidence to locate the failing boundary.
Observability in practice
Record a trace ID across retrieval, model calls, tool execution, approvals, and final scoring while applying redaction and access controls.
What is the common confusion about Observability?
Logging collects events. Observability makes those events structured and connected enough to answer operational questions.
Learn Observability in the course
Start with
- Agent Observability: Langfuse, Phoenix, Opik
Three open-source agent observability platforms dominate 2026. Langfuse (MIT) — 6M+ installs/month, tracing + prompt management + evals + session replay.
Lessons that name Observability in a title or section
- LLM Observability Stack Selection
The 2026 observability market splits into two categories. Development platforms (LangSmith, Langfuse, Comet Opik) bundle monitoring with evals, prompt management, session replays.
- Capstone 11 — LLM Observability & Eval Dashboard
Langfuse went open-core. Arize Phoenix published the 2026 GenAI semconv mappings. Helicone and Braintrust both doubled down on per-user cost attribution.
- Capstone Lesson 28: Observability with OTel GenAI Spans and Prometheus Metrics
An agent harness without observability is a black box that costs money. This lesson hand-rolls a span builder that emits records compliant with the OpenTelemetry GenAI semantic conventions, writes…
- Building a Production LLM Application
You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.
- Agent Framework Tradeoffs — Graph, Role, and Actor Orchestration
Every framework sells the same demo (research agent builds a report) and hides the same bug (state schema fights with the orchestration layer).
- The Actor Model for Agents — Async Messages and Typed Runtimes
Agents as actors: async message exchange, event-driven handlers, fault isolation, natural concurrency. AutoGen v0.4 (Microsoft Research, Jan 2025) redesigned agent orchestration around this model;…
- Production Runtimes: Queue, Event, Cron
Production agents run on six runtime shapes: request-response, streaming, durable execution, queue-based background, event-driven, and scheduled. Pick the shape before you pick the framework.
- AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost
A gateway sits between your apps and model providers. Core features are provider routing, fallback, retries, rate limiting, secret references, observability, guardrails.
Taught in Phase 14: Agent Engineering.
Also covered in Phase 11: LLM Engineering, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.
Related terms
- TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
- Canary ReleaseA deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
- Incident ResponseThe coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens…
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
- PostmortemA durable incident record that explains impact, detection, response, contributing conditions, recovery, and owned follow-up actions…
- SaturationThe degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
- Service Level Indicator (SLI)A quantitative measure of service behavior at a defined user-relevant boundary, such as successful request ratio or latency below a…
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
More terms in AI-native development
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.