AI-native development · Glossary term
What is Trace?
A correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
Why does Trace matter?
It lets you reconstruct where time, cost, and failure entered a multi-step workflow.
Trace in practice
Propagate one trace identifier through the agent harness and attach redacted spans for each model and tool operation.
What is the common confusion about Trace?
A trace should record operational evidence, not expose hidden model reasoning, secrets, or unredacted sensitive content.
Learn Trace in the course
Start with
- OpenTelemetry GenAI Semantic Conventions
OpenTelemetry's GenAI SIG (launched April 2024) defines the standard schema for agent telemetry. Span names, attributes, and content-capture rules converge across vendors so agent traces mean the…
Lessons that name Trace in a title or section
- Capstone: Stateless Tool Ecosystem
A production agent system is a set of boundaries, not a pile of features. This capstone separates a readable in-process simulation from the protocol clients, authorization server, sandbox, and…
- The Harness as a Library — Subagents and Session Store
A harness you can import: built-in tools, subagents for context isolation, hooks, W3C trace propagation, session persistence.
- FinOps for LLMs — Unit Economics and Multi-Tenant Attribution
Traditional FinOps breaks on LLM spend. Costs are token-transactions, not resource-uptime. Tags don't map — an API call is a transaction, not an asset.
Taught in Phase 14: Agent Engineering.
Also covered in Phase 13: Tools & Protocols and Phase 17: Infrastructure & Production.
Related terms
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
- Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
Sources
More terms in AI-native development
- Backpressure
- Circuit Breaker
- Coding Agent
- Cost per Successful Task
- Flaky Test
- Handoff
- Idempotency
- Model Router
- Observability
- Patch
- Progressive Disclosure
- Rate Limit
- Regression Test
- Repository Instructions
- Repository Map
- Reproducible Build
- Retry with Backoff
- Reviewer Agent
- Scope Contract
- Semantic Cache
- Test Oracle
- Worktree
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.