Infrastructure & serving · Glossary term

What is Time per Output Token (TPOT)?

For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`. System distributions then aggregate those per-request averages.

Why does Time per Output Token (TPOT) matter?

Users can receive the first token quickly while the rest of the answer streams slowly, so startup latency alone does not describe generation responsiveness.

Time per Output Token (TPOT) in practice

Compute TPOT separately for each request, report percentiles across requests by output length and concurrency, and avoid pooling all token intervals or comparing systems with different tokenizers and measurement boundaries.

What is the common confusion about Time per Output Token (TPOT)?

TPOT is a per-request average. An individual inter-token latency is one gap between consecutive tokens, while time to first token includes the wait before output starts.

Learn Time per Output Token (TPOT) in the course

Start with

  • Inference Metrics — TTFT, TPOT, ITL, Goodput, P99

    Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.

    Phase 17: Infrastructure & Production

Taught in Phase 17: Infrastructure & Production.

  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
  • Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
  • Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.