Reliability & operations · Glossary term

What is Tail Latency?

The latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time window.

Why does Tail Latency matter?

Averages can look healthy while a meaningful group of users waits much longer because of queueing, contention, retries, or variable request cost.

Tail Latency in practice

Report several percentiles by route and workload, retain timeouts as censored or failed observations according to a documented rule, and trace slow requests across dependencies.

What is the common confusion about Tail Latency?

Tail latency is not the single slowest request and has no meaning without the percentile, population, and measurement boundary.

Learn Tail Latency in the course

Start with

  • Inference Metrics — TTFT, TPOT, ITL, Goodput, P99

    Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.

    Phase 17: Infrastructure & Production

Taught in Phase 17: Infrastructure & Production.

  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
  • SaturationThe degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
  • GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
  • Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
  • Deadline PropagationPassing the remaining end-to-end time budget to downstream calls so each dependency knows how long the original request can still usefully…
  • Dynamic BatchingA runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed…
  • Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
  • Service Level Indicator (SLI)A quantitative measure of service behavior at a defined user-relevant boundary, such as successful request ratio or latency below a…

Sources

More terms in Reliability & operations

Open the Reliability & operations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.