Infrastructure & serving · Glossary term
What is Inter-Token Latency (ITL)?
The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token after the first.
Why does Inter-Token Latency (ITL) matter?
Individual gaps expose decode stalls and streaming jitter that a per-request average can hide, especially under batching, preemption, or mixed workloads.
Inter-Token Latency (ITL) in practice
Record each post-first-token interval with its request and token position, then report distributions by workload, output length, and concurrency without pooling away request boundaries.
What is the common confusion about Inter-Token Latency (ITL)?
ITL is one interval between consecutive tokens. Time per output token is a per-request average across those intervals, while time to first token covers the wait before streaming begins.
Learn Inter-Token Latency (ITL) in the course
Start with
- Inference Metrics — TTFT, TPOT, ITL, Goodput, P99
Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.
Taught in Phase 17: Infrastructure & Production.
Related terms
- Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.