Models & inference · Glossary term
What is Time to First Token (TTFT)?
The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined measurement boundary.
Also called TTFT.
Why does Time to First Token (TTFT) matter?
TTFT strongly affects perceived responsiveness and can reveal queueing, prompt processing, cache, or network delays.
Time to First Token (TTFT) in practice
Record client-side TTFT by model, prompt length, region, and cache status, then separate it from total completion time.
What is the common confusion about Time to First Token (TTFT)?
TTFT is not tokens per second. One measures startup latency; the other measures generation throughput after output begins.
Learn Time to First Token (TTFT) in the course
Lessons that name Time to First Token (TTFT) in a title or section
- Inference Metrics — TTFT, TPOT, ITL, Goodput, P99
Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.
- Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…
Covered in Phase 17: Infrastructure & Production.
Related terms
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
- GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
- Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
- PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
- Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…
- Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
- Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
- TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.