Models & inference · Glossary term

What is Time to First Token (TTFT)?

The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined measurement boundary.

Also called TTFT.

Why does Time to First Token (TTFT) matter?

TTFT strongly affects perceived responsiveness and can reveal queueing, prompt processing, cache, or network delays.

Time to First Token (TTFT) in practice

Record client-side TTFT by model, prompt length, region, and cache status, then separate it from total completion time.

What is the common confusion about Time to First Token (TTFT)?

TTFT is not tokens per second. One measures startup latency; the other measures generation throughput after output begins.

Learn Time to First Token (TTFT) in the course

Lessons that name Time to First Token (TTFT) in a title or section

Covered in Phase 17: Infrastructure & Production.

  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
  • GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
  • Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…
  • Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
  • Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
  • TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.