Infrastructure & serving · Glossary term

What is Tokens per Second (TPS)?

A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.

Also called TPS and output token throughput.

Why does Tokens per Second (TPS) matter?

It complements startup latency by showing how quickly generation proceeds after output begins and how serving behaves under load.

Tokens per Second (TPS) in practice

State whether TPS is per request or aggregate, exclude or identify prefill, and report batch, concurrency, sequence lengths, hardware, and percentile latency.

What is the common confusion about Tokens per Second (TPS)?

TPS is not directly comparable across different tokenizers, workloads, quality settings, or measurement boundaries.

Learn Tokens per Second (TPS) in the course

No lesson links to this term yet. Search the course catalog for it.

  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.