Infrastructure & serving · Glossary term

What is Decode Phase?

The iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.

Why does Decode Phase matter?

Decode work has different compute, memory, and scheduling behavior from prefill, so one aggregate latency number can hide the actual serving bottleneck.

Decode Phase in practice

Measure inter-token latency and output throughput separately, account for KV-cache occupancy, and test mixed workloads where active decodes share capacity with new prefills.

What is the common confusion about Decode Phase?

Decode phase is not the decoder component of an encoder-decoder model. It names the runtime generation stage.

Learn Decode Phase in the course

Start with

  • Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d

    Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…

    Phase 17: Infrastructure & Production

Taught in Phase 17: Infrastructure & Production.

  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
  • Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
  • Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
  • Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
  • Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.