Infrastructure & serving · Glossary term

What is Disaggregated Serving?

A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.

Why does Disaggregated Serving matter?

Prefill and decode stress hardware differently, so independent pools can be sized and scheduled for their own bottlenecks instead of competing in one queue.

Disaggregated Serving in practice

Measure state-transfer cost, route requests through compatible model versions, scale each pool from its own demand signal, and test failure recovery between phases.

What is the common confusion about Disaggregated Serving?

Disaggregation separates runtime stages. It does not split one model into tensor or pipeline-parallel shards within a stage.

Learn Disaggregated Serving in the course

Start with

  • Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d

    Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…

    Phase 17: Infrastructure & Production

Lessons that name Disaggregated Serving in a title or section

Taught in Phase 17: Infrastructure & Production.

  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
  • GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.