Infrastructure & serving · Glossary term
What is Disaggregated Serving?
A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.
Why does Disaggregated Serving matter?
Prefill and decode stress hardware differently, so independent pools can be sized and scheduled for their own bottlenecks instead of competing in one queue.
Disaggregated Serving in practice
Measure state-transfer cost, route requests through compatible model versions, scale each pool from its own demand signal, and test failure recovery between phases.
What is the common confusion about Disaggregated Serving?
Disaggregation separates runtime stages. It does not split one model into tensor or pipeline-parallel shards within a stage.
Learn Disaggregated Serving in the course
Start with
- Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…
Lessons that name Disaggregated Serving in a title or section
- Production Serving Stack — KV Offloading and Cache-Aware Routing
A production serving stack wires router, engines, and observability into one Kubernetes deployment — and treats KV cache as a resource that can leave the GPU.
Taught in Phase 17: Infrastructure & Production.
Related terms
- PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
- GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.