Infrastructure & serving · Glossary term

What is Chunked Prefill?

A serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with decode work from other requests.

Why does Chunked Prefill matter?

One long prompt can otherwise occupy the accelerator and delay active generations, producing poor tail latency even when total throughput looks healthy.

Chunked Prefill in practice

Choose a chunk policy from measured workloads, account for scheduling overhead, and compare prefill completion, decode latency, and goodput under mixed prompt lengths.

What is the common confusion about Chunked Prefill?

Chunked prefill changes how prompt computation is scheduled. It does not split the user's context into independent semantic chunks or change the model's context window.

Learn Chunked Prefill in the course

Start with

Taught in Phase 17: Infrastructure & Production.

  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • Dynamic BatchingA runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed…
  • Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.