Infrastructure & serving · Glossary term
What is Prefill?
The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.
Also called Prefill Phase.
Why does Prefill matter?
Prompt shape, queueing, and cache reuse affect prefill cost, and prefill competes differently for compute than decode, so it strongly influences startup latency and serving schedules.
Prefill in practice
Record prompt tokens and prefill latency, separate queue time from execution time, compare cached and uncached prefixes, and test long prompts beside active decode traffic.
What is the common confusion about Prefill?
Prefill is the runtime prompt-processing stage, not the first generated token itself. The first token appears only after prefill and any queueing complete.
Learn Prefill in the course
Start with
- Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…
Lessons that name Prefill in a title or section
- Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…
- Pre-Training a Mini GPT (124M Parameters)
GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.
- Inference Optimization
Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.
- Prompt Engineering: Techniques & Patterns
Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.
- GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling
Three layers, not one. Karpenter provisions nodes dynamically (under one minute, 40% faster than Cluster Autoscaler). KAI Scheduler handles gang scheduling, topology awareness, and hierarchical…
Taught in Phase 17: Infrastructure & Production.
Also covered in Phase 10: LLMs from Scratch and Phase 11: LLM Engineering.
Related terms
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
- Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
- Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.