Infrastructure & serving · Glossary term

What is Prefill?

The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.

Also called Prefill Phase.

Why does Prefill matter?

Prompt shape, queueing, and cache reuse affect prefill cost, and prefill competes differently for compute than decode, so it strongly influences startup latency and serving schedules.

Prefill in practice

Record prompt tokens and prefill latency, separate queue time from execution time, compare cached and uncached prefixes, and test long prompts beside active decode traffic.

What is the common confusion about Prefill?

Prefill is the runtime prompt-processing stage, not the first generated token itself. The first token appears only after prefill and any queueing complete.

Learn Prefill in the course

Start with

  • Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d

    Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…

    Phase 17: Infrastructure & Production

Lessons that name Prefill in a title or section

  • Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill

    Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…

    Phase 17: Infrastructure & Production

  • Pre-Training a Mini GPT (124M Parameters)

    GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.

    Phase 10: LLMs from Scratch

  • Inference Optimization

    Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.

    Phase 10: LLMs from Scratch

  • Prompt Engineering: Techniques & Patterns

    Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.

    Phase 11: LLM Engineering

  • GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling

    Three layers, not one. Karpenter provisions nodes dynamically (under one minute, 40% faster than Cluster Autoscaler). KAI Scheduler handles gang scheduling, topology awareness, and hierarchical…

    Phase 17: Infrastructure & Production

Taught in Phase 17: Infrastructure & Production.

Also covered in Phase 10: LLMs from Scratch and Phase 11: LLM Engineering.

  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
  • Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
  • Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.