Infrastructure & serving · Glossary term
What is Paged KV Cache?
A KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead of requiring one contiguous allocation per sequence.
Why does Paged KV Cache matter?
Variable sequence lengths create fragmentation and unpredictable growth, so block-based allocation can improve usable memory and enable flexible sharing.
Paged KV Cache in practice
Select block size from workload measurements, track allocation and eviction, isolate state between requests, and test cancellation and prefix sharing under memory pressure.
What is the common confusion about Paged KV Cache?
Paged KV cache manages runtime attention-state memory. It does not move model parameters to disk or extend the model's trained context limit.
Learn Paged KV Cache in the course
Start with
- Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…
Taught in Phase 17: Infrastructure & Production.
Related terms
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
- Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.