Infrastructure & serving · Glossary term

What is Paged KV Cache?

A KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead of requiring one contiguous allocation per sequence.

Why does Paged KV Cache matter?

Variable sequence lengths create fragmentation and unpredictable growth, so block-based allocation can improve usable memory and enable flexible sharing.

Paged KV Cache in practice

Select block size from workload measurements, track allocation and eviction, isolate state between requests, and test cancellation and prefix sharing under memory pressure.

What is the common confusion about Paged KV Cache?

Paged KV cache manages runtime attention-state memory. It does not move model parameters to disk or extend the model's trained context limit.

Learn Paged KV Cache in the course

Start with

Taught in Phase 17: Infrastructure & Production.

  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
  • Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.