Infrastructure & serving · Glossary term

What is Prefix Caching?

Reusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix computation.

Why does Prefix Caching matter?

Shared system instructions, templates, or documents can consume substantial prefill work, but reuse only helps when token sequences and cache eligibility match.

Prefix Caching in practice

Place stable tokens before request-specific content, include model and tokenizer versions in cache identity, isolate tenant-sensitive state, monitor hit rate, and treat eviction as normal.

What is the common confusion about Prefix Caching?

Prefix caching reuses runtime attention state for exact token prefixes. Prompt caching is a broader provider or application contract, while semantic caching reuses a prior result for a similar request.

Learn Prefix Caching in the course

Start with

  • Inference Optimization

    Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.

    Phase 10: LLMs from Scratch

Lessons that name Prefix Caching in a title or section

  • Prompt Caching and Semantic Caching Economics

    Pricing snapshot dated 2026-04. Numeric claims below reflect vendor rate cards captured at this lesson's publication; verify against the linked docs before quoting them downstream.

    Phase 17: Infrastructure & Production

Taught in Phase 10: LLMs from Scratch.

Also covered in Phase 17: Infrastructure & Production.

  • Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
  • Semantic CacheA cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.