Prompting & context · Glossary term

What is Prompt Cache?

Reuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some preprocessing work.

Why does Prompt Cache matter?

Stable instructions and large shared documents can become cheaper or faster across repeated calls when the provider's cache contract is satisfied.

Prompt Cache in practice

Place stable policy text before request-specific content, monitor cache-hit metadata, and treat misses as normal because eligibility and lifetime vary by provider.

What is the common confusion about Prompt Cache?

A prompt cache is a provider or application reuse contract and may use prefix caching internally. Prefix caching specifically reuses eligible exact-token KV state, while semantic caching reuses a prior result for a sufficiently similar request.

Learn Prompt Cache in the course

Start with

  • Prompt Caching and Context Caching

    Your system prompt is 4,000 tokens. Your RAG context is 20,000 tokens. You send both with every request. You also pay for both — every time.

    Phase 11: LLM Engineering

Taught in Phase 11: LLM Engineering.

  • Semantic CacheA cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
  • Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.

More terms in Prompting & context

Open the Prompting & context list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.