Models & inference · Glossary term

What is KV Cache?

Stored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections for the unchanged prefix at every decoding step.

What people say

“A cache that makes token generation faster.”

Why does KV Cache matter?

It reduces repeated computation but consumes memory that grows with sequence length, layers, batch, and model configuration.

What is the common confusion about KV Cache?

A KV cache is runtime attention state for a sequence. Prefix caching reuses eligible KV state across requests, while prompt caching is a broader provider or application reuse contract.

Learn KV Cache in the course

Start with

  • KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Phase 07: Transformers Deep Dive

Lessons that name KV Cache in a title or section

  • Multi-Region LLM Serving and KV Cache Locality

    Round-robin load balancing is actively harmful for cached LLM inference. A request that does not land on the node holding its prefix pays full prefill cost — roughly 800 ms at P50 on a long prompt…

    Phase 17: Infrastructure & Production

  • Attention Variants — Sliding Window, Sparse, Differential

    Full attention is a circle. Every token sees every token, and memory pays the price. Four variants bend the shape of the circle and recover half the cost.

    Phase 07: Transformers Deep Dive

  • Speculative Decoding — Draft, Verify, Repeat

    Autoregressive decoding is serial. Each token waits for the previous one. Speculative decoding breaks the chain: a cheap model drafts N tokens, the expensive model verifies all N in one forward pass.

    Phase 07: Transformers Deep Dive

  • Pre-Training a Mini GPT (124M Parameters)

    GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.

    Phase 10: LLMs from Scratch

  • Inference Optimization

    Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.

    Phase 10: LLMs from Scratch

  • Open Models: Architecture Walkthroughs

    You built a GPT-2 Small from scratch in Lesson 04. Frontier open models in 2026 are the same family with five or six concrete changes. RMSNorm instead of LayerNorm. SwiGLU instead of GELU.

    Phase 10: LLMs from Scratch

  • Speculative Decoding and EAGLE-3

    Phase 7 · Lesson 16 proved the math: the Leviathan rejection rule preserves the verifier's distribution exactly. This lesson is the training-stack view of 2026 production speculative decoding.

    Phase 10: LLMs from Scratch

  • Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell

    Hardware-specialized inference compilation trades portability for throughput, and TensorRT-LLM — NVIDIA-only, tuned for Blackwell — is the clearest example of the trade paying off.

    Phase 17: Infrastructure & Production

Taught in Phase 07: Transformers Deep Dive.

Also covered in Phase 10: LLMs from Scratch and Phase 17: Infrastructure & Production.

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
  • Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…
  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.