Models & inference · Glossary term

What is Speculative Decoding?

An inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel. In exact sampling variants, an acceptance and correction rule preserves the target model's output distribution.

Why does Speculative Decoding matter?

It can reduce serial target-model decoding work when drafts are accepted, without requiring a change to the target model's trained weights.

Speculative Decoding in practice

Measure acceptance rate and end-to-end latency on real prompts, include draft-model overhead, and verify that the implementation preserves the intended decoding distribution.

What is the common confusion about Speculative Decoding?

Speculative decoding is not ordinary model routing or unverified autocomplete. Exact variants preserve the target distribution through acceptance and correction, while approximate variants may trade that guarantee for speed.

Learn Speculative Decoding in the course

Lessons that name Speculative Decoding in a title or section

  • Speculative Decoding — Draft, Verify, Repeat

    Autoregressive decoding is serial. Each token waits for the previous one. Speculative decoding breaks the chain: a cheap model drafts N tokens, the expensive model verifies all N in one forward pass.

    Phase 07: Transformers Deep Dive

  • Speculative Decoding and EAGLE-3

    Phase 7 · Lesson 16 proved the math: the Leviathan rejection rule preserves the verifier's distribution exactly. This lesson is the training-stack view of 2026 production speculative decoding.

    Phase 10: LLMs from Scratch

  • Speculative Decoding and EAGLE

    A frontier LLM generating one token requires a full forward pass over billions of parameters. That forward pass is massively over-provisioned: most of the time a much smaller model can guess the…

    Phase 10: LLMs from Scratch

  • EAGLE-3 Speculative Decoding in Production

    Speculative decoding pairs a fast draft model with the target model. The draft proposes K tokens; the target verifies in a single forward; accepted tokens are free.

    Phase 17: Infrastructure & Production

  • KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Phase 07: Transformers Deep Dive

  • Inference Optimization

    Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.

    Phase 10: LLMs from Scratch

Covered in Phase 07: Transformers Deep Dive, Phase 10: LLMs from Scratch and Phase 17: Infrastructure & Production.

  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
  • Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.

Sources

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.