Models & inference · Glossary term
What is Decoding Strategy?
The algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
Why does Decoding Strategy matter?
Greedy selection, sampling, truncation, and search can produce different quality, diversity, latency, and repeatability from the same logits.
Decoding Strategy in practice
Define the task's decoding settings, stop rules, and seed behavior in the eval configuration so results can be compared fairly.
What is the common confusion about Decoding Strategy?
Decoding changes how outputs are selected; it does not change the model's trained parameters or add knowledge.
Learn Decoding Strategy in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
- Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
- Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
- Stop SequenceAn application-specified token or text pattern that causes generation to stop when the decoding system encounters it.
Sources
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.