Models & inference · Glossary term
What is Nucleus Sampling (Top-p)?
A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
Also called Top-p sampling.
Why does Nucleus Sampling (Top-p) matter?
The candidate-set size adapts to the distribution, retaining more options when uncertainty is broad and fewer when probability is concentrated.
Nucleus Sampling (Top-p) in practice
Evaluate the threshold with temperature and stop settings held constant, and record the complete decoding configuration with every result.
What is the common confusion about Nucleus Sampling (Top-p)?
Top-p is a probability-mass threshold, while top-k always keeps a fixed maximum number of candidates.
Learn Nucleus Sampling (Top-p) in the course
Lessons that name Nucleus Sampling (Top-p) in a title or section
- Sampling Methods
Sampling is how AI explores the space of possibilities. Language: Python Implement inverse CDF, rejection, and importance sampling from scratch using only uniform random numbers.
Covered in Phase 01: Math Foundations.
Related terms
- Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
Sources
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.