Models & inference · Glossary term
What is Temperature?
A decoding parameter that rescales logits before a probability distribution is formed. Higher positive values usually flatten the distribution; lower positive values sharpen it.
“A creativity setting.”
Why does Temperature matter?
Temperature changes sampling behavior, not the model's knowledge or factuality.
What is the common confusion about Temperature?
A zero setting is often implemented as greedy decoding, but exact behavior and determinism depend on the provider, sampler, seed support, and serving system.
Learn Temperature in the course
Lessons that name Temperature in a title or section
- Sampling Methods
Sampling is how AI explores the space of possibilities. Language: Python Implement inverse CDF, rejection, and importance sampling from scratch using only uniform random numbers.
- Prompt Engineering: Techniques & Patterns
Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.
- CLIP and Contrastive Vision-Language Pretraining
OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a…
- Emu3: Next-Token Prediction for Image and Video Generation
BAAI's Emu3 (Wang et al., September 2024) is the 2024 result that should have ended the diffusion-versus-autoregressive debate.
- GPT Model Assembly
Twelve blocks stacked, a token embedding, a learned position embedding, a final LayerNorm, and a tied language model head. That is the entire 124 million parameter GPT model.
- Vision-Language Pretraining
The encoder, projection, and decoder are wired. Now train them together. Two objectives drive learning: a contrastive image-text loss (InfoNCE) that pulls matching pairs together in the joint…
Covered in Phase 01: Math Foundations, Phase 11: LLM Engineering, Phase 12: Multimodal AI and Phase 19: Capstone Projects.
Related terms
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
- Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
- Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.