Models & inference · Glossary term

What is Temperature?

A decoding parameter that rescales logits before a probability distribution is formed. Higher positive values usually flatten the distribution; lower positive values sharpen it.

What people say

“A creativity setting.”

Why does Temperature matter?

Temperature changes sampling behavior, not the model's knowledge or factuality.

What is the common confusion about Temperature?

A zero setting is often implemented as greedy decoding, but exact behavior and determinism depend on the provider, sampler, seed support, and serving system.

Learn Temperature in the course

Lessons that name Temperature in a title or section

  • Sampling Methods

    Sampling is how AI explores the space of possibilities. Language: Python Implement inverse CDF, rejection, and importance sampling from scratch using only uniform random numbers.

    Phase 01: Math Foundations

  • Prompt Engineering: Techniques & Patterns

    Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.

    Phase 11: LLM Engineering

  • CLIP and Contrastive Vision-Language Pretraining

    OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a…

    Phase 12: Multimodal AI

  • Emu3: Next-Token Prediction for Image and Video Generation

    BAAI's Emu3 (Wang et al., September 2024) is the 2024 result that should have ended the diffusion-versus-autoregressive debate.

    Phase 12: Multimodal AI

  • GPT Model Assembly

    Twelve blocks stacked, a token embedding, a learned position embedding, a final LayerNorm, and a tied language model head. That is the entire 124 million parameter GPT model.

    Phase 19: Capstone Projects

  • Vision-Language Pretraining

    The encoder, projection, and decoder are wired. Now train them together. Two objectives drive learning: a contrastive image-text loss (InfoNCE) that pulls matching pairs together in the joint…

    Phase 19: Capstone Projects

Covered in Phase 01: Math Foundations, Phase 11: LLM Engineering, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
  • HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
  • Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
  • Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.