Math & training · Glossary term

What is Softmax?

A function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization. Its outputs are positive and sum to one, so they can parameterize a categorical distribution.

What people say

“A function that turns logits into normalized positive values.”

What is the common confusion about Softmax?

Softmax values are not automatically calibrated probabilities about real-world correctness.

Learn Softmax in the course

Lessons that name Softmax in a title or section

  • Probability and Distributions

    Probability is the language AI uses to express uncertainty. Language: Python Implement PMFs and PDFs from scratch for Bernoulli, categorical, Poisson, uniform, and normal distributions.

    Phase 01: Math Foundations

  • Numerical Stability

    Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…

    Phase 01: Math Foundations

  • Logistic Regression

    Logistic regression bends a straight line into an S-curve to answer yes-or-no questions with probabilities. Implement logistic regression from scratch using the sigmoid function and binary…

    Phase 02: ML Fundamentals

  • Activation Functions

    Without nonlinearity, your 100-layer network is a fancy matrix multiply. Activations are the gates that let neural networks think in curves.

    Phase 03: Deep Learning Core

  • Loss Functions

    Your network makes a prediction. The ground truth says otherwise. How wrong is it? That number is the loss. Pick the wrong loss function and your model optimizes for the wrong thing entirely.

    Phase 03: Deep Learning Core

  • Image Classification

    A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…

    Phase 04: Computer Vision

  • Self-Attention from Scratch

    Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…

    Phase 07: Transformers Deep Dive

  • KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Phase 07: Transformers Deep Dive

Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 07: Transformers Deep Dive, Phase 09: Reinforcement Learning and Phase 10: LLMs from Scratch.

  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
  • Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.