Math & training · Glossary term
What is Softmax?
A function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization. Its outputs are positive and sum to one, so they can parameterize a categorical distribution.
“A function that turns logits into normalized positive values.”
What is the common confusion about Softmax?
Softmax values are not automatically calibrated probabilities about real-world correctness.
Learn Softmax in the course
Lessons that name Softmax in a title or section
- Probability and Distributions
Probability is the language AI uses to express uncertainty. Language: Python Implement PMFs and PDFs from scratch for Bernoulli, categorical, Poisson, uniform, and normal distributions.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
- Logistic Regression
Logistic regression bends a straight line into an S-curve to answer yes-or-no questions with probabilities. Implement logistic regression from scratch using the sigmoid function and binary…
- Activation Functions
Without nonlinearity, your 100-layer network is a fancy matrix multiply. Activations are the gates that let neural networks think in curves.
- Loss Functions
Your network makes a prediction. The ground truth says otherwise. How wrong is it? That number is the loss. Pick the wrong loss function and your model optimizes for the wrong thing entirely.
- Image Classification
A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…
- Self-Attention from Scratch
Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…
- KV Cache, Flash Attention & Inference Optimization
Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…
Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 07: Transformers Deep Dive, Phase 09: Reinforcement Learning and Phase 10: LLMs from Scratch.
Related terms
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
- Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.