Math & training · Glossary term
What is Cross-Entropy?
A loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it assigns low probability to the observed next token.
“The classification loss.”
What is the common confusion about Cross-Entropy?
Perplexity is the exponentiated average cross-entropy only when the averaging and logarithm base are defined consistently.
Learn Cross-Entropy in the course
Lessons that name Cross-Entropy in a title or section
- Information Theory
Information theory measures surprise. Loss functions are built on it. Language: Python Compute entropy, cross-entropy, and KL divergence from scratch and explain their relationship.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
- Logistic Regression
Logistic regression bends a straight line into an S-curve to answer yes-or-no questions with probabilities. Implement logistic regression from scratch using the sigmoid function and binary…
- Loss Functions
Your network makes a prediction. The ground truth says otherwise. How wrong is it? That number is the loss. Pick the wrong loss function and your model optimizes for the wrong thing entirely.
- Image Classification
A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…
- Semantic Segmentation — U-Net
Segmentation is classification at every pixel. U-Net makes it work by pairing a downsampling encoder with an upsampling decoder and wiring skip connections between them.
- Instruction Tuning (SFT)
A base model predicts the next token. That's it. It doesn't follow instructions, answer questions, or refuse harmful requests. SFT is the bridge between a token predictor and a useful assistant.
Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision and Phase 10: LLMs from Scratch.
Related terms
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.