Math & training · Glossary term

What is Cross-Entropy?

A loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it assigns low probability to the observed next token.

What people say

“The classification loss.”

What is the common confusion about Cross-Entropy?

Perplexity is the exponentiated average cross-entropy only when the averaging and logarithm base are defined consistently.

Learn Cross-Entropy in the course

Lessons that name Cross-Entropy in a title or section

  • Information Theory

    Information theory measures surprise. Loss functions are built on it. Language: Python Compute entropy, cross-entropy, and KL divergence from scratch and explain their relationship.

    Phase 01: Math Foundations

  • Numerical Stability

    Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…

    Phase 01: Math Foundations

  • Logistic Regression

    Logistic regression bends a straight line into an S-curve to answer yes-or-no questions with probabilities. Implement logistic regression from scratch using the sigmoid function and binary…

    Phase 02: ML Fundamentals

  • Loss Functions

    Your network makes a prediction. The ground truth says otherwise. How wrong is it? That number is the loss. Pick the wrong loss function and your model optimizes for the wrong thing entirely.

    Phase 03: Deep Learning Core

  • Image Classification

    A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…

    Phase 04: Computer Vision

  • Semantic Segmentation — U-Net

    Segmentation is classification at every pixel. U-Net makes it work by pairing a downsampling encoder with an upsampling decoder and wiring skip connections between them.

    Phase 04: Computer Vision

  • Instruction Tuning (SFT)

    A base model predicts the next token. That's it. It doesn't follow instructions, answer questions, or refuse harmful requests. SFT is the bridge between a token predictor and a useful assistant.

    Phase 10: LLMs from Scratch

Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision and Phase 10: LLMs from Scratch.

  • Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…

More terms in Math & training

Open the Math & training list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.