Models & inference · Glossary term
What is Perplexity?
The exponentiated average negative log-likelihood under a stated tokenization and logarithm convention. Lower values mean the model assigned higher probability to the evaluated sequence.
“How surprised a language model is by a dataset.”
What is the common confusion about Perplexity?
Perplexity is not comparable across different tokenizers or evaluation setups and does not directly measure factuality or usefulness.
Learn Perplexity in the course
Lessons that name Perplexity in a title or section
- Perplexity and Calibration
If your model says 90 percent confident on a thousand answers and gets six hundred right, it is not well calibrated. Calibration is half of trustworthy eval.
- Information Theory
Information theory measures surprise. Loss functions are built on it. Language: Python Compute entropy, cross-entropy, and KL divergence from scratch and explain their relationship.
- Text Generation Before Transformers — N-gram Language Models
If a word is surprising, the model is bad. Perplexity makes surprise a number. Smoothing keeps it finite. Before transformers, before RNNs, before word embeddings, a language model predicted the…
- Evaluation: Benchmarks, Evals, LM Harness
Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Every frontier lab games benchmarks. MMLU scores go up while models still can't reliably count the number of R's in…
Covered in Phase 01: Math Foundations, Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.
Related terms
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.