Models & inference · Glossary term
What are Logits?
The model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into selections.
Why do Logits matter?
Temperature, softmax, top-k, and top-p operate on or derive from logits, so logits connect model computation to generated tokens.
Logits in practice
Inspect logits or log probabilities when the API exposes them, apply masks before sampling, and avoid interpreting raw magnitude as calibrated confidence.
What is the common confusion about Logits?
Logits are not probabilities and are not comparable across unrelated positions, models, or tasks without a defined transformation.
Learn Logits in the course
Lessons that name Logits in a title or section
- Image Classification
A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…
Covered in Phase 04: Computer Vision.
Related terms
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- Knowledge DistillationTraining a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary…
- Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
Sources
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.