Math & training · Glossary term
What is Overfitting?
A generalization gap in which performance on training data is substantially better than performance on representative unseen data. Memorization can contribute, but the operational symptom is poor generalization.
“The model memorized the training data.”
Overfitting in practice
Compare training and held-out metrics, inspect subgroup failures, and test changes such as data quality, regularization, early stopping, or model capacity.
Learn Overfitting in the course
Lessons that name Overfitting in a title or section
- What Is Machine Learning
Machine learning is teaching computers to find patterns in data instead of writing rules by hand. Explain the difference between supervised, unsupervised, and reinforcement learning and identify…
- Regularization
Your model gets 99% on training data and 60% on test data. It memorized instead of learning. Regularization is the tax you impose on complexity to force generalization.
Covered in Phase 02: ML Fundamentals and Phase 03: Deep Learning Core.
Related terms
- UnderfittingA model or training setup has insufficient effective capacity, optimization, features, or training signal to capture useful patterns in…
- DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.
- Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- Data AugmentationCreating modified examples, such as transformed images, perturbed audio, or paraphrased text, to increase training diversity without…
- Data DeduplicationDetecting and removing exact and near-duplicate examples within or across datasets.
- Dataset SplitA documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
- Distribution ShiftA difference between the data distribution used to build or evaluate a system and the distribution it encounters after deployment.
- EpochOne traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.