Math & training · Glossary term
What is Stochastic Gradient Descent (SGD)?
An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training dataset.
Also called SGD.
Why does Stochastic Gradient Descent (SGD) matter?
It is the baseline for understanding gradient noise, momentum, batch scaling, and the adaptive optimizers used in modern training.
Stochastic Gradient Descent (SGD) in practice
Record batch sampling, learning rate, momentum if used, and schedule, then compare validation behavior under equal update or token budgets.
What is the common confusion about Stochastic Gradient Descent (SGD)?
In current practice, SGD usually means minibatch SGD, and its useful learning rate does not follow one universal batch-scaling rule.
Learn Stochastic Gradient Descent (SGD) in the course
Lessons that name Stochastic Gradient Descent (SGD) in a title or section
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- Build Your Own Mini Framework
You have built neurons, layers, networks, backprop, activations, loss functions, optimizers, regularization, initialization, and LR schedules. All as separate pieces.
Covered in Phase 01: Math Foundations and Phase 03: Deep Learning Core.
Related terms
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
Sources
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.