Math & training · Glossary term
What is Hyperparameter?
A configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary model parameter. Examples include learning rate, batch size, layer count, and decoding settings.
“A setting you tune.”
What is the common confusion about Hyperparameter?
Some hyperparameters are selected before training, while others can be changed during a schedule or at inference time.
Learn Hyperparameter in the course
Lessons that name Hyperparameter in a title or section
- Hyperparameter Tuning
Hyperparameters are the knobs you turn before training starts. Turning them well is the difference between a mediocre model and a great one.
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- MDPs, States, Actions & Rewards
A Markov Decision Process is five things: states, actions, transitions, rewards, a discount. Everything in RL — Q-learning, PPO, DPO, GRPO — optimizes over this shape.
- Instruction Tuning (SFT)
A base model predicts the next token. That's it. It doesn't follow instructions, answer questions, or refuse harmful requests. SFT is the bridge between a token predictor and a useful assistant.
Covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 09: Reinforcement Learning and Phase 10: LLMs from Scratch.
Related terms
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- UnderfittingA model or training setup has insufficient effective capacity, optimization, features, or training signal to capture useful patterns in…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.