Math & training · Glossary term
What is Weight?
A trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the training objective.
“A learned number inside a model.”
What is the common confusion about Weight?
Not every parameter is called a weight; biases, embeddings, and normalization scales are parameters too.
Learn Weight in the course
Lessons that name Weight in a title or section
- Weight Initialization and Training Stability
Initialize wrong and training never starts. Initialize right and 50 layers train as smoothly as 3. Implement zero, random, Xavier/Glorot, and Kaiming/He initialization strategies and measure their…
- Loading Pretrained Weights
Training a 124 million parameter model from scratch is a budget decision; loading a published checkpoint is a Tuesday. This lesson loads pretrained GPT-2 style weights from a safetensors file into…
- Handling Imbalanced Data
When 99% of your data is "normal," accuracy is a lie. Language: Python Implement SMOTE from scratch and explain how synthetic oversampling differs from random duplication.
- Multi-Layer Networks and Forward Pass
One neuron draws a line. Stack them, and you can draw anything. Build a multi-layer network from scratch with Layer and Network classes that perform a complete forward pass.
- Optimizers
Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.
- Regularization
Your model gets 99% on training data and 60% on test data. It memorized instead of learning. Regularization is the tax you impose on complexity to force generalization.
- Debugging Neural Networks
Your network compiled. It ran. It produced a number. The number is wrong and nothing crashed. Welcome to the hardest kind of debugging -- the kind where there is no error message.
- Self-Attention from Scratch
Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…
Covered in Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 07: Transformers Deep Dive, Phase 11: LLM Engineering and Phase 19: Capstone Projects.
Related terms
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
More terms in Math & training
- Activation Checkpointing
- Activation Function
- Adam (Optimizer)
- AdamW
- Autograd
- Backpropagation
- Batch Size
- Contrastive Learning
- Cross-Entropy
- Data Augmentation
- DPO (Direct Preference Optimization)
- Dropout
- Eigenvalue
- Epoch
- Fine-tuning
- Gradient
- Gradient Accumulation
- Gradient Clipping
- Gradient Descent
- Hyperparameter
- JAX
- Knowledge Distillation
- Learning Rate
- Learning Rate Schedule
- LoRA (Low-Rank Adaptation)
- Loss Function
- Mixed Precision
- NaN (Not a Number)
- Normalization
- Optimizer
- Overfitting
- QLoRA
- ReLU
- RLHF (Reinforcement Learning from Human Feedback)
- SFT (Supervised Fine-Tuning)
- Softmax
- Stochastic Gradient Descent (SGD)
- Transfer Learning
- Underfitting
- Warmup
- Weight Decay
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.