Models & inference · Glossary term
What is Parameter?
A value learned during training, commonly a weight, bias, embedding element, or normalization parameter. Parameter count is one measure of model capacity, but it does not directly determine quality, memory, or serving cost.
“A number used to describe model size.”
What is the common confusion about Parameter?
Memory per parameter depends on numerical format, quantization metadata, sharding, optimizer state, activations, and runtime overhead.
Learn Parameter in the course
Lessons that name Parameter in a title or section
- Pre-Training a Mini GPT (124M Parameters)
GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.
- Tool Schema Design — Naming, Descriptions, Parameter Constraints
A correct tool fails silently when the model cannot tell when to use it. Naming, descriptions, and parameter shapes drive 10 to 20 percentage-point swings in tool-selection accuracy on benchmarks…
- Support Vector Machines
Find the widest street between two classes. That is the entire idea. Language: Python Implement a linear SVM from scratch using hinge loss and gradient descent on the primal formulation.
- Hyperparameter Tuning
Hyperparameters are the knobs you turn before training starts. Turning them well is the difference between a mediocre model and a great one.
- Anomaly Detection
Normal is easy to define. Abnormal is whatever doesn't fit. Language: Python Implement Z-score, IQR, and Isolation Forest anomaly detection methods from scratch.
- Introduction to JAX
PyTorch mutates tensors. TensorFlow builds graphs. JAX compiles pure functions. That last one changes how you think about deep learning.
- Real-Time Vision — Edge Deployment
Edge inference is the discipline of getting a 90-accuracy model to run at 30 fps on a device with 2 GB of RAM. Every percentage point of accuracy is traded against milliseconds of latency.
- Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick
Every voice agent lives or dies on two decisions: is the user speaking now, and are they done? VAD answers the first. Turn-detection (VAD + silence-hangover + semantic endpoint model) answers the…
Covered in Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 10: LLMs from Scratch, Phase 11: LLM Engineering, Phase 12: Multimodal AI and Phase 13: Tools & Protocols.
Related terms
- WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
- MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
- QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- Tensor ParallelismPartitioning tensor operations within a model layer across devices, with collective communication combining partial results during the…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.