Models & inference · Glossary term

What is Parameter?

A value learned during training, commonly a weight, bias, embedding element, or normalization parameter. Parameter count is one measure of model capacity, but it does not directly determine quality, memory, or serving cost.

What people say

“A number used to describe model size.”

What is the common confusion about Parameter?

Memory per parameter depends on numerical format, quantization metadata, sharding, optimizer state, activations, and runtime overhead.

Learn Parameter in the course

Lessons that name Parameter in a title or section

  • Pre-Training a Mini GPT (124M Parameters)

    GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.

    Phase 10: LLMs from Scratch

  • Tool Schema Design — Naming, Descriptions, Parameter Constraints

    A correct tool fails silently when the model cannot tell when to use it. Naming, descriptions, and parameter shapes drive 10 to 20 percentage-point swings in tool-selection accuracy on benchmarks…

    Phase 13: Tools & Protocols

  • Support Vector Machines

    Find the widest street between two classes. That is the entire idea. Language: Python Implement a linear SVM from scratch using hinge loss and gradient descent on the primal formulation.

    Phase 02: ML Fundamentals

  • Hyperparameter Tuning

    Hyperparameters are the knobs you turn before training starts. Turning them well is the difference between a mediocre model and a great one.

    Phase 02: ML Fundamentals

  • Anomaly Detection

    Normal is easy to define. Abnormal is whatever doesn't fit. Language: Python Implement Z-score, IQR, and Isolation Forest anomaly detection methods from scratch.

    Phase 02: ML Fundamentals

  • Introduction to JAX

    PyTorch mutates tensors. TensorFlow builds graphs. JAX compiles pure functions. That last one changes how you think about deep learning.

    Phase 03: Deep Learning Core

  • Real-Time Vision — Edge Deployment

    Edge inference is the discipline of getting a 90-accuracy model to run at 30 fps on a device with 2 GB of RAM. Every percentage point of accuracy is traded against milliseconds of latency.

    Phase 04: Computer Vision

  • Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick

    Every voice agent lives or dies on two decisions: is the user speaking now, and are they done? VAD answers the first. Turn-detection (VAD + silence-hangover + semantic endpoint model) answers the…

    Phase 06: Speech & Audio

Covered in Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 10: LLMs from Scratch, Phase 11: LLM Engineering, Phase 12: Multimodal AI and Phase 13: Tools & Protocols.

  • WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
  • MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • Tensor ParallelismPartitioning tensor operations within a model layer across devices, with collective communication combining partial results during the…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.