Models & inference · Glossary term
What is MoE (Mixture of Experts)?
An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token. Sparse activation can increase total parameter capacity without using every expert on every forward pass.
“A large model that activates only part of its parameters for each token.”
Why does MoE (Mixture of Experts) matter?
Compute, memory, communication, routing balance, and quality depend on the specific architecture and serving system.
What is the common confusion about MoE (Mixture of Experts)?
Product names do not prove an MoE architecture unless the model developer discloses it.
Learn MoE (Mixture of Experts) in the course
Start with
- Mixture of Experts (MoE)
A dense 70B transformer activates every parameter for every token. A 671B MoE activates only 37B per token and beats it on every benchmark. Sparsity is the most important scaling idea of the decade.
Lessons that name MoE (Mixture of Experts) in a title or section
- Open Models: Architecture Walkthroughs
You built a GPT-2 Small from scratch in Lesson 04. Frontier open models in 2026 are the same family with five or six concrete changes. RMSNorm instead of LayerNorm. SwiGLU instead of GELU.
- Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…
Taught in Phase 07: Transformers Deep Dive.
Also covered in Phase 10: LLMs from Scratch and Phase 17: Infrastructure & Production.
Related terms
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.