Infrastructure & serving · Glossary term
What is Expert Parallelism?
Distributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected experts.
Why does Expert Parallelism matter?
Sparse experts increase model capacity without executing every expert for every token, but routing introduces communication, load-balance, and placement constraints.
Expert Parallelism in practice
Measure token distribution by expert, provision communication bandwidth, cap or route overflow deliberately, and test quality when traffic produces uneven expert demand.
What is the common confusion about Expert Parallelism?
Expert parallelism partitions experts selected by a router. Tensor parallelism partitions the tensor operations inside layers.
Learn Expert Parallelism in the course
Start with
- Mixture of Experts (MoE)
A dense 70B transformer activates every parameter for every token. A 671B MoE activates only 37B per token and beats it on every benchmark. Sparsity is the most important scaling idea of the decade.
Taught in Phase 07: Transformers Deep Dive.
Related terms
- MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
- Tensor ParallelismPartitioning tensor operations within a model layer across devices, with collective communication combining partial results during the…
- Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.