Models & inference · Glossary term

What is MoE (Mixture of Experts)?

An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token. Sparse activation can increase total parameter capacity without using every expert on every forward pass.

What people say

“A large model that activates only part of its parameters for each token.”

Why does MoE (Mixture of Experts) matter?

Compute, memory, communication, routing balance, and quality depend on the specific architecture and serving system.

What is the common confusion about MoE (Mixture of Experts)?

Product names do not prove an MoE architecture unless the model developer discloses it.

Learn MoE (Mixture of Experts) in the course

Start with

  • Mixture of Experts (MoE)

    A dense 70B transformer activates every parameter for every token. A 671B MoE activates only 37B per token and beats it on every benchmark. Sparsity is the most important scaling idea of the decade.

    Phase 07: Transformers Deep Dive

Lessons that name MoE (Mixture of Experts) in a title or section

  • Open Models: Architecture Walkthroughs

    You built a GPT-2 Small from scratch in Lesson 04. Frontier open models in 2026 are the same family with five or six concrete changes. RMSNorm instead of LayerNorm. SwiGLU instead of GELU.

    Phase 10: LLMs from Scratch

  • Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d

    Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…

    Phase 17: Infrastructure & Production

Taught in Phase 07: Transformers Deep Dive.

Also covered in Phase 10: LLMs from Scratch and Phase 17: Infrastructure & Production.

  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.