Infrastructure & serving · Glossary term

What is Pipeline Parallelism?

Partitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.

Why does Pipeline Parallelism matter?

It lets models exceed one device's memory, but stage imbalance, pipeline bubbles, activation transfers, and failure coordination affect usable performance.

Pipeline Parallelism in practice

Balance stage cost, choose a microbatch schedule, measure idle time and interconnect traffic, and keep model and checkpoint partition metadata versioned.

What is the common confusion about Pipeline Parallelism?

Pipeline parallelism divides layers by depth. Tensor parallelism divides tensor operations within a layer.

Learn Pipeline Parallelism in the course

Start with

  • Scaling: Distributed Training, FSDP, DeepSpeed

    Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.

    Phase 10: LLMs from Scratch

Lessons that name Pipeline Parallelism in a title or section

  • DualPipe Parallelism

    DeepSeek-V3 was trained on 2,048 H800 GPUs with MoE experts scattered across nodes. Cross-node expert all-to-all communication cost 1 GPU-hour of comm for every 1 GPU-hour of compute.

    Phase 10: LLMs from Scratch

Taught in Phase 10: LLMs from Scratch.

  • Tensor ParallelismPartitioning tensor operations within a model layer across devices, with collective communication combining partial results during the…
  • Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
  • Batch SizeThe number of examples whose losses contribute to one gradient estimate before an optimizer update.
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.