Infrastructure & serving · Glossary term

What is Tensor Parallelism?

Partitioning tensor operations within a model layer across devices, with collective communication combining partial results during the layer computation.

Why does Tensor Parallelism matter?

It lets one layer use memory and compute from several devices, but frequent communication can dominate when the interconnect or partition is unsuitable.

Tensor Parallelism in practice

Match partition dimensions to model shapes, benchmark collective traffic, keep ranks on a fast interconnect, and record the sharding layout with checkpoints and serving configuration.

What is the common confusion about Tensor Parallelism?

Tensor parallelism splits work inside layers. Pipeline parallelism places different layer groups on different devices.

Learn Tensor Parallelism in the course

Start with

  • Scaling: Distributed Training, FSDP, DeepSpeed

    Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.

    Phase 10: LLMs from Scratch

Taught in Phase 10: LLMs from Scratch.

  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
  • Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.