Infrastructure & serving · Glossary term
What is Tensor Parallelism?
Partitioning tensor operations within a model layer across devices, with collective communication combining partial results during the layer computation.
Why does Tensor Parallelism matter?
It lets one layer use memory and compute from several devices, but frequent communication can dominate when the interconnect or partition is unsuitable.
Tensor Parallelism in practice
Match partition dimensions to model shapes, benchmark collective traffic, keep ranks on a fast interconnect, and record the sharding layout with checkpoints and serving configuration.
What is the common confusion about Tensor Parallelism?
Tensor parallelism splits work inside layers. Pipeline parallelism places different layer groups on different devices.
Learn Tensor Parallelism in the course
Start with
- Scaling: Distributed Training, FSDP, DeepSpeed
Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.
Taught in Phase 10: LLMs from Scratch.
Related terms
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
- Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.