Multimodal systems · Glossary term

What is Early Fusion?

Combining raw or low-level representations from several modalities before most task-specific modeling occurs.

Why does Early Fusion matter?

Early interaction can expose fine-grained cross-modal relationships, but it also requires compatible representations and careful handling of alignment and missing inputs.

Early Fusion in practice

Convert each modality into a declared token or feature representation, preserve source and position markers, fuse them before the shared backbone, and compare against single-modality and late-fusion baselines.

What is the common confusion about Early Fusion?

Early fusion describes where streams are combined in the architecture. It does not guarantee that the model learns useful alignment between them.

Learn Early Fusion in the course

Start with

Taught in Phase 12: Multimodal AI.

  • Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
  • Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
  • Modality AlignmentLearning or establishing correspondences between representations from different modalities so semantically or temporally related items can…
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.