Multimodal systems · Glossary term
What is Early Fusion?
Combining raw or low-level representations from several modalities before most task-specific modeling occurs.
Why does Early Fusion matter?
Early interaction can expose fine-grained cross-modal relationships, but it also requires compatible representations and careful handling of alignment and missing inputs.
Early Fusion in practice
Convert each modality into a declared token or feature representation, preserve source and position markers, fuse them before the shared backbone, and compare against single-modality and late-fusion baselines.
What is the common confusion about Early Fusion?
Early fusion describes where streams are combined in the architecture. It does not guarantee that the model learns useful alignment between them.
Learn Early Fusion in the course
Start with
- Chameleon and Early-Fusion Token-Only Multimodal Models
Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.
Taught in Phase 12: Multimodal AI.
Related terms
- Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
- Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
- Modality AlignmentLearning or establishing correspondences between representations from different modalities so semantically or temporally related items can…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.