Multimodal systems · Glossary term

What is Late Fusion?

Processing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near the task output.

Why does Late Fusion matter?

Separate branches can use modality-specific architectures and tolerate missing inputs, though they may miss fine-grained interactions available to earlier fusion.

Late Fusion in practice

Calibrate each branch, define how missing modalities affect the merge, compare score-level and feature-level combinations, and evaluate each branch alone as an ablation.

What is the common confusion about Late Fusion?

Late fusion describes the position of combination. It does not mean simple averaging or guarantee that the modalities contribute equally.

Learn Late Fusion in the course

Start with

  • Cross-Attention Fusion

    The projection layer aligns one image vector with one caption vector. A real vision-language decoder needs every text token to attend to every patch token, so the model can ground each word in a…

    Phase 19: Capstone Projects

Taught in Phase 19: Capstone Projects.

  • Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.
  • Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
  • ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.