Multimodal systems · Glossary term

What is Modality Alignment?

Learning or establishing correspondences between representations from different modalities so semantically or temporally related items can be matched.

Why does Modality Alignment matter?

Fusion and cross-modal retrieval fail when the system cannot connect the same event, object, or concept across differently structured inputs.

Modality Alignment in practice

Define positive and negative pairs, preserve time or spatial metadata, evaluate mismatched examples, and measure alignment separately from downstream task accuracy.

What is the common confusion about Modality Alignment?

Alignment makes representations comparable or corresponding. It does not require them to become identical or erase modality-specific information.

Learn Modality Alignment in the course

Start with

  • Projection Layer for Modality Alignment

    A vision encoder produces image tokens. A text decoder consumes text tokens. The two live in different vector spaces. A small two-layer MLP projects image tokens into the text embedding space, and a…

    Phase 19: Capstone Projects

Taught in Phase 19: Capstone Projects.

  • Shared Embedding SpaceA common vector space in which representations from different modalities can be compared with the same similarity function.
  • Contrastive LearningTraining by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space.
  • GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
  • Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
  • Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.