Multimodal systems · Glossary term

What is Shared Embedding Space?

A common vector space in which representations from different modalities can be compared with the same similarity function.

Why does Shared Embedding Space matter?

It enables cross-modal retrieval and matching, such as finding images from text, without requiring both items to share a raw representation.

Shared Embedding Space in practice

Train paired and unpaired negatives deliberately, normalize vectors when the objective requires it, evaluate both retrieval directions, and inspect subgroup and language performance.

What is the common confusion about Shared Embedding Space?

Sharing a vector dimension does not create a shared semantic space. The training objective and data must establish cross-modal comparability.

Learn Shared Embedding Space in the course

Start with

  • CLIP and Contrastive Vision-Language Pretraining

    OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a…

    Phase 12: Multimodal AI

Taught in Phase 12: Multimodal AI.

  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • Modality AlignmentLearning or establishing correspondences between representations from different modalities so semantically or temporally related items can…
  • Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.