Multimodal systems · Glossary term

What is Cross-Attention?

Attention in which the query representation comes from one sequence or representation while keys and values come from another.

Why does Cross-Attention matter?

It gives one stream a learnable way to retrieve information from another, such as language tokens attending to visual features.

Cross-Attention in practice

State which stream supplies queries, keys, and values, apply masks for missing or invalid positions, and inspect whether the model still performs when one modality is ablated.

What is the common confusion about Cross-Attention?

Cross-attention is not intrinsically multimodal. It can connect two text sequences or other representations; self-attention instead derives queries, keys, and values from the same sequence representation.

Learn Cross-Attention in the course

Lessons that name Cross-Attention in a title or section

  • Flamingo and Gated Cross-Attention for Few-Shot VLMs

    DeepMind's Flamingo (2022) did two things before anyone else. It showed a single model could process arbitrarily interleaved sequences of images, videos, and text.

    Phase 12: Multimodal AI

  • Cross-Attention Fusion

    The projection layer aligns one image vector with one caption vector. A real vision-language decoder needs every text token to attend to every patch token, so the model can ground each word in a…

    Phase 19: Capstone Projects

  • From CLIP to BLIP-2 — Q-Former as Modality Bridge

    CLIP aligns image and text but cannot generate captions, answer questions, or hold a conversation. BLIP-2 (Salesforce, 2023) solved that with a small trainable bridge: 32 learnable query vectors…

    Phase 12: Multimodal AI

Covered in Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
  • Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
  • Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.