Multimodal systems · Glossary term

What is Modality?

A form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.

Why does Modality matter?

Different modalities have different sampling rates, noise, spatial or temporal structure, and missing-data behavior, so one preprocessing assumption rarely fits all of them.

Modality in practice

Document each modality's source, units, resolution, timing, preprocessing, and missing-value policy before designing alignment or fusion.

What is the common confusion about Modality?

A modality is not merely a file extension or feature column. Several encodings can represent one modality, and one sample can contain several modalities.

Learn Modality in the course

Start with

Lessons that name Modality in a title or section

  • From CLIP to BLIP-2 — Q-Former as Modality Bridge

    CLIP aligns image and text but cannot generate captions, answer questions, or hold a conversation. BLIP-2 (Salesforce, 2023) solved that with a small trainable bridge: 32 learnable query vectors…

    Phase 12: Multimodal AI

  • Projection Layer for Modality Alignment

    A vision encoder produces image tokens. A text decoder consumes text tokens. The two live in different vector spaces. A small two-layer MLP projects image tokens into the text embedding space, and a…

    Phase 19: Capstone Projects

Taught in Phase 12: Multimodal AI.

Also covered in Phase 19: Capstone Projects.

  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.