Multimodal systems · Glossary term
What is Modality?
A form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
Why does Modality matter?
Different modalities have different sampling rates, noise, spatial or temporal structure, and missing-data behavior, so one preprocessing assumption rarely fits all of them.
Modality in practice
Document each modality's source, units, resolution, timing, preprocessing, and missing-value policy before designing alignment or fusion.
What is the common confusion about Modality?
A modality is not merely a file extension or feature column. Several encodings can represent one modality, and one sample can contain several modalities.
Learn Modality in the course
Start with
- MIO and Any-to-Any Streaming Multimodal Models
GPT-4o ships a product most open models cannot replicate: an agent that hears voice, sees video, and speaks back in real time.
Lessons that name Modality in a title or section
- From CLIP to BLIP-2 — Q-Former as Modality Bridge
CLIP aligns image and text but cannot generate captions, answer questions, or hold a conversation. BLIP-2 (Salesforce, 2023) solved that with a small trainable bridge: 32 learnable query vectors…
- Projection Layer for Modality Alignment
A vision encoder produces image tokens. A text decoder consumes text tokens. The two live in different vector spaces. A small two-layer MLP projects image tokens into the text embedding space, and a…
Taught in Phase 12: Multimodal AI.
Also covered in Phase 19: Capstone Projects.
Related terms
- Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.