Multimodal systems · Glossary term
What is Multimodal Model?
A model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or coordinated prediction.
Why does Multimodal Model matter?
Multimodal capability depends on how modalities interact, not simply on accepting several input types, and failures can occur at each representation boundary.
Multimodal Model in practice
Document supported input and output combinations, evaluate each modality alone and together, test missing or conflicting inputs, and track preprocessing versions with the model.
What is the common confusion about Multimodal Model?
A pipeline with separate image and text models is multimodal at the system level, but it is not necessarily one jointly trained multimodal model.
Learn Multimodal Model in the course
Start with
- MIO and Any-to-Any Streaming Multimodal Models
GPT-4o ships a product most open models cannot replicate: an agent that hears voice, sees video, and speaks back in real time.
Lessons that name Multimodal Model in a title or section
- Chameleon and Early-Fusion Token-Only Multimodal Models
Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.
- Janus-Pro: Decoupled Encoders for Unified Multimodal Models
Unified multimodal models have an unavoidable tension. Understanding wants semantic features — SigLIP or DINOv2 output vectors rich with concept-level information.
Taught in Phase 12: Multimodal AI.
Related terms
- ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
- Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
- Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.