Multimodal systems · Glossary term
What is Vision-Language Model (VLM)?
A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.
Why does Vision-Language Model (VLM) matter?
VLM performance depends on the visual encoder, language component, connection mechanism, training data, and resolution policy rather than one generic capability label.
Vision-Language Model (VLM) in practice
Evaluate text-only and vision-only controls, vary image resolution and layout, require evidence localization where possible, and report failures by visual skill and language.
What is the common confusion about Vision-Language Model (VLM)?
Accepting an image does not prove the model uses it correctly, and a VLM is not necessarily able to generate images.
Learn Vision-Language Model (VLM) in the course
Start with
- Vision-Language Models — The ViT-MLP-LLM Pattern
A vision encoder converts an image into tokens. An MLP projector maps those tokens into the LLM's embedding space. A language model does the rest.
Lessons that name Vision-Language Model (VLM) in a title or section
- Flamingo and Gated Cross-Attention for Few-Shot VLMs
DeepMind's Flamingo (2022) did two things before anyone else. It showed a single model could process arbitrarily interleaved sequences of images, videos, and text.
- Open-Weight VLM Recipes: What Actually Matters
The 2024-2026 open-weight VLM literature is a forest of ablation tables. Apple's MM1 tested 13 combinations of image encoder, connector, and data mix.
- World Models & Video Diffusion
A video model that predicts the next seconds of a scene is a world simulator. Condition that prediction on actions and you have a learned game engine.
Taught in Phase 04: Computer Vision.
Also covered in Phase 12: Multimodal AI.
Related terms
- Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
- Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
- Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
- Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.