Multimodal systems · Glossary term

What is Vision-Language Model (VLM)?

A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.

Why does Vision-Language Model (VLM) matter?

VLM performance depends on the visual encoder, language component, connection mechanism, training data, and resolution policy rather than one generic capability label.

Vision-Language Model (VLM) in practice

Evaluate text-only and vision-only controls, vary image resolution and layout, require evidence localization where possible, and report failures by visual skill and language.

What is the common confusion about Vision-Language Model (VLM)?

Accepting an image does not prove the model uses it correctly, and a VLM is not necessarily able to generate images.

Learn Vision-Language Model (VLM) in the course

Start with

Lessons that name Vision-Language Model (VLM) in a title or section

  • Flamingo and Gated Cross-Attention for Few-Shot VLMs

    DeepMind's Flamingo (2022) did two things before anyone else. It showed a single model could process arbitrarily interleaved sequences of images, videos, and text.

    Phase 12: Multimodal AI

  • Open-Weight VLM Recipes: What Actually Matters

    The 2024-2026 open-weight VLM literature is a forest of ablation tables. Apple's MM1 tested 13 combinations of image encoder, connector, and data mix.

    Phase 12: Multimodal AI

  • World Models & Video Diffusion

    A video model that predicts the next seconds of a scene is a world simulator. Condition that prediction on actions and you have a learned game engine.

    Phase 04: Computer Vision

Taught in Phase 04: Computer Vision.

Also covered in Phase 12: Multimodal AI.

  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
  • Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
  • Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.