Multimodal systems · Glossary term
What is Visual Grounding?
Connecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
Why does Visual Grounding matter?
A fluent visual answer can be unsupported, while grounding makes the claimed referent inspectable and enables region-level evaluation.
Visual Grounding in practice
Require a box, mask, or temporal segment with the answer, test ambiguous and absent referents, and score localization separately from language correctness.
What is the common confusion about Visual Grounding?
Visual grounding identifies where the referenced evidence is. General image captioning can describe a scene without localizing each claim.
Learn Visual Grounding in the course
Start with
- Cross-Attention Fusion
The projection layer aligns one image vector with one caption vector. A real vision-language decoder needs every text token to attend to every patch token, so the model can ground each word in a…
Taught in Phase 19: Capstone Projects.
Related terms
- GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
- Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.