Multimodal systems · Glossary term
What is Vision Transformer (ViT)?
A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.
Why does Vision Transformer (ViT) matter?
It provides a sequence-model interface for visual data, but performance and compute depend on patch size, resolution, pretraining, and inductive biases.
Vision Transformer (ViT) in practice
Keep patching and normalization consistent with training, account for position-embedding behavior at new resolutions, and compare against a suitable visual baseline on the target dataset.
What is the common confusion about Vision Transformer (ViT)?
ViT is an architecture family, not every transformer that accepts images, and its patches are not inherently semantic objects.
Learn Vision Transformer (ViT) in the course
Start with
- Vision Transformers (ViT)
Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…
Lessons that name Vision Transformer (ViT) in a title or section
- Vision Transformers (ViT)
An image is a grid of patches. A sentence is a grid of tokens. The same transformer eats both. Before 2020, computer vision meant convolutions.
- Vision Transformers and the Patch-Token Primitive
Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.
- Vision Transformer Encoder
Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…
Taught in Phase 04: Computer Vision.
Also covered in Phase 07: Transformers Deep Dive, Phase 12: Multimodal AI and Phase 19: Capstone Projects.
Related terms
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
- Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
- EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
- Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
- Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.