Multimodal systems · Glossary term

What is Vision Transformer (ViT)?

A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.

Why does Vision Transformer (ViT) matter?

It provides a sequence-model interface for visual data, but performance and compute depend on patch size, resolution, pretraining, and inductive biases.

Vision Transformer (ViT) in practice

Keep patching and normalization consistent with training, account for position-embedding behavior at new resolutions, and compare against a suitable visual baseline on the target dataset.

What is the common confusion about Vision Transformer (ViT)?

ViT is an architecture family, not every transformer that accepts images, and its patches are not inherently semantic objects.

Learn Vision Transformer (ViT) in the course

Start with

  • Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Phase 04: Computer Vision

Lessons that name Vision Transformer (ViT) in a title or section

  • Vision Transformers (ViT)

    An image is a grid of patches. A sentence is a grid of tokens. The same transformer eats both. Before 2020, computer vision meant convolutions.

    Phase 07: Transformers Deep Dive

  • Vision Transformers and the Patch-Token Primitive

    Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.

    Phase 12: Multimodal AI

  • Vision Transformer Encoder

    Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…

    Phase 19: Capstone Projects

Taught in Phase 04: Computer Vision.

Also covered in Phase 07: Transformers Deep Dive, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
  • Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.