Multimodal systems · Glossary term

What is Patch Embedding?

A learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.

Why does Patch Embedding matter?

It creates the interface between a spatial image grid and a sequence model, with patch size controlling token count and retained local detail.

Patch Embedding in practice

Record patch and image dimensions, handle padding or resizing explicitly, add position information, and measure how resolution changes affect both accuracy and token cost.

What is the common confusion about Patch Embedding?

A patch embedding is the vector representation of a patch, not a semantic object detector or a guarantee that patch boundaries match visual entities.

Learn Patch Embedding in the course

Start with

  • Vision Transformers and the Patch-Token Primitive

    Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.

    Phase 12: Multimodal AI

Lessons that name Patch Embedding in a title or section

  • Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Phase 04: Computer Vision

Taught in Phase 12: Multimodal AI.

Also covered in Phase 04: Computer Vision.

  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
  • Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.