Multimodal systems · Glossary term

What is Image Token?

A model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned visual-codebook entry.

Why does Image Token matter?

Turning visual input into a sequence lets transformer-style components process images together with text or other tokenized modalities.

Image Token in practice

Document whether tokens are continuous patches or discrete codes, preserve spatial position, test resolution and aspect-ratio changes, and count visual tokens in the model's input budget.

What is the common confusion about Image Token?

An image token is not necessarily one pixel, one object, or one fixed physical area. Its scope follows the visual encoder or tokenizer.

Learn Image Token in the course

Start with

Taught in Phase 04: Computer Vision.

  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.