Data & representations · Glossary term

What is Token?

An integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation. A token can be a whole word, part of a word, punctuation, whitespace, a byte sequence, or a special control symbol.

What people say

“A word-sized piece of model input or output.”

What is the common confusion about Token?

Character-to-token ratios vary by language, content, and tokenizer, so count with the target model's tokenizer or provider tools.

Learn Token in the course

Start with

Lessons that name Token in a title or section

Taught in Phase 10: LLMs from Scratch.

Also covered in Phase 04: Computer Vision, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 09: Reinforcement Learning, Phase 11: LLM Engineering, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems, Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.

  • Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
  • Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
  • Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
  • Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.
  • Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
  • ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
  • Stop SequenceAn application-specified token or text pattern that causes generation to stop when the decoding system encounters it.
  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.

More terms in Data & representations

Open the Data & representations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.