Data & representations · Glossary term

What is Vocabulary?

The finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.

Why does Vocabulary matter?

Vocabulary design affects sequence length, multilingual coverage, code representation, embedding size, and compatibility between tokenizers and model weights.

Vocabulary in practice

Version the vocabulary and special-token assignments with the model, test encode-decode round trips, and never substitute a tokenizer with merely similar token names.

What is the common confusion about Vocabulary?

A model vocabulary is not a dictionary of human words; many entries are fragments, bytes, whitespace patterns, or control symbols.

Learn Vocabulary in the course

Lessons that name Vocabulary in a title or section

  • Docker for AI

    Containers make "works on my machine" a thing of the past. Build a GPU-enabled Docker image with CUDA, PyTorch, and AI libraries from a Dockerfile.

    Phase 00: Setup & Tooling

  • Bag of Words, TF-IDF, and Text Representation

    Count first, think later. TF-IDF still beats embeddings on well-defined tasks in 2026. The model needs numbers. You have strings. Every NLP pipeline has to answer the same question.

    Phase 05: NLP: Foundations to Advanced

  • GloVe, FastText, and Subword Embeddings

    Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.

    Phase 05: NLP: Foundations to Advanced

  • Tokenizers: BPE, WordPiece, SentencePiece

    Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.

    Phase 10: LLMs from Scratch

  • Chameleon and Early-Fusion Token-Only Multimodal Models

    Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.

    Phase 12: Multimodal AI

  • Mesa-Optimization and Deceptive Alignment

    Hubinger et al. (arXiv:1906.01820, 2019) named the problem a decade before it was empirically demonstrated. When you train a learned optimizer to minimize a base objective, the learned optimizer's…

    Phase 18: Ethics, Safety & Alignment

Covered in Phase 00: Setup & Tooling, Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 18: Ethics, Safety & Alignment.

  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together

Sources

More terms in Data & representations

Open the Data & representations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.