Data & representations · Glossary term
What is Vocabulary?
The finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
Why does Vocabulary matter?
Vocabulary design affects sequence length, multilingual coverage, code representation, embedding size, and compatibility between tokenizers and model weights.
Vocabulary in practice
Version the vocabulary and special-token assignments with the model, test encode-decode round trips, and never substitute a tokenizer with merely similar token names.
What is the common confusion about Vocabulary?
A model vocabulary is not a dictionary of human words; many entries are fragments, bytes, whitespace patterns, or control symbols.
Learn Vocabulary in the course
Lessons that name Vocabulary in a title or section
- Docker for AI
Containers make "works on my machine" a thing of the past. Build a GPU-enabled Docker image with CUDA, PyTorch, and AI libraries from a Dockerfile.
- Bag of Words, TF-IDF, and Text Representation
Count first, think later. TF-IDF still beats embeddings on well-defined tasks in 2026. The model needs numbers. You have strings. Every NLP pipeline has to answer the same question.
- GloVe, FastText, and Subword Embeddings
Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.
- Tokenizers: BPE, WordPiece, SentencePiece
Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.
- Chameleon and Early-Fusion Token-Only Multimodal Models
Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.
- Mesa-Optimization and Deceptive Alignment
Hubinger et al. (arXiv:1906.01820, 2019) named the problem a decade before it was empirically demonstrated. When you train a learned optimizer to minimize a base objective, the learned optimizer's…
Covered in Phase 00: Setup & Tooling, Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 18: Ethics, Safety & Alignment.
Related terms
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
Sources
More terms in Data & representations
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.