Data & representations · Glossary term
What is Byte Pair Encoding (BPE)?
A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
Why does Byte Pair Encoding (BPE) matter?
It balances vocabulary size with the ability to represent rare or unseen words as smaller units.
Byte Pair Encoding (BPE) in practice
Train the tokenizer only on approved corpus splits, version its merge rules with the model, and inspect how it segments code, multilingual text, and whitespace.
What is the common confusion about Byte Pair Encoding (BPE)?
BPE is one tokenizer family, not a universal description of how every model creates tokens.
Learn Byte Pair Encoding (BPE) in the course
Lessons that name Byte Pair Encoding (BPE) in a title or section
- Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece
Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.
- Tokenizers: BPE, WordPiece, SentencePiece
Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.
- BPE Tokenizer From Scratch
Bytes in, ids out, ids back to the same bytes. Build the tokenizer that every modern text model still starts from. Train a Byte-Pair Encoding vocabulary from a raw text corpus by repeatedly merging…
- GloVe, FastText, and Subword Embeddings
Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.
- Building a Tokenizer from Scratch
Lesson 01 gave you a toy. This lesson gives you a weapon. Build a production-grade BPE tokenizer that handles Unicode, whitespace normalization, and special tokens.
Covered in Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.
Related terms
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
Sources
More terms in Data & representations
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.