Data & representations · Glossary term

What is Byte Pair Encoding (BPE)?

A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.

Why does Byte Pair Encoding (BPE) matter?

It balances vocabulary size with the ability to represent rare or unseen words as smaller units.

Byte Pair Encoding (BPE) in practice

Train the tokenizer only on approved corpus splits, version its merge rules with the model, and inspect how it segments code, multilingual text, and whitespace.

What is the common confusion about Byte Pair Encoding (BPE)?

BPE is one tokenizer family, not a universal description of how every model creates tokens.

Learn Byte Pair Encoding (BPE) in the course

Lessons that name Byte Pair Encoding (BPE) in a title or section

  • Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece

    Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.

    Phase 05: NLP: Foundations to Advanced

  • Tokenizers: BPE, WordPiece, SentencePiece

    Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.

    Phase 10: LLMs from Scratch

  • BPE Tokenizer From Scratch

    Bytes in, ids out, ids back to the same bytes. Build the tokenizer that every modern text model still starts from. Train a Byte-Pair Encoding vocabulary from a raw text corpus by repeatedly merging…

    Phase 19: Capstone Projects

  • GloVe, FastText, and Subword Embeddings

    Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.

    Phase 05: NLP: Foundations to Advanced

  • Building a Tokenizer from Scratch

    Lesson 01 gave you a toy. This lesson gives you a weapon. Build a production-grade BPE tokenizer that handles Unicode, whitespace normalization, and special tokens.

    Phase 10: LLMs from Scratch

Covered in Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.

  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together

Sources

More terms in Data & representations

Open the Data & representations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.