Data & representations · Glossary term

What is Tokenization?

Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.

Why does Tokenization matter?

Tokenization determines sequence length, vocabulary boundaries, cost accounting, truncation behavior, and how text or code is represented before embedding.

Tokenization in practice

Use the exact tokenizer for the target model, version it with artifacts, and test multilingual text, code, whitespace, and special tokens.

What is the common confusion about Tokenization?

Tokenization is not always word splitting, and two models can assign different token counts and IDs to the same input.

Learn Tokenization in the course

Lessons that name Tokenization in a title or section

  • Text Processing — Tokenization, Stemming, Lemmatization

    Language is continuous. Models are discrete. Preprocessing is the bridge. A model cannot read "The cats were running." It reads integers. Every NLP system opens with the same three questions.

    Phase 05: NLP: Foundations to Advanced

  • Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece

    Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.

    Phase 05: NLP: Foundations to Advanced

  • Multilingual NLP

    One model, 100+ languages, zero training data for most of them. Cross-lingual transfer is the practical miracle of the 2020s. English has billions of labeled examples. Urdu has thousands.

    Phase 05: NLP: Foundations to Advanced

  • Embodied VLAs: RT-2, OpenVLA, π0, GR00T

    The first time a model read a recipe off a website and executed it in a kitchen robot was RT-2 (Google DeepMind, July 2023).

    Phase 12: Multimodal AI

Covered in Phase 05: NLP: Foundations to Advanced and Phase 12: Multimodal AI.

  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
  • Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…

Sources

More terms in Data & representations

Open the Data & representations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.