Data & representations · Glossary term
What is Tokenization?
Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
Why does Tokenization matter?
Tokenization determines sequence length, vocabulary boundaries, cost accounting, truncation behavior, and how text or code is represented before embedding.
Tokenization in practice
Use the exact tokenizer for the target model, version it with artifacts, and test multilingual text, code, whitespace, and special tokens.
What is the common confusion about Tokenization?
Tokenization is not always word splitting, and two models can assign different token counts and IDs to the same input.
Learn Tokenization in the course
Lessons that name Tokenization in a title or section
- Text Processing — Tokenization, Stemming, Lemmatization
Language is continuous. Models are discrete. Preprocessing is the bridge. A model cannot read "The cats were running." It reads integers. Every NLP system opens with the same three questions.
- Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece
Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.
- Multilingual NLP
One model, 100+ languages, zero training data for most of them. Cross-lingual transfer is the practical miracle of the 2020s. English has billions of labeled examples. Urdu has thousands.
- Embodied VLAs: RT-2, OpenVLA, π0, GR00T
The first time a model read a recipe off a website and executed it in a kitchen robot was RT-2 (Google DeepMind, July 2023).
Covered in Phase 05: NLP: Foundations to Advanced and Phase 12: Multimodal AI.
Related terms
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
Sources
More terms in Data & representations
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.