Data & representations · Glossary term
What is Token?
An integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation. A token can be a whole word, part of a word, punctuation, whitespace, a byte sequence, or a special control symbol.
“A word-sized piece of model input or output.”
What is the common confusion about Token?
Character-to-token ratios vary by language, content, and tokenizer, so count with the target model's tokenizer or provider tools.
Learn Token in the course
Start with
- Tokenizers: BPE, WordPiece, SentencePiece
Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.
Lessons that name Token in a title or section
- Video-Language Models: Temporal Tokens and Grounding
Video is not a stack of photos. A 5-second clip has causal ordering, action verbs, and event timing that an image model cannot represent.
- MCP Auth in Production: Issuer-Bound Enrollment and Tokens
Lesson 16 built the OAuth 2.1 state machine. This lesson hardens its production boundaries for MCP 2026-07-28: Client ID Metadata Documents first, deprecated dynamic registration only for…
- Kill Switches, Circuit Breakers, and Canary Tokens
A kill switch is a boolean held outside the agent's edit surface — a Redis key, a feature flag, a signed config — that disables the agent entirely.
- Agent Economies, Token Incentives, Reputation
Long-horizon autonomous agents (METR's 1-hour to 8-hour work-curve) need economic agency. The emerging 5-layer stack is: DePIN (physical compute) → Identity (W3C DIDs + reputation capital) →…
- Token and Positional Embeddings
Ids are integers. The model wants vectors. Two lookup tables sit between them, and the choice of the positional one shapes what the model can learn.
- Vision Transformers (ViT)
Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…
- Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake
2026 music generation: Suno v5 and Udio v4 dominate commercial; MusicGen, Stable Audio Open, and ACE-Step lead open-source. The technical problem is mostly solved.
- Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split
2026 audio generation is almost all tokens. EnCodec, SNAC, Mimi, and DAC turn continuous waveforms into discrete sequences that a transformer can predict.
Taught in Phase 10: LLMs from Scratch.
Also covered in Phase 04: Computer Vision, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 09: Reinforcement Learning, Phase 11: LLM Engineering, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems, Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.
Related terms
- Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
- Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.
- Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
- LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
- ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
- Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- Stop SequenceAn application-specified token or text pattern that causes generation to stop when the decoding system encounters it.
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
More terms in Data & representations
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.