Multimodal systems · Glossary term
What is Audio Token?
A discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several codebooks.
Why does Audio Token matter?
Discrete audio representations let sequence models process, predict, store, or generate sound using token-oriented architectures.
Audio Token in practice
Version the codec with the model, preserve sample-rate and codebook metadata, measure reconstruction quality, and distinguish semantic audio tokens from waveform-compression tokens.
What is the common confusion about Audio Token?
An audio token is not a fixed duration, phoneme, or word. Its meaning and time span depend on the tokenizer and codebook design.
Learn Audio Token in the course
Start with
- Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split
2026 audio generation is almost all tokens. EnCodec, SNAC, Mimi, and DAC turn continuous waveforms into discrete sequences that a transformer can predict.
Lessons that name Audio Token in a title or section
- Audio Generation
Audio is a 1-D signal at 16-48 kHz. A five-second clip is 80-240k samples. No transformer attends to that sequence directly.
Taught in Phase 06: Speech & Audio.
Also covered in Phase 08: Generative AI.
Related terms
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
- Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.