Multimodal systems · Glossary term

What is Audio Token?

A discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several codebooks.

Why does Audio Token matter?

Discrete audio representations let sequence models process, predict, store, or generate sound using token-oriented architectures.

Audio Token in practice

Version the codec with the model, preserve sample-rate and codebook metadata, measure reconstruction quality, and distinguish semantic audio tokens from waveform-compression tokens.

What is the common confusion about Audio Token?

An audio token is not a fixed duration, phoneme, or word. Its meaning and time span depend on the tokenizer and codebook design.

Learn Audio Token in the course

Start with

Lessons that name Audio Token in a title or section

  • Audio Generation

    Audio is a 1-D signal at 16-48 kHz. A five-second clip is 80-240k samples. No transformer attends to that sequence directly.

    Phase 08: Generative AI

Taught in Phase 06: Speech & Audio.

Also covered in Phase 08: Generative AI.

  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.