Multimodal systems · Glossary term

What is Automatic Speech Recognition (ASR)?

The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.

Why does Automatic Speech Recognition (ASR) matter?

Speech interfaces depend on more than language modeling. Acoustic variation, segmentation, decoding, vocabulary, and domain conditions all affect the final transcript.

Automatic Speech Recognition (ASR) in practice

Evaluate word or character errors by language, speaker, noise, and domain, retain timestamps when downstream grounding needs them, and test the exact audio preprocessing used in production.

What is the common confusion about Automatic Speech Recognition (ASR)?

ASR transcribes what was said. Determining who spoke requires diarization or speaker recognition, while translation and intent understanding are separate tasks.

Learn Automatic Speech Recognition (ASR) in the course

Start with

  • Speech Recognition (ASR) — CTC, RNN-T, Attention

    Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.

    Phase 06: Speech & Audio

Lessons that name Automatic Speech Recognition (ASR) in a title or section

  • Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS)

    A voice agent that feels right has end-to-end latency under 800ms, knows when you have stopped talking, handles barge-in, and can call a tool without stalling.

    Phase 19: Capstone Projects

  • Real-Time Audio Processing

    Batch pipelines process a file. Real-time pipelines process the next 20 milliseconds before the next 20 arrive. Every conversational AI, broadcast studio, and telephony bot lives and dies by this…

    Phase 06: Speech & Audio

  • Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards

    You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…

    Phase 06: Speech & Audio

  • Many-Shot Jailbreaking

    Anil, Durmus, Panickssery, Sharma, et al. (Anthropic, NeurIPS 2024). Many-shot jailbreaking (MSJ) exploits long context windows: stuff hundreds of faux user-assistant turns where the assistant…

    Phase 18: Ethics, Safety & Alignment

Taught in Phase 06: Speech & Audio.

Also covered in Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.

  • Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…

Sources

More terms in Multimodal systems

Open the Multimodal systems list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.