Multimodal systems · Glossary term
What is Automatic Speech Recognition (ASR)?
The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.
Why does Automatic Speech Recognition (ASR) matter?
Speech interfaces depend on more than language modeling. Acoustic variation, segmentation, decoding, vocabulary, and domain conditions all affect the final transcript.
Automatic Speech Recognition (ASR) in practice
Evaluate word or character errors by language, speaker, noise, and domain, retain timestamps when downstream grounding needs them, and test the exact audio preprocessing used in production.
What is the common confusion about Automatic Speech Recognition (ASR)?
ASR transcribes what was said. Determining who spoke requires diarization or speaker recognition, while translation and intent understanding are separate tasks.
Learn Automatic Speech Recognition (ASR) in the course
Start with
- Speech Recognition (ASR) — CTC, RNN-T, Attention
Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.
Lessons that name Automatic Speech Recognition (ASR) in a title or section
- Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS)
A voice agent that feels right has end-to-end latency under 800ms, knows when you have stopped talking, handles barge-in, and can call a tool without stalling.
- Real-Time Audio Processing
Batch pipelines process a file. Real-time pipelines process the next 20 milliseconds before the next 20 arrive. Every conversational AI, broadcast studio, and telephony bot lives and dies by this…
- Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards
You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…
- Many-Shot Jailbreaking
Anil, Durmus, Panickssery, Sharma, et al. (Anthropic, NeurIPS 2024). Many-shot jailbreaking (MSJ) exploits long context windows: stuff hundreds of faux user-assistant turns where the assistant…
Taught in Phase 06: Speech & Audio.
Also covered in Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
- EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
Sources
More terms in Multimodal systems
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.