Phase 06 · Speech & Audio

Learn Speech and Audio AI from Scratch: 17 Free Lessons

The other half of human communication. Hear, understand, speak.

  • 17 lessons
  • 13 build
  • 4 learn
  • ~19 hours
  • Python

Start Phase 06

First lesson Audio Fundamentals — Waveforms, Sampling, Fourier Transform

Run this command from the repository root:

python3 phases/06-speech-and-audio/01-audio-fundamentals/code/main.py

Keep the command, exit code, detected frequency peaks, alias frequency, and an explanation of why a low-pass filter must run before downsampling.

All 17 lessons in Phase 06

  1. Audio Fundamentals — Waveforms, Sampling, Fourier Transform

    Waveforms are the raw signal. Spectrograms are the representation. Mel features are the ML-friendly form. Every modern ASR and TTS pipeline walks this ladder, and the first rung is understanding…

    Learn · Python · ~45 min

  2. Spectrograms, Mel Scale & Audio Features

    Neural nets do not consume raw waveforms well. They consume spectrograms. They consume mel spectrograms even better. Every ASR, TTS, and audio classifier in 2026 lives or dies by this single…

    Build · Python · ~45 min

  3. Audio Classification — From k-NN on MFCCs to AST and BEATs

    Everything from "dog barking vs siren" to "which language is this" is audio classification. The features are mels. The architecture moves each decade.

    Build · Python · ~75 min

  4. Speech Recognition (ASR) — CTC, RNN-T, Attention

    Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.

    Build · Python · ~45 min

  5. Whisper — Architecture & Fine-Tuning

    Whisper is a 30-second-window transformer encoder-decoder, trained on 680k hours of multilingual weakly-supervised audio-text pairs. One architecture, multiple tasks, robust across 99 languages.

    Build · Python · ~75 min

  6. Speaker Recognition & Verification

    ASR asks "what did they say?" Speaker recognition asks "who said it?" The math looks the same — embeddings plus cosine — but every production decision hinges on a single EER number.

    Build · Python · ~45 min

  7. Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro

    ASR inverts speech to text; TTS inverts text to speech. The 2026 stack is three parts: text → tokens, tokens → mel, mel → waveform. Each part has a default model that fits in a laptop.

    Build · Python · ~75 min

  8. Voice Cloning & Voice Conversion

    Voice cloning reads your text in someone else's voice. Voice conversion rewrites your voice into someone else's while preserving what you said.

    Build · Python · ~75 min

  9. Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake

    2026 music generation: Suno v5 and Udio v4 dominate commercial; MusicGen, Stable Audio Open, and ACE-Step lead open-source. The technical problem is mostly solved.

    Build · Python · ~75 min

  10. Audio-Language Models — Qwen2.5-Omni, Audio Flamingo, GPT-4o Audio

    2026 audio-language models reason over speech + environmental sound + music. Qwen2.5-Omni-7B matches GPT-4o Audio on MMAU-Pro. Audio Flamingo Next beats Gemini 2.5 Pro on LongAudioBench.

    Build · Python · ~45 min

  11. Real-Time Audio Processing

    Batch pipelines process a file. Real-time pipelines process the next 20 milliseconds before the next 20 arrive. Every conversational AI, broadcast studio, and telephony bot lives and dies by this…

    Build · Python · ~75 min

  12. Build a Voice Assistant Pipeline — The Phase 6 Capstone

    Everything from lessons 01-11, stitched together. Build a voice assistant that listens, reasons, and talks back. In 2026 that is a solved engineering problem, not a research problem — but the…

    Build · Python · ~120 min

  13. Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split

    2026 audio generation is almost all tokens. EnCodec, SNAC, Mimi, and DAC turn continuous waveforms into discrete sequences that a transformer can predict.

    Learn · Python · ~60 min

  14. Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick

    Every voice agent lives or dies on two decisions: is the user speaking now, and are they done? VAD answers the first. Turn-detection (VAD + silence-hangover + semantic endpoint model) answers the…

    Build · Python · ~45 min

  15. Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue

    2024-2026 redefined voice AI. Moshi ships a single model that listens and speaks simultaneously at 200 ms latency. Hibiki does speech-to-speech translation chunk-by-chunk.

    Learn · Python · ~75 min

  16. Voice Anti-Spoofing & Audio Watermarking — ASVspoof 5, AudioSeal, WaveVerify

    Voice cloning shipped faster than defenses. 2026 production voice systems need two things: a detector (AASIST, RawNet2) that classifies real vs fake speech, and a watermark (AudioSeal) that survives…

    Build · Python · ~75 min

  17. Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards

    You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…

    Learn · Python · ~60 min

Glossary terms in this phase

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
  • Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
  • ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
  • CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.

Frequently asked questions

How many lessons are in Phase 06: Speech & Audio?

Phase 06 has 17 lessons: 13 Build lessons and 4 Learn lessons. The lesson code uses Python.

What should I know before I start Phase 06?

The phase guide gives these prerequisites: Phase 1 vectors, matrices, and probability. The first demo uses only the Python standard library. In the course roadmap, this phase builds on Phase 03: Deep Learning Core.

Is Phase 06 free?

Yes. All 17 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 06 take?

The time estimates of all 17 lessons add up to about 19 hours.