Phase 08 · Generative AI

Learn Generative AI from Scratch: 15 Free Lessons

Create images, video, audio, 3D, and more.

  • 15 lessons
  • 14 build
  • 1 learn
  • ~16 hours
  • Python

Start Phase 08

First lesson Generative Models — Taxonomy & History

Run this command from the repository root:

python3 phases/08-generative-ai/01-generative-models-taxonomy-history/code/main.py

Keep the command, exit code, density estimates, generated samples, and one sentence explaining what an implicit generator cannot answer about p(x).

All 15 lessons in Phase 08

  1. Generative Models — Taxonomy & History

    Every image model, text model, video model, and 3D model fits in one of five buckets. Pick the wrong bucket and you will fight the math for weeks.

    Learn · Python · ~45 min

  2. Autoencoders & Variational Autoencoders (VAE)

    A plain autoencoder compresses then reconstructs. It memorizes. It does not generate. Add one trick — force the code to look Gaussian — and you get a sampler.

    Build · Python · ~75 min

  3. GANs — Generator vs Discriminator

    Goodfellow's trick in 2014 was to skip density entirely. Two networks. One makes fakes. One catches them. They fight until the fakes are indistinguishable from real. It shouldn't work.

    Build · Python · ~75 min

  4. Conditional GANs & Pix2Pix

    The first big unlock of 2014-2017 was controlling what a GAN makes. Attach a label, or an image, or a sentence. Pix2Pix did the image version and it still beats every generic text-to-image model on…

    Build · Python · ~75 min

  5. StyleGAN

    Most generators stir z into every layer at the same time. StyleGAN split it apart: first map z to an intermediate w, then inject w at every resolution level through AdaIN.

    Build · Python · ~45 min

  6. Diffusion Models — DDPM from Scratch

    Ho, Jain, Abbeel (2020) gave the field a recipe it could not quit. Destroy the data with noise over a thousand small steps. Train one neural net to predict the noise. Reverse the process at inference.

    Build · Python · ~75 min

  7. Latent Diffusion & Stable Diffusion

    Pixel-space diffusion on 512×512 images is a computational war crime. Rombach et al. (2022) noticed that you do not need all 786k dimensions to generate an image — you need enough to capture…

    Build · Python · ~75 min

  8. ControlNet, LoRA & Conditioning

    Text alone is a clumsy control signal. ControlNet lets you clone a pretrained diffusion model and steer it with a depth map, pose skeleton, scribble, or edge image.

    Build · Python · ~75 min

  9. Inpainting, Outpainting & Image Editing

    Text-to-image makes new things. Inpainting fixes old ones. In production, 70% of billable image work is editing — swap a background, remove a logo, extend the canvas, regenerate a hand.

    Build · Python · ~75 min

  10. Video Generation

    An image is a 2-D tensor. A video is a 3-D one. The theory is the same; the compute is 10-100x harder. OpenAI's Sora (Feb 2024) proved it was possible.

    Build · Python · ~45 min

  11. Audio Generation

    Audio is a 1-D signal at 16-48 kHz. A five-second clip is 80-240k samples. No transformer attends to that sequence directly.

    Build · Python · ~45 min

  12. 3D Generation

    3D is the modality where 2D-to-3D leverage is strongest. The 2023 breakthrough was 3D Gaussian Splatting. The 2024-2026 generative push layers multi-view diffusion + 3D reconstruction on top to…

    Build · Python · ~45 min

  13. Flow Matching & Rectified Flows

    Diffusion models take 20-50 sampling steps because they walk a curved path from noise to data. Flow matching (Lipman et al., 2023) and rectified flow (Liu et al., 2022) trained straight paths.

    Build · Python · ~45 min

  14. Evaluation — FID, CLIP Score, Human Preference

    Every generative model leaderboard cites FID, CLIP score, and a win rate from a human-preference arena. Each number has a failure mode a determined researcher can game.

    Build · Python · ~45 min

  15. Visual Autoregressive Modeling (VAR): Next-Scale Prediction

    Diffusion models sample iteratively in time (denoising steps). VAR samples iteratively in scale — it predicts a 1x1 token, then 2x2, then 4x4, up to the final resolution, each scale conditioning on…

    Build · Python · ~90 min

Glossary terms in this phase

  • Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • GAN (Generative Adversarial Network)A generator network tries to create realistic data while a discriminator network tries to tell real from fake.
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…

Frequently asked questions

How many lessons are in Phase 08: Generative AI?

Phase 08 has 15 lessons: 14 Build lessons and 1 Learn lesson. The lesson code uses Python.

What should I know before I start Phase 08?

The phase guide gives these prerequisites: Phase 2 ML Fundamentals, Phase 3 Deep Learning Core, and Phase 7 Lesson 14, Build a Transformer from Scratch. In the course roadmap, this phase builds on Phase 07: Transformers Deep Dive.

Is Phase 08 free?

Yes. All 15 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 08 take?

The time estimates of all 15 lessons add up to about 16 hours.