Phase 12 · Multimodal AI

Learn Multimodal AI from Scratch: 25 Free Lessons

Models that see, hear, read, and reason across modalities.

  • 25 lessons
  • 16 build
  • 9 learn
  • ~67 hours
  • Python

Start Phase 12

First lesson Vision Transformers and the Patch-Token Primitive

Run this command from the repository root:

python3 phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.py

Keep the command, exit code, patch grid and sequence lengths, parameter counts, and one sentence explaining why higher resolution creates more visual tokens.

All 25 lessons in Phase 12

  1. Vision Transformers and the Patch-Token Primitive

    Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.

    Learn · Python · ~120 min

  2. CLIP and Contrastive Vision-Language Pretraining

    OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a…

    Build · Python · ~180 min

  3. From CLIP to BLIP-2 — Q-Former as Modality Bridge

    CLIP aligns image and text but cannot generate captions, answer questions, or hold a conversation. BLIP-2 (Salesforce, 2023) solved that with a small trainable bridge: 32 learnable query vectors…

    Build · Python · ~180 min

  4. Flamingo and Gated Cross-Attention for Few-Shot VLMs

    DeepMind's Flamingo (2022) did two things before anyone else. It showed a single model could process arbitrarily interleaved sequences of images, videos, and text.

    Learn · Python · ~120 min

  5. LLaVA and Visual Instruction Tuning

    LLaVA (April 2023) is the most copied multimodal architecture on the planet. It replaced BLIP-2's Q-Former with a 2-layer MLP, replaced Flamingo's gated cross-attention with naive token…

    Build · Python · ~180 min

  6. Any-Resolution Vision: Patch-n'-Pack and NaFlex

    Real images are not 224x224 squares. A receipt is 9:16, a chart is 16:9, a medical scan might be 4096x4096, a mobile screenshot is 9:19.5.

    Build · Python · ~120 min

  7. Open-Weight VLM Recipes: What Actually Matters

    The 2024-2026 open-weight VLM literature is a forest of ablation tables. Apple's MM1 tested 13 combinations of image encoder, connector, and data mix.

    Learn · Python · ~180 min

  8. LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model

    Before LLaVA-OneVision (Li et al., August 2024) the open-VLM world had separate lineages: LLaVA-1.5 for single images, multi-image models like Mantis and VILA, video models like Video-LLaVA and…

    Build · Python · ~180 min

  9. Qwen-VL Family and Dynamic-FPS Video

    The Qwen-VL family — Qwen-VL (2023), Qwen2-VL (2024), Qwen2.5-VL (2025), Qwen3-VL (2025) — is the most influential open vision-language model lineage in 2026.

    Learn · Python · ~120 min

  10. InternVL3: Native Multimodal Pretraining

    Every open VLM before InternVL3 followed the same three-step recipe: take a text LLM trained on trillions of text tokens, bolt on a vision encoder, then fine-tune the seams.

    Learn · Python · ~120 min

  11. Chameleon and Early-Fusion Token-Only Multimodal Models

    Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.

    Build · Python · ~180 min

  12. Emu3: Next-Token Prediction for Image and Video Generation

    BAAI's Emu3 (Wang et al., September 2024) is the 2024 result that should have ended the diffusion-versus-autoregressive debate.

    Learn · Python · ~120 min

  13. Transfusion: Autoregressive Text + Diffusion Image in One Transformer

    Chameleon and Emu3 bet everything on discrete tokens. They work, but the quantization bottleneck is visible — the image quality plateaus below continuous-space diffusion models.

    Build · Python · ~180 min

  14. Show-o and Discrete-Diffusion Unified Models

    Transfusion mixes continuous and discrete representations. Show-o (Xie et al., August 2024) goes the other way: text tokens use causal next-token prediction, image tokens use masked discrete…

    Learn · Python · ~120 min

  15. Janus-Pro: Decoupled Encoders for Unified Multimodal Models

    Unified multimodal models have an unavoidable tension. Understanding wants semantic features — SigLIP or DINOv2 output vectors rich with concept-level information.

    Build · Python · ~120 min

  16. MIO and Any-to-Any Streaming Multimodal Models

    GPT-4o ships a product most open models cannot replicate: an agent that hears voice, sees video, and speaks back in real time.

    Learn · Python · ~120 min

  17. Video-Language Models: Temporal Tokens and Grounding

    Video is not a stack of photos. A 5-second clip has causal ordering, action verbs, and event timing that an image model cannot represent.

    Build · Python · ~180 min

  18. Long-Video Understanding at Million-Token Context

    A 1-hour 4K video at 24 FPS, patched and embedded, produces on the order of 60 million tokens. A 2-hour podcast episode transcribed is 30,000 tokens.

    Build · Python · ~180 min

  19. Audio-Language Models: the Whisper to Audio Flamingo 3 Arc

    Whisper (Radford et al., December 2022) settled speech recognition — 680k hours of weakly-supervised multilingual speech, a simple encoder-decoder transformer, a benchmark that made every subsequent…

    Build · Python · ~180 min

  20. Omni Models: Qwen2.5-Omni and the Thinker-Talker Split

    GPT-4o's product demo in May 2024 was disruptive not because of the underlying model but because of the product shape — a voice interface where you talk, the model sees what the camera sees, and it…

    Build · Python · ~180 min

  21. Embodied VLAs: RT-2, OpenVLA, π0, GR00T

    The first time a model read a recipe off a website and executed it in a kitchen robot was RT-2 (Google DeepMind, July 2023).

    Learn · Python · ~180 min

  22. Document and Diagram Understanding

    Documents are not photos. A PDF, scientific paper, invoice, or handwritten form has layout, tables, diagrams, footnotes, headers, and semantic structure that plain image understanding cannot capture.

    Build · Python · ~180 min

  23. ColPali and Vision-Native Document RAG

    Traditional RAG parses PDFs into text, splits into chunks, embeds chunks, stores vectors. Every step loses signal: OCR drops chart data, chunking breaks table rows, text embeddings ignore figures.

    Build · Python · ~180 min

  24. Multimodal RAG and Cross-Modal Retrieval

    Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with…

    Build · Python · ~180 min

  25. Multimodal Agents and Computer-Use (Capstone)

    The 2026 frontier product is a multimodal agent that reads screenshots, clicks buttons, navigates web UIs, fills forms, and completes workflows end-to-end.

    Build · Python · ~240 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
  • DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.
  • Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
  • GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • PatchA reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
  • Shared Embedding SpaceA common vector space in which representations from different modalities can be compared with the same similarity function.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
  • Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
  • Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.

Frequently asked questions

How many lessons are in Phase 12: Multimodal AI?

Phase 12 has 25 lessons: 16 Build lessons and 9 Learn lessons. The lesson code uses Python.

What should I know before I start Phase 12?

The phase guide gives these prerequisites: Phase 7 Transformers and Phase 4 Computer Vision. In the course roadmap, this phase builds on Phase 10: LLMs from Scratch.

Is Phase 12 free?

Yes. All 25 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 12 take?

The time estimates of all 25 lessons add up to about 67 hours.