Phase 12 · Multimodal AI
Learn Multimodal AI from Scratch: 25 Free Lessons
Models that see, hear, read, and reason across modalities.
- 25 lessons
- 16 build
- 9 learn
- ~67 hours
- Python
Start Phase 12
First lesson Vision Transformers and the Patch-Token Primitive
Run this command from the repository root:
python3 phases/12-multimodal-ai/01-vision-transformer-patch-tokens/code/main.pyKeep the command, exit code, patch grid and sequence lengths, parameter counts, and one sentence explaining why higher resolution creates more visual tokens.
All 25 lessons in Phase 12
- Vision Transformers and the Patch-Token Primitive
Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.
- CLIP and Contrastive Vision-Language Pretraining
OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a…
- From CLIP to BLIP-2 — Q-Former as Modality Bridge
CLIP aligns image and text but cannot generate captions, answer questions, or hold a conversation. BLIP-2 (Salesforce, 2023) solved that with a small trainable bridge: 32 learnable query vectors…
- Flamingo and Gated Cross-Attention for Few-Shot VLMs
DeepMind's Flamingo (2022) did two things before anyone else. It showed a single model could process arbitrarily interleaved sequences of images, videos, and text.
- LLaVA and Visual Instruction Tuning
LLaVA (April 2023) is the most copied multimodal architecture on the planet. It replaced BLIP-2's Q-Former with a 2-layer MLP, replaced Flamingo's gated cross-attention with naive token…
- Any-Resolution Vision: Patch-n'-Pack and NaFlex
Real images are not 224x224 squares. A receipt is 9:16, a chart is 16:9, a medical scan might be 4096x4096, a mobile screenshot is 9:19.5.
- Open-Weight VLM Recipes: What Actually Matters
The 2024-2026 open-weight VLM literature is a forest of ablation tables. Apple's MM1 tested 13 combinations of image encoder, connector, and data mix.
- LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model
Before LLaVA-OneVision (Li et al., August 2024) the open-VLM world had separate lineages: LLaVA-1.5 for single images, multi-image models like Mantis and VILA, video models like Video-LLaVA and…
- Qwen-VL Family and Dynamic-FPS Video
The Qwen-VL family — Qwen-VL (2023), Qwen2-VL (2024), Qwen2.5-VL (2025), Qwen3-VL (2025) — is the most influential open vision-language model lineage in 2026.
- InternVL3: Native Multimodal Pretraining
Every open VLM before InternVL3 followed the same three-step recipe: take a text LLM trained on trillions of text tokens, bolt on a vision encoder, then fine-tune the seams.
- Chameleon and Early-Fusion Token-Only Multimodal Models
Every VLM we have seen so far keeps images and text separate. Visual tokens come from a vision encoder, flow into a projector, then meet text inside the LLM.
- Emu3: Next-Token Prediction for Image and Video Generation
BAAI's Emu3 (Wang et al., September 2024) is the 2024 result that should have ended the diffusion-versus-autoregressive debate.
- Transfusion: Autoregressive Text + Diffusion Image in One Transformer
Chameleon and Emu3 bet everything on discrete tokens. They work, but the quantization bottleneck is visible — the image quality plateaus below continuous-space diffusion models.
- Show-o and Discrete-Diffusion Unified Models
Transfusion mixes continuous and discrete representations. Show-o (Xie et al., August 2024) goes the other way: text tokens use causal next-token prediction, image tokens use masked discrete…
- Janus-Pro: Decoupled Encoders for Unified Multimodal Models
Unified multimodal models have an unavoidable tension. Understanding wants semantic features — SigLIP or DINOv2 output vectors rich with concept-level information.
- MIO and Any-to-Any Streaming Multimodal Models
GPT-4o ships a product most open models cannot replicate: an agent that hears voice, sees video, and speaks back in real time.
- Video-Language Models: Temporal Tokens and Grounding
Video is not a stack of photos. A 5-second clip has causal ordering, action verbs, and event timing that an image model cannot represent.
- Long-Video Understanding at Million-Token Context
A 1-hour 4K video at 24 FPS, patched and embedded, produces on the order of 60 million tokens. A 2-hour podcast episode transcribed is 30,000 tokens.
- Audio-Language Models: the Whisper to Audio Flamingo 3 Arc
Whisper (Radford et al., December 2022) settled speech recognition — 680k hours of weakly-supervised multilingual speech, a simple encoder-decoder transformer, a benchmark that made every subsequent…
- Omni Models: Qwen2.5-Omni and the Thinker-Talker Split
GPT-4o's product demo in May 2024 was disruptive not because of the underlying model but because of the product shape — a voice interface where you talk, the model sees what the camera sees, and it…
- Embodied VLAs: RT-2, OpenVLA, π0, GR00T
The first time a model read a recipe off a website and executed it in a kitchen robot was RT-2 (Google DeepMind, July 2023).
- Document and Diagram Understanding
Documents are not photos. A PDF, scientific paper, invoice, or handwritten form has layout, tables, diagrams, footnotes, headers, and semantic structure that plain image understanding cannot capture.
- ColPali and Vision-Native Document RAG
Traditional RAG parses PDFs into text, splits into chunks, embeds chunks, stores vectors. Every step loses signal: OCR drops chart data, chunking breaks table rows, text embeddings ignore figures.
- Multimodal RAG and Cross-Modal Retrieval
Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with…
- Multimodal Agents and Computer-Use (Capstone)
The 2026 frontier product is a multimodal agent that reads screenshots, clicks buttons, navigates web UIs, fills forms, and completes workflows end-to-end.
Glossary terms in this phase
- AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
- DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.
- Early FusionCombining raw or low-level representations from several modalities before most task-specific modeling occurs.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
- GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
- Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- PatchA reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
- Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
- RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
- Shared Embedding SpaceA common vector space in which representations from different modalities can be compared with the same similarity function.
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
- Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
- Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.
Frequently asked questions
How many lessons are in Phase 12: Multimodal AI?
Phase 12 has 25 lessons: 16 Build lessons and 9 Learn lessons. The lesson code uses Python.
What should I know before I start Phase 12?
The phase guide gives these prerequisites: Phase 7 Transformers and Phase 4 Computer Vision. In the course roadmap, this phase builds on Phase 10: LLMs from Scratch.
Is Phase 12 free?
Yes. All 25 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 12 take?
The time estimates of all 25 lessons add up to about 67 hours.