Phase 04 · Computer Vision

Learn Computer Vision from Scratch: 28 Free Lessons

From pixels to understanding across image, video, and 3D.

  • 28 lessons
  • 27 build
  • 1 learn
  • ~31 hours
  • Python

Start Phase 04

First lesson Image Fundamentals — Pixels, Channels, Color Spaces

Run this command from the repository root:

python3 phases/04-computer-vision/01-image-fundamentals/code/main.py

Keep the command, exit code, HWC and CHW shapes, normalized channel statistics, round-trip pixel difference, and interpolation roughness. The demo generates a deterministic synthetic image and does not use the network.

All 28 lessons in Phase 04

  1. Image Fundamentals — Pixels, Channels, Color Spaces

    An image is a tensor of light samples. Every vision model you will ever use starts from this one fact. Explain how a continuous scene gets discretized into pixels and why sampling/quantization…

    Learn · Python · ~45 min

  2. Convolutions from Scratch

    A convolution is a tiny dense layer you slide across an image, sharing the same weights at every location. Implement 2D convolution from scratch using only NumPy, including the nested-loop version…

    Build · Python · ~75 min

  3. CNNs — LeNet to ResNet

    Every major CNN of the last thirty years is the same conv–nonlinearity–downsample recipe with one new idea bolted on. Learn the ideas in order.

    Build · Python · ~75 min

  4. Image Classification

    A classifier is a function from pixels to a probability distribution over classes. Everything else is plumbing. Build an end-to-end image classification pipeline on CIFAR-10: dataset, augmentation,…

    Build · Python · ~75 min

  5. Transfer Learning & Fine-Tuning

    Somebody else spent a million GPU hours teaching a network what edges, textures, and object parts look like. You should borrow those features before training your own.

    Build · Python · ~75 min

  6. Object Detection — YOLO from Scratch

    Detection is classification plus regression, run at every position in a feature map, then cleaned up with non-maximum suppression.

    Build · Python · ~75 min

  7. Semantic Segmentation — U-Net

    Segmentation is classification at every pixel. U-Net makes it work by pairing a downsampling encoder with an upsampling decoder and wiring skip connections between them.

    Build · Python · ~75 min

  8. Instance Segmentation — Mask R-CNN

    Add a tiny mask branch to a Faster R-CNN detector and you have instance segmentation. The hard part is RoIAlign, and it is harder than it looks.

    Build · Python · ~75 min

  9. Image Generation — GANs

    A GAN is two neural networks in a fixed game. One draws, one critiques. They get better together until the drawings fool the critic.

    Build · Python · ~75 min

  10. Image Generation — Diffusion Models

    A diffusion model learns to denoise. Train it to remove a tiny bit of noise from a noisy image, repeat that backwards a thousand times, and you have an image generator.

    Build · Python · ~75 min

  11. Stable Diffusion — Architecture & Fine-Tuning

    Stable Diffusion is a DDPM that runs in the latent space of a pretrained VAE, conditioned on text via cross-attention, sampled with a fast deterministic ODE solver, and steered by classifier-free…

    Build · Python · ~75 min

  12. Video Understanding — Temporal Modeling

    A video is a sequence of images plus the physics that connects them. Every video model either treats time as an extra axis (3D conv), a sequence to attend over (transformer), or a feature to extract…

    Build · Python · ~45 min

  13. 3D Vision — Point Clouds & NeRFs

    3D vision comes in two flavours. Point clouds are the sensor's raw output. NeRFs are the learned volumetric field. Both answer "what is where in space." Distinguish explicit (point cloud, mesh,…

    Build · Python · ~45 min

  14. Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Build · Python · ~45 min

  15. Real-Time Vision — Edge Deployment

    Edge inference is the discipline of getting a 90-accuracy model to run at 30 fps on a device with 2 GB of RAM. Every percentage point of accuracy is traded against milliseconds of latency.

    Build · Python · ~75 min

  16. Build a Complete Vision Pipeline — Capstone

    A production vision system is a chain of models and rules stitched with data contracts. The pieces are already in this phase; the capstone wires them together end-to-end.

    Build · Python · ~120 min

  17. Self-Supervised Vision — SimCLR, DINO, MAE

    Labels are the bottleneck of supervised vision. Self-supervised pretraining removes them: learn visual features from 100M unlabelled images, fine-tune on 10k labelled ones.

    Build · Python · ~75 min

  18. Open-Vocabulary Vision — CLIP

    Train an image encoder and a text encoder together so that matching (image, caption) pairs land at the same point in a shared space. That is the whole trick.

    Build · Python · ~45 min

  19. OCR & Document Understanding

    OCR is a three-stage pipeline — detect text boxes, recognise the characters, then lay them out. Every modern OCR system reorders these stages or merges them.

    Build · Python · ~45 min

  20. Image Retrieval & Metric Learning

    A retrieval system ranks candidates by a distance in embedding space. Metric learning is the discipline of shaping that space so the distances mean what you want.

    Build · Python · ~45 min

  21. Keypoint Detection & Pose Estimation

    A pose is a set of ordered keypoints. A keypoint detector is a heatmap regressor. Everything else is bookkeeping. Distinguish top-down and bottom-up pose estimation and state when each is used.

    Build · Python · ~45 min

  22. 3D Gaussian Splatting from Scratch

    A scene is a cloud of millions of 3D Gaussians. Each one has a position, orientation, scale, opacity, and a colour that depends on viewing direction.

    Build · Python · ~90 min

  23. Diffusion Transformers & Rectified Flow

    The U-Net is not the secret of diffusion. Replace it with a transformer, swap the noise schedule for a straight-line flow, and suddenly you have SD3, FLUX, and every 2026 text-to-image model.

    Build · Python · ~75 min

  24. SAM 3 & Open-Vocabulary Segmentation

    Give a model a text prompt and an image and get masks for every matching object. SAM 3 made that a single forward pass. Distinguish SAM (visual prompts only), Grounded SAM / SAM 2 (detector + SAM),…

    Build · Python · ~60 min

  25. Vision-Language Models — The ViT-MLP-LLM Pattern

    A vision encoder converts an image into tokens. An MLP projector maps those tokens into the LLM's embedding space. A language model does the rest.

    Build · Python · ~75 min

  26. Monocular Depth & Geometry Estimation

    A depth map is a single-channel image where each pixel is a distance from the camera. Predicting it from one RGB frame used to be impossible without stereo or LiDAR.

    Build · Python · ~60 min

  27. Multi-Object Tracking & Video Memory

    Tracking is detection plus association. Detect every frame. Match this frame's detections to last frame's tracks by ID. Distinguish tracking-by-detection from query-based tracking and name the…

    Build · Python · ~60 min

  28. World Models & Video Diffusion

    A video model that predicts the next seconds of a scene is a world simulator. Condition that prediction on actions and you have a learned game engine.

    Build · Python · ~75 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
  • Contrastive LearningTraining by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space.
  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • GAN (Generative Adversarial Network)A generator network tries to create realistic data while a discriminator network tries to tell real from fake.
  • Image TokenA model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • Latent SpaceA learned representation space whose coordinates encode factors useful to a model. It may be lower-dimensional than the input, but…
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • LogitsThe model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into…
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • PatchA reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • QLoRAA parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training…
  • Recall@KFor one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`.
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • Transfer LearningStarting from representations or parameters learned on one data distribution or objective and adapting them for another.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
  • Vision-Language Model (VLM)A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval,…
  • Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.

Frequently asked questions

How many lessons are in Phase 04: Computer Vision?

Phase 04 has 28 lessons: 27 Build lessons and 1 Learn lesson. The lesson code uses Python.

What should I know before I start Phase 04?

The phase guide gives these prerequisites: Phase 1 Lesson 12, Tensor Operations, and Phase 3 Lesson 11, Introduction to PyTorch. The first demo needs only NumPy. In the course roadmap, this phase builds on Phase 03: Deep Learning Core.

Is Phase 04 free?

Yes. All 28 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 04 take?

The time estimates of all 28 lessons add up to about 31 hours.