Phase 03 · Deep Learning Core

Learn Deep Learning from Scratch: 13 Free Lessons

Neural networks from first principles. No frameworks until you build one yourself.

  • 13 lessons
  • ~19 hours
  • Python

Start Phase 03

First lesson The Perceptron

Run this command from the repository root:

python3 phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py

Keep the command, exit code, converged gate results, XOR failure of a single perceptron, and the final two-layer XOR predictions.

All 13 lessons in Phase 03

  1. The Perceptron

    The perceptron is the atom of neural networks. Split it open and you find weights, a bias, and a decision. Implement a perceptron from scratch in Python, including the weight update rule and step…

    Build · Python · ~60 min

  2. Multi-Layer Networks and Forward Pass

    One neuron draws a line. Stack them, and you can draw anything. Build a multi-layer network from scratch with Layer and Network classes that perform a complete forward pass.

    Build · Python · ~90 min

  3. Backpropagation from Scratch

    Backpropagation is the algorithm that makes learning possible. Without it, neural networks are just expensive random number generators.

    Build · Python · ~120 min

  4. Activation Functions

    Without nonlinearity, your 100-layer network is a fancy matrix multiply. Activations are the gates that let neural networks think in curves.

    Build · Python · ~75 min

  5. Loss Functions

    Your network makes a prediction. The ground truth says otherwise. How wrong is it? That number is the loss. Pick the wrong loss function and your model optimizes for the wrong thing entirely.

    Build · Python · ~75 min

  6. Optimizers

    Gradient descent tells you which direction to move. It says nothing about how far or how fast. SGD is a compass. Adam is GPS with traffic data.

    Build · Python · ~75 min

  7. Regularization

    Your model gets 99% on training data and 60% on test data. It memorized instead of learning. Regularization is the tax you impose on complexity to force generalization.

    Build · Python · ~75 min

  8. Weight Initialization and Training Stability

    Initialize wrong and training never starts. Initialize right and 50 layers train as smoothly as 3. Implement zero, random, Xavier/Glorot, and Kaiming/He initialization strategies and measure their…

    Build · Python · ~90 min

  9. Learning Rate Schedules and Warmup

    The learning rate is the single most important hyperparameter. Not the architecture. Not the dataset size. Not the activation function. The learning rate. If you tune nothing else, tune this.

    Build · Python · ~90 min

  10. Build Your Own Mini Framework

    You have built neurons, layers, networks, backprop, activations, loss functions, optimizers, regularization, initialization, and LR schedules. All as separate pieces.

    Build · Python · ~120 min

  11. Introduction to PyTorch

    You built the engine from pistons and crankshafts. Now learn the one everyone actually drives. Build and train neural networks using PyTorch's nn.Module, nn.Sequential, and autograd.

    Build · Python · ~75 min

  12. Introduction to JAX

    PyTorch mutates tensors. TensorFlow builds graphs. JAX compiles pure functions. That last one changes how you think about deep learning.

    Build · Python · ~90 min

  13. Debugging Neural Networks

    Your network compiled. It ran. It produced a number. The number is wrong and nothing crashed. Welcome to the hardest kind of debugging -- the kind where there is no error message.

    Build · Python · ~90 min

Glossary terms in this phase

  • Activation FunctionA function applied after a linear or affine layer that introduces nonlinearity. Without it, composing layers with weights and biases…
  • Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
  • AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
  • AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
  • BackpropagationAn efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph.
  • Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
  • Data AugmentationCreating modified examples, such as transformed images, perturbed audio, or paraphrased text, to increase training diversity without…
  • DropoutDuring training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path.
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
  • HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
  • JAXA Python library for transforming numerical functions with automatic differentiation, compilation, vectorization, and parallel execution…
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
  • Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
  • NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
  • NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
  • OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
  • OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • ReLURectified Linear Unit, defined as `f(x) = max(0, x)`. It is inexpensive and has a non-saturating positive branch, though zero gradients on…
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
  • WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
  • Weight DecayAn update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from…

Frequently asked questions

How many lessons are in Phase 03: Deep Learning Core?

Phase 03 has 13 lessons, all Build lessons. The lesson code uses Python.

What should I know before I start Phase 03?

The phase guide gives these prerequisites: Phase 1 Linear Algebra Intuition. Phase 2 is recommended for model evaluation vocabulary. In the course roadmap, this phase builds on Phase 02: ML Fundamentals.

Is Phase 03 free?

Yes. All 13 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 03 take?

The time estimates of all 13 lessons add up to about 19 hours.

What comes after Phase 03?

Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio and Phase 09: Reinforcement Learning build on this phase.