Phase 01 · Math Foundations

Learn Math for Machine Learning: 22 Free Lessons

The intuition behind every AI algorithm, through code, not textbooks.

  • 22 lessons
  • 16 build
  • 6 learn
  • ~32 hours
  • Python, Julia

Start Phase 01

First lesson Linear Algebra Intuition

Run this command from the repository root:

python3 phases/01-math-foundations/01-linear-algebra-intuition/code/vectors.py

Keep the command, exit code, normalized-vector output, projection residual, and one sentence explaining why a matrix-vector product is a neural-network layer.

All 22 lessons in Phase 01

  1. Linear Algebra Intuition

    Every AI model is just matrix math wearing a fancy hat. Implement vector and matrix operations (addition, dot product, matrix multiply) from scratch in Python.

    Learn · Python, Julia · ~60 min

  2. Vectors, Matrices & Operations

    Every neural network is just matrix multiplication with extra steps. Build a Matrix class with element-wise operations, matrix multiplication, transpose, determinant, and inverse.

    Build · Python, Julia · ~60 min

  3. Matrix Transformations

    A matrix is a machine that reshapes space. Learn what it does to every point, and you understand the whole transformation.

    Build · Python, Julia · ~75 min

  4. Calculus for Machine Learning

    Derivatives tell you which way is downhill. That is all a neural network needs to learn. Language: Python Compute numerical and analytical derivatives for common ML functions (x^2, sigmoid,…

    Learn · Python · ~60 min

  5. Chain Rule & Automatic Differentiation

    The chain rule is the engine behind every neural network that learns. Language: Python Build a minimal autograd engine (Value class) that records operations and computes gradients via reverse-mode…

    Build · Python · ~90 min

  6. Probability and Distributions

    Probability is the language AI uses to express uncertainty. Language: Python Implement PMFs and PDFs from scratch for Bernoulli, categorical, Poisson, uniform, and normal distributions.

    Learn · Python · ~75 min

  7. Bayes' Theorem

    Probability is about what you expect. Bayes' theorem is about what you learn. Language: Python Apply Bayes' theorem to compute posterior probabilities from priors, likelihoods, and evidence.

    Build · Python · ~75 min

  8. Optimization

    Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.

    Build · Python · ~75 min

  9. Information Theory

    Information theory measures surprise. Loss functions are built on it. Language: Python Compute entropy, cross-entropy, and KL divergence from scratch and explain their relationship.

    Learn · Python · ~60 min

  10. Dimensionality Reduction

    High-dimensional data has structure. You find it by looking from the right angle. Language: Python Implement PCA from scratch: center data, compute the covariance matrix, eigendecompose, and project.

    Build · Python · ~90 min

  11. Singular Value Decomposition

    SVD is the Swiss Army knife of linear algebra. Every matrix has one. Every data scientist needs one. Implement SVD via power iteration and explain the geometric meaning of U, Sigma, and V^T.

    Build · Python, Julia · ~120 min

  12. Tensor Operations

    Tensors are the common language between data and deep learning. Every image, every sentence, every gradient flows through them.

    Build · Python · ~90 min

  13. Numerical Stability

    Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…

    Build · Python · ~120 min

  14. Norms and Distances

    Your distance function defines what "similar" means. Choose wrong and everything downstream breaks. Language: Python Implement L1, L2, cosine, Mahalanobis, Jaccard, and edit distance functions from…

    Build · Python · ~90 min

  15. Statistics for Machine Learning

    Statistics is how you know if your model actually works or just got lucky. Language: Python Compute descriptive statistics, Pearson/Spearman correlation, and covariance matrices from scratch.

    Build · Python · ~120 min

  16. Sampling Methods

    Sampling is how AI explores the space of possibilities. Language: Python Implement inverse CDF, rejection, and importance sampling from scratch using only uniform random numbers.

    Build · Python · ~120 min

  17. Linear Systems

    Solving Ax = b is the oldest problem in mathematics that still runs your neural network. Language: Python Solve Ax = b using Gaussian elimination with partial pivoting and back substitution.

    Build · Python · ~120 min

  18. Convex Optimization

    Convex problems have one valley. Neural networks have millions. Knowing the difference matters. Language: Python Test whether a function is convex using the definition, second derivative, and…

    Build · Python · ~90 min

  19. Complex Numbers for AI

    The square root of -1 is not imaginary. It is the key to rotations, frequencies, and half of signal processing. Language: Python Perform complex arithmetic (add, multiply, divide, conjugate) in both…

    Learn · Python · ~60 min

  20. The Fourier Transform

    Every signal is a sum of sine waves. The Fourier transform tells you which ones. Language: Python Implement the DFT from scratch and verify it against the O(N log N) Cooley-Tukey FFT.

    Build · Python · ~90 min

  21. Graph Theory for Machine Learning

    Graphs are the data structure of relationships. If your data has connections, you need graph theory. Language: Python Build a graph class with adjacency matrix/list representations and implement BFS…

    Build · Python · ~90 min

  22. Stochastic Processes

    Randomness with structure. The math behind random walks, Markov chains, and diffusion models. Language: Python Simulate 1D and 2D random walks and verify the sqrt(n) scaling of displacement.

    Learn · Python · ~75 min

Glossary terms in this phase

  • Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
  • CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
  • Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
  • EigenvalueA scalar that describes how a linear transformation scales a corresponding nonzero eigenvector without changing its direction.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
  • Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
  • HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
  • Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
  • Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
  • Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
  • NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
  • NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
  • Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
  • PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
  • Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…

Frequently asked questions

How many lessons are in Phase 01: Math Foundations?

Phase 01 has 22 lessons: 16 Build lessons and 6 Learn lessons. The lesson code uses Python and Julia.

What should I know before I start Phase 01?

The phase guide gives these prerequisites: Complete Phase 0, or confirm that Python 3.11+ and Git work from the repository root. In the course roadmap, this phase builds on Phase 00: Setup & Tooling.

Is Phase 01 free?

Yes. All 22 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 01 take?

The time estimates of all 22 lessons add up to about 32 hours.

What comes after Phase 01?

Phase 02: ML Fundamentals builds on this phase.