Phase 01 · Math Foundations
Learn Math for Machine Learning: 22 Free Lessons
The intuition behind every AI algorithm, through code, not textbooks.
- 22 lessons
- 16 build
- 6 learn
- ~32 hours
- Python, Julia
Start Phase 01
First lesson Linear Algebra Intuition
Run this command from the repository root:
python3 phases/01-math-foundations/01-linear-algebra-intuition/code/vectors.pyKeep the command, exit code, normalized-vector output, projection residual, and one sentence explaining why a matrix-vector product is a neural-network layer.
All 22 lessons in Phase 01
- Linear Algebra Intuition
Every AI model is just matrix math wearing a fancy hat. Implement vector and matrix operations (addition, dot product, matrix multiply) from scratch in Python.
- Vectors, Matrices & Operations
Every neural network is just matrix multiplication with extra steps. Build a Matrix class with element-wise operations, matrix multiplication, transpose, determinant, and inverse.
- Matrix Transformations
A matrix is a machine that reshapes space. Learn what it does to every point, and you understand the whole transformation.
- Calculus for Machine Learning
Derivatives tell you which way is downhill. That is all a neural network needs to learn. Language: Python Compute numerical and analytical derivatives for common ML functions (x^2, sigmoid,…
- Chain Rule & Automatic Differentiation
The chain rule is the engine behind every neural network that learns. Language: Python Build a minimal autograd engine (Value class) that records operations and computes gradients via reverse-mode…
- Probability and Distributions
Probability is the language AI uses to express uncertainty. Language: Python Implement PMFs and PDFs from scratch for Bernoulli, categorical, Poisson, uniform, and normal distributions.
- Bayes' Theorem
Probability is about what you expect. Bayes' theorem is about what you learn. Language: Python Apply Bayes' theorem to compute posterior probabilities from priors, likelihoods, and evidence.
- Optimization
Training a neural network is nothing more than finding the bottom of a valley. Language: Python Implement vanilla gradient descent, SGD with momentum, and Adam from scratch.
- Information Theory
Information theory measures surprise. Loss functions are built on it. Language: Python Compute entropy, cross-entropy, and KL divergence from scratch and explain their relationship.
- Dimensionality Reduction
High-dimensional data has structure. You find it by looking from the right angle. Language: Python Implement PCA from scratch: center data, compute the covariance matrix, eigendecompose, and project.
- Singular Value Decomposition
SVD is the Swiss Army knife of linear algebra. Every matrix has one. Every data scientist needs one. Implement SVD via power iteration and explain the geometric meaning of U, Sigma, and V^T.
- Tensor Operations
Tensors are the common language between data and deep learning. Every image, every sentence, every gradient flows through them.
- Numerical Stability
Floating point is a leaky abstraction. It will bite you during training, and you will not see it coming. Language: Python Implement numerically stable softmax and log-sum-exp using the…
- Norms and Distances
Your distance function defines what "similar" means. Choose wrong and everything downstream breaks. Language: Python Implement L1, L2, cosine, Mahalanobis, Jaccard, and edit distance functions from…
- Statistics for Machine Learning
Statistics is how you know if your model actually works or just got lucky. Language: Python Compute descriptive statistics, Pearson/Spearman correlation, and covariance matrices from scratch.
- Sampling Methods
Sampling is how AI explores the space of possibilities. Language: Python Implement inverse CDF, rejection, and importance sampling from scratch using only uniform random numbers.
- Linear Systems
Solving Ax = b is the oldest problem in mathematics that still runs your neural network. Language: Python Solve Ax = b using Gaussian elimination with partial pivoting and back substitution.
- Convex Optimization
Convex problems have one valley. Neural networks have millions. Knowing the difference matters. Language: Python Test whether a function is convex using the definition, second derivative, and…
- Complex Numbers for AI
The square root of -1 is not imaginary. It is the key to rotations, frequencies, and half of signal processing. Language: Python Perform complex arithmetic (add, multiply, divide, conjugate) in both…
- The Fourier Transform
Every signal is a sum of sine waves. The Fourier transform tells you which ones. Language: Python Implement the DFT from scratch and verify it against the O(N log N) Cooley-Tukey FFT.
- Graph Theory for Machine Learning
Graphs are the data structure of relationships. If your data has connections, you need graph theory. Language: Python Build a graph class with adjacency matrix/list representations and implement BFS…
- Stochastic Processes
Randomness with structure. The math behind random walks, Markov chains, and diffusion models. Language: Python Simulate 1D and 2D random walks and verify the sqrt(n) scaling of displacement.
Glossary terms in this phase
- Adam (Optimizer)Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- AutogradA system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation.
- CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
- Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
- EigenvalueA scalar that describes how a linear transformation scales a corresponding nonzero eigenvector without changing its direction.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- Learning Rate ScheduleA policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- NormalizationA family of transformations that rescale or recenter inputs, activations, or features using defined statistics.
- Nucleus Sampling (Top-p)A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- Stochastic Gradient Descent (SGD)An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training…
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- Top-k SamplingA decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and…
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…
Frequently asked questions
How many lessons are in Phase 01: Math Foundations?
Phase 01 has 22 lessons: 16 Build lessons and 6 Learn lessons. The lesson code uses Python and Julia.
What should I know before I start Phase 01?
The phase guide gives these prerequisites: Complete Phase 0, or confirm that Python 3.11+ and Git work from the repository root. In the course roadmap, this phase builds on Phase 00: Setup & Tooling.
Is Phase 01 free?
Yes. All 22 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 01 take?
The time estimates of all 22 lessons add up to about 32 hours.
What comes after Phase 01?
Phase 02: ML Fundamentals builds on this phase.