Phase 02 · ML Fundamentals
Learn Machine Learning from Scratch: 18 Free Lessons
Classical machine learning is still the backbone of most production AI.
- 18 lessons
- 16 build
- 2 learn
- ~26 hours
- Python
Start Phase 02
First lesson What Is Machine Learning
Run this command from the repository root:
python3 phases/02-ml-fundamentals/01-what-is-machine-learning/code/ml_intro.pyKeep the command, exit code, test accuracy, random baseline, and one sentence explaining why the learned classifier beats that baseline.
All 18 lessons in Phase 02
- What Is Machine Learning
Machine learning is teaching computers to find patterns in data instead of writing rules by hand. Explain the difference between supervised, unsupervised, and reinforcement learning and identify…
- Linear Regression
Linear regression draws the best straight line through your data. It is the "hello world" of machine learning. Derive the gradient descent update rules for mean squared error and implement linear…
- Logistic Regression
Logistic regression bends a straight line into an S-curve to answer yes-or-no questions with probabilities. Implement logistic regression from scratch using the sigmoid function and binary…
- Decision Trees and Random Forests
A decision tree is just a flowchart. But a forest of them is one of the most powerful tools in ML. Language: Python Implement Gini impurity, entropy, and information gain calculations to find…
- Support Vector Machines
Find the widest street between two classes. That is the entire idea. Language: Python Implement a linear SVM from scratch using hinge loss and gradient descent on the primal formulation.
- K-Nearest Neighbors and Distances
Store everything. Predict by looking at your neighbors. The simplest algorithm that actually works. Language: Python Implement KNN classification and regression from scratch with configurable K and…
- Unsupervised Learning
No labels, no teacher. The algorithm finds structure on its own. Implement K-Means, DBSCAN, and Gaussian Mixture Models from scratch and compare their clustering behavior.
- Feature Engineering & Selection
A good feature is worth a thousand data points. Implement numerical transforms (standardization, min-max scaling, log transform, binning) and explain when each is appropriate.
- Model Evaluation
A model is only as good as the way you measure it. Implement K-fold and stratified K-fold cross-validation from scratch and explain why stratification matters for imbalanced data.
- Bias-Variance Tradeoff
Every model error comes from one of three sources: bias, variance, or noise. You can only control the first two. Language: Python Derive the bias-variance decomposition of expected prediction error…
- Ensemble Methods
A group of weak learners, combined correctly, becomes a strong learner. This is not a metaphor. It is a theorem. Language: Python Implement AdaBoost and gradient boosting from scratch and explain…
- Hyperparameter Tuning
Hyperparameters are the knobs you turn before training starts. Turning them well is the difference between a mediocre model and a great one.
- ML Pipelines
A model is not a product. A pipeline is. The pipeline is everything from raw data to deployed prediction, and every step must be reproducible.
- Naive Bayes
The "naive" assumption is wrong, and it works anyway. That's the beauty of it. Language: Python Implement Multinomial Naive Bayes from scratch with Laplace smoothing for text classification.
- Time Series Fundamentals
Past performance does predict future results -- if you check for stationarity first. Language: Python Decompose a time series into trend, seasonality, and residual components and test for…
- Anomaly Detection
Normal is easy to define. Abnormal is whatever doesn't fit. Language: Python Implement Z-score, IQR, and Isolation Forest anomaly detection methods from scratch.
- Handling Imbalanced Data
When 99% of your data is "normal," accuracy is a lie. Language: Python Implement SMOTE from scratch and explain how synthetic oversampling differs from random duplication.
- Feature Selection
More features is not better. The right features is better. Language: Python Implement filter methods (variance threshold, mutual information, chi-squared) and wrapper methods (RFE, forward…
Glossary terms in this phase
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- Data LeakageUnintended use of information during training or feature construction that would not be available at the real prediction point or belongs…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- Gradient DescentA family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather…
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- Learning RateA scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that…
- Loss FunctionAn objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce.
- OverfittingA generalization gap in which performance on training data is substantially better than performance on representative unseen data.
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- UnderfittingA model or training setup has insufficient effective capacity, optimization, features, or training signal to capture useful patterns in…
- WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
Frequently asked questions
How many lessons are in Phase 02: ML Fundamentals?
Phase 02 has 18 lessons: 16 Build lessons and 2 Learn lessons. The lesson code uses Python.
What should I know before I start Phase 02?
The phase guide gives these prerequisites: Phase 1 Math Foundations and NumPy. Check the route with python3 phases/00-setup-and-tooling/01-dev-environment/code/verify.py --route ml-foundations. In the course roadmap, this phase builds on Phase 01: Math Foundations.
Is Phase 02 free?
Yes. All 18 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 02 take?
The time estimates of all 18 lessons add up to about 26 hours.
What comes after Phase 02?
Phase 03: Deep Learning Core builds on this phase.