Phase 10 · LLMs from Scratch
Build an LLM from Scratch: 24 Free Lessons
Build, train, and understand large language models.
- 24 lessons
- 20 build
- 4 learn
- ~33 hours
- Python, Rust
Start Phase 10
First lesson Tokenizers: BPE, WordPiece, SentencePiece
Run this command from the repository root:
python3 phases/10-llms-from-scratch/01-tokenizers/code/main.pyKeep the command, exit code, encode/decode round-trip results, learned merge count, and compression ratios. tiktoken is an optional comparison.
All 24 lessons in Phase 10
- Tokenizers: BPE, WordPiece, SentencePiece
Your LLM does not read English. It reads integers. The tokenizer decides whether those integers carry meaning or waste it.
- Building a Tokenizer from Scratch
Lesson 01 gave you a toy. This lesson gives you a weapon. Build a production-grade BPE tokenizer that handles Unicode, whitespace normalization, and special tokens.
- Data Pipelines for Pre-Training
The model is a mirror. It reflects whatever data you feed it. Feed it garbage, it reflects garbage with perfect fluency.
- Pre-Training a Mini GPT (124M Parameters)
GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.
- Scaling: Distributed Training, FSDP, DeepSpeed
Your 124M model trained on one GPU. Now try 7 billion parameters. The model doesn't fit in memory. The data takes weeks on a single machine. Distributed training isn't optional at scale.
- Instruction Tuning (SFT)
A base model predicts the next token. That's it. It doesn't follow instructions, answer questions, or refuse harmful requests. SFT is the bridge between a token predictor and a useful assistant.
- RLHF: Reward Model + PPO
SFT teaches the model to follow instructions. But it doesn't teach the model which response is BETTER. Two grammatically correct, factually accurate answers can differ enormously in helpfulness.
- DPO: Direct Preference Optimization
RLHF works. It also requires training three models (SFT, reward model, policy), managing PPO's instability, and tuning a KL penalty. DPO asks: what if you could skip all of that?
- Constitutional AI and Self-Improvement
RLHF needs humans in the loop. Constitutional AI replaces most of them with the model itself. Write a list of principles, have the model critique its own outputs against those principles, and train…
- Evaluation: Benchmarks, Evals, LM Harness
Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Every frontier lab games benchmarks. MMLU scores go up while models still can't reliably count the number of R's in…
- Quantization: Making Models Fit
A 70B model in FP16 needs 140GB. Two A100s just for weights. Quantize to FP8: one 80GB GPU. INT4: a MacBook. Implement symmetric and asymmetric quantization from FP16 to INT8 and INT4, including…
- Inference Optimization
Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.
- Building a Complete LLM Pipeline
Everything from Lessons 01 to 12 is one stage of one pipeline. This lesson is the scaffold that turns those stages into a single end-to-end run: tokenize, pre-train, scale, SFT, align, evaluate,…
- Open Models: Architecture Walkthroughs
You built a GPT-2 Small from scratch in Lesson 04. Frontier open models in 2026 are the same family with five or six concrete changes. RMSNorm instead of LayerNorm. SwiGLU instead of GELU.
- Speculative Decoding and EAGLE-3
Phase 7 · Lesson 16 proved the math: the Leviathan rejection rule preserves the verifier's distribution exactly. This lesson is the training-stack view of 2026 production speculative decoding.
- Differential Attention (V2)
Softmax attention spreads a small amount of probability over every non-matching token. Over 100k tokens that noise adds up and drowns the signal.
- Native Sparse Attention (DeepSeek NSA)
At 64k tokens, attention eats 70-80% of decode latency. Every open-model lab has a plan to fix it. DeepSeek's NSA (ACL 2025 best paper) is the one that stuck: three parallel attention branches —…
- Multi-Token Prediction (MTP)
Every autoregressive LLM from GPT-2 to Llama 3 trains on one loss per position: predict the next token. DeepSeek-V3 added a second loss per position: predict the token after that.
- DualPipe Parallelism
DeepSeek-V3 was trained on 2,048 H800 GPUs with MoE experts scattered across nodes. Cross-node expert all-to-all communication cost 1 GPU-hour of comm for every 1 GPU-hour of compute.
- DeepSeek-V3 Architecture Walkthrough
Phase 10 · Lesson 14 named the six architectural knobs every open model turns. DeepSeek-V3 (December 2024, 671B parameters total, 37B active) turns all six and adds four more: Multi-Head Latent…
- Jamba — Hybrid SSM-Transformer
State space models (SSMs) and transformers want different things. Transformers buy quality via attention at quadratic cost.
- Async and Hogwild! Inference
Speculative decoding (Phase 10 · 15) parallelizes tokens within one sequence. Multi-agent frameworks parallelize across whole sequences but force explicit coordination (voting, sub-task splitting).
- Speculative Decoding and EAGLE
A frontier LLM generating one token requires a full forward pass over billions of parameters. That forward pass is massively over-provisioned: most of the time a much smaller model can guess the…
- Gradient Checkpointing and Activation Recomputation
Backprop keeps every intermediate activation. At 70B parameters and 128K context that is 3 TB of activations per rank. Checkpointing trades FLOPs for memory: recompute instead of save.
Glossary terms in this phase
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
- Cross-EntropyA loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it…
- DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- HyperparameterA configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
- ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
- PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
- Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
- QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
- RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
- RLHF (Reinforcement Learning from Human Feedback)A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal.
- RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
- Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
- SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- Tensor ParallelismPartitioning tensor operations within a model layer across devices, with collective communication combining partial results during the…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
Frequently asked questions
How many lessons are in Phase 10: LLMs from Scratch?
Phase 10 has 24 lessons: 20 Build lessons and 4 Learn lessons. The lesson code uses Python and Rust.
What should I know before I start Phase 10?
The phase guide gives these prerequisites: Phase 5 NLP Foundations. Phase 7 Transformers is strongly recommended before the model-building lessons. In the course roadmap, this phase builds on Phase 07: Transformers Deep Dive.
Is Phase 10 free?
Yes. All 24 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 10 take?
The time estimates of all 24 lessons add up to about 33 hours.
What comes after Phase 10?
Phase 11: LLM Engineering and Phase 12: Multimodal AI build on this phase.