Phase 07 · Transformers Deep Dive

Learn Transformers from Scratch: 16 Free Lessons

The architecture that changed everything. Understand every layer.

  • 16 lessons
  • 12 build
  • 4 learn
  • ~17 hours
  • Python

Start Phase 07

First lesson Why Transformers — The Problems with RNNs

Run this command from the repository root:

python3 phases/07-transformers-deep-dive/01-why-transformers/code/main.py

Keep the command, exit code, serial and parallel depth table, equivalence check, and one sentence describing the speed-versus-memory tradeoff.

All 16 lessons in Phase 07

  1. Why Transformers — The Problems with RNNs

    RNNs process tokens one at a time. Transformers process all tokens at once. That single architectural bet changed every scaling curve in deep learning after 2017.

    Learn · Python · ~45 min

  2. Self-Attention from Scratch

    Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…

    Build · Python · ~90 min

  3. Multi-Head Attention

    One attention head learns one relation at a time. Eight heads learn eight. Heads are free. Take more of them. A single self-attention head computes one attention matrix.

    Build · Python · ~75 min

  4. Positional Encoding — Sinusoidal, RoPE, ALiBi

    Attention is permutation-invariant. "The cat sat on the mat" and "mat the on sat cat the" produce the same output without positional signal.

    Build · Python · ~45 min

  5. The Full Transformer — Encoder + Decoder

    Attention is the star. Everything else — residuals, normalization, feed-forward, cross-attention — is the scaffolding that lets you stack it deep.

    Build · Python · ~75 min

  6. BERT — Masked Language Modeling

    GPT predicts the next word. BERT predicts a missing word. One sentence of difference — and half a decade of everything embedding-shaped.

    Build · Python · ~45 min

  7. GPT — Causal Language Modeling

    BERT sees both sides. GPT sees only the past. The triangle mask is the most consequential single line of code in modern AI.

    Build · Python · ~75 min

  8. T5, BART — Encoder-Decoder Models

    Encoders understand. Decoders generate. Put them back together and you get a model built for input → output tasks: translate, summarize, rewrite, transcribe.

    Learn · Python · ~45 min

  9. Vision Transformers (ViT)

    An image is a grid of patches. A sentence is a grid of tokens. The same transformer eats both. Before 2020, computer vision meant convolutions.

    Build · Python · ~45 min

  10. Audio Transformers — Whisper Architecture

    Audio is an image of frequency over time. Whisper is a ViT that eats mel spectrograms and speaks back. Before Whisper (OpenAI, Radford et al.

    Learn · Python · ~45 min

  11. Mixture of Experts (MoE)

    A dense 70B transformer activates every parameter for every token. A 671B MoE activates only 37B per token and beats it on every benchmark. Sparsity is the most important scaling idea of the decade.

    Build · Python · ~45 min

  12. KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Build · Python · ~75 min

  13. Scaling Laws

    The 2020 Kaplan paper said: bigger model, lower loss. The 2022 Hoffmann paper said: you were under-training. Compute goes into two buckets — parameters and tokens — and the split is not obvious.

    Learn · Python · ~45 min

  14. Build a Transformer from Scratch — The Capstone

    Thirteen lessons. One model. No shortcuts. You've read every paper. You've implemented attention, multi-head splits, positional encodings, encoder and decoder blocks, BERT and GPT losses, MoE, KV…

    Build · Python · ~120 min

  15. Attention Variants — Sliding Window, Sparse, Differential

    Full attention is a circle. Every token sees every token, and memory pays the price. Four variants bend the shape of the circle and recover half the cost.

    Build · Python · ~60 min

  16. Speculative Decoding — Draft, Verify, Repeat

    Autoregressive decoding is serial. Each token waits for the previous one. Speculative decoding breaks the chain: a cheap model drafts N tokens, the expensive model verifies all N in one forward pass.

    Build · Python · ~60 min

Glossary terms in this phase

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
  • FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
  • GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
  • WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…

Frequently asked questions

How many lessons are in Phase 07: Transformers Deep Dive?

Phase 07 has 16 lessons: 12 Build lessons and 4 Learn lessons. The lesson code uses Python.

What should I know before I start Phase 07?

The phase guide gives these prerequisites: Phase 3 Deep Learning Core, Phase 5 Lesson 09 on sequence-to-sequence models, and Phase 5 Lesson 10 on attention. In the course roadmap, this phase builds on Phase 05: NLP: Foundations to Advanced.

Is Phase 07 free?

Yes. All 16 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 07 take?

The time estimates of all 16 lessons add up to about 17 hours.

What comes after Phase 07?

Phase 08: Generative AI and Phase 10: LLMs from Scratch build on this phase.