Phase 05 · NLP: Foundations to Advanced

Learn NLP from Scratch: 29 Free Lessons

Language is the interface to intelligence. Master every layer.

  • 29 lessons
  • 24 build
  • 5 learn
  • ~31 hours
  • Python

Start Phase 05

First lesson Text Processing — Tokenization, Stemming, Lemmatization

Run this command from the repository root:

python3 phases/05-nlp-foundations-to-advanced/01-text-processing/code/main.py

Keep the command, exit code, token list, stems, lemmas, and one example where stemming loses meaning but lemmatization preserves it.

All 29 lessons in Phase 05

  1. Text Processing — Tokenization, Stemming, Lemmatization

    Language is continuous. Models are discrete. Preprocessing is the bridge. A model cannot read "The cats were running." It reads integers. Every NLP system opens with the same three questions.

    Build · Python · ~45 min

  2. Bag of Words, TF-IDF, and Text Representation

    Count first, think later. TF-IDF still beats embeddings on well-defined tasks in 2026. The model needs numbers. You have strings. Every NLP pipeline has to answer the same question.

    Build · Python · ~75 min

  3. Word Embeddings — Word2Vec from Scratch

    A word is the company it keeps. Train a shallow net on that idea and geometry falls out. TF-IDF knows dog and puppy are different words. It does not know they mean nearly the same thing.

    Build · Python · ~75 min

  4. GloVe, FastText, and Subword Embeddings

    Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.

    Build · Python · ~45 min

  5. Sentiment Analysis

    The canonical NLP task. Most of what you need to know about classical text classification shows up here. "The food was not great." Positive or negative? Sentiment sounds simple.

    Build · Python · ~75 min

  6. Named Entity Recognition

    Pull the names out. Sounds easy until you deal with ambiguous boundaries, nested entities, and domain jargon. "Apple sued Google over its iPhone search deal in the US." Five entities: Apple (ORG),…

    Build · Python · ~75 min

  7. POS Tagging and Syntactic Parsing

    Grammar was unfashionable for a while. Then every LLM pipeline needed to validate structured extraction, and it came back. Lesson 01 promised that lemmatization needs a part-of-speech tag.

    Build · Python · ~45 min

  8. CNNs and RNNs for Text

    Convolutions learn n-grams. Recurrences remember. Both are superseded by attention. Both still matter on constrained hardware. TF-IDF and Word2Vec produced flat vectors that ignored word order.

    Build · Python · ~75 min

  9. Sequence-to-Sequence Models

    Two RNNs pretending to be a translator. The bottleneck they hit is the reason attention exists. Classification maps a variable-length sequence to a single label.

    Build · Python · ~75 min

  10. Attention Mechanism — The Breakthrough

    The decoder stops squinting at a compressed summary and starts looking at the whole source. Everything after this is attention plus engineering. Lesson 09 ended on a measured failure.

    Build · Python · ~45 min

  11. Machine Translation

    Translation is the task that paid for NLP research for thirty years and keeps paying now. A model reads a sentence in one language and produces a sentence in another. Length varies. Word order varies.

    Build · Python · ~75 min

  12. Text Summarization

    Extractive systems tell you what the document said. Abstractive systems tell you what the author meant. Different tasks, different pitfalls. A 2,000-word news article lands in your feed.

    Build · Python · ~75 min

  13. Question Answering Systems

    Three systems shaped modern QA. Extractive found spans. Retrieval-augmented grounded them in documents. Generative produced answers. Every modern AI assistant is a mix of the three.

    Build · Python · ~75 min

  14. Information Retrieval and Search

    BM25 is precise but brittle. Dense casts a wide net but misses keywords. Hybrid is the 2026 default. Everything else is tuning.

    Build · Python · ~75 min

  15. Topic Modeling — LDA and BERTopic

    LDA: documents are mixtures of topics, topics are distributions over words. BERTopic: documents cluster in embedding space, clusters are topics. Same goal, different decompositions.

    Build · Python · ~45 min

  16. Text Generation Before Transformers — N-gram Language Models

    If a word is surprising, the model is bad. Perplexity makes surprise a number. Smoothing keeps it finite. Before transformers, before RNNs, before word embeddings, a language model predicted the…

    Build · Python · ~45 min

  17. Chatbots — Rule-Based to Neural to LLM Agents

    ELIZA replied with pattern matches. DialogFlow mapped intents. GPT answered from weights. Claude runs tools and verifies. Each era solved the previous one's worst failure.

    Build · Python · ~75 min

  18. Multilingual NLP

    One model, 100+ languages, zero training data for most of them. Cross-lingual transfer is the practical miracle of the 2020s. English has billions of labeled examples. Urdu has thousands.

    Build · Python · ~45 min

  19. Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece

    Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.

    Learn · Python · ~60 min

  20. Structured Outputs & Constrained Decoding

    Ask an LLM for JSON. Get JSON most of the time. In production, "most" is the problem. Constrained decoding turns "most" into "always" by editing the logits before sampling.

    Build · Python · ~60 min

  21. Natural Language Inference — Textual Entailment

    "t entails h" means a human reading t would conclude h is true. NLI is the task of predicting entailment / contradiction / neutral. Boring on the surface, load-bearing in production.

    Learn · Python · ~60 min

  22. Embedding Models — The 2026 Deep Dive

    Word2Vec gave you a vector per word. Modern embedding models give you a vector per passage, cross-lingual, with sparse, dense, and multi-vector views, sized to fit your index.

    Learn · Python · ~60 min

  23. Chunking Strategies for RAG

    Chunking configuration influences retrieval quality as much as the choice of embedding model (Vectara NAACL 2025). Get chunking wrong and no amount of reranking saves you.

    Build · Python · ~60 min

  24. Coreference Resolution

    "She called him. He did not answer. The doctor was at lunch." Three references to two people and nobody is named. Coreference resolution figures out who is who. Extract every mention of Apple Inc.

    Learn · Python · ~60 min

  25. Entity Linking & Disambiguation

    NER found "Paris." Entity linking decides: Paris, France? Paris Hilton? Paris, Texas? Paris (the Trojan prince)? Without linking, your knowledge graph stays ambiguous.

    Build · Python · ~60 min

  26. Relation Extraction & Knowledge Graph Construction

    NER found the entities. Entity linking anchored them. Relation extraction finds the edges between them. A knowledge graph is the sum of nodes, edges, and their provenance.

    Build · Python · ~60 min

  27. LLM Evaluation — RAGAS, DeepEval, G-Eval

    Exact-match and F1 miss semantic equivalence. Human review does not scale. LLM-as-judge is the production answer — with enough calibration to trust the number.

    Build · Python · ~75 min

  28. Long-Context Evaluation — NIAH, RULER, LongBench, MRCR

    Gemini 3 Pro advertises 10M tokens of context. At 1M tokens, 8-needle MRCR drops to 26.3%. Advertised ≠ usable. Long-context evaluation tells you the actual capacity of the model you are shipping on.

    Learn · Python · ~60 min

  29. Dialogue State Tracking

    "I want a cheap restaurant in the north... actually make it moderate... and add Italian." Three turns, three state updates. DST keeps the slot-value dict in sync so the booking works.

    Build · Python · ~75 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
  • Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
  • ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
  • CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • Dense RetrievalFirst-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
  • RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
  • Reciprocal Rank Fusion (RRF)A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
  • ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
  • Structured OutputModel output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form…
  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
  • Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.

Frequently asked questions

How many lessons are in Phase 05: NLP: Foundations to Advanced?

Phase 05 has 29 lessons: 24 Build lessons and 5 Learn lessons. The lesson code uses Python.

What should I know before I start Phase 05?

The phase guide gives these prerequisites: Phase 2 Lesson 14, Naive Bayes. Python 3.11+ is enough for the first lesson. In the course roadmap, this phase builds on Phase 03: Deep Learning Core.

Is Phase 05 free?

Yes. All 29 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 05 take?

The time estimates of all 29 lessons add up to about 31 hours.

What comes after Phase 05?

Phase 07: Transformers Deep Dive builds on this phase.