Phase 05 · NLP: Foundations to Advanced
Learn NLP from Scratch: 29 Free Lessons
Language is the interface to intelligence. Master every layer.
- 29 lessons
- 24 build
- 5 learn
- ~31 hours
- Python
Start Phase 05
First lesson Text Processing — Tokenization, Stemming, Lemmatization
Run this command from the repository root:
python3 phases/05-nlp-foundations-to-advanced/01-text-processing/code/main.pyKeep the command, exit code, token list, stems, lemmas, and one example where stemming loses meaning but lemmatization preserves it.
All 29 lessons in Phase 05
- Text Processing — Tokenization, Stemming, Lemmatization
Language is continuous. Models are discrete. Preprocessing is the bridge. A model cannot read "The cats were running." It reads integers. Every NLP system opens with the same three questions.
- Bag of Words, TF-IDF, and Text Representation
Count first, think later. TF-IDF still beats embeddings on well-defined tasks in 2026. The model needs numbers. You have strings. Every NLP pipeline has to answer the same question.
- Word Embeddings — Word2Vec from Scratch
A word is the company it keeps. Train a shallow net on that idea and geometry falls out. TF-IDF knows dog and puppy are different words. It does not know they mean nearly the same thing.
- GloVe, FastText, and Subword Embeddings
Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.
- Sentiment Analysis
The canonical NLP task. Most of what you need to know about classical text classification shows up here. "The food was not great." Positive or negative? Sentiment sounds simple.
- Named Entity Recognition
Pull the names out. Sounds easy until you deal with ambiguous boundaries, nested entities, and domain jargon. "Apple sued Google over its iPhone search deal in the US." Five entities: Apple (ORG),…
- POS Tagging and Syntactic Parsing
Grammar was unfashionable for a while. Then every LLM pipeline needed to validate structured extraction, and it came back. Lesson 01 promised that lemmatization needs a part-of-speech tag.
- CNNs and RNNs for Text
Convolutions learn n-grams. Recurrences remember. Both are superseded by attention. Both still matter on constrained hardware. TF-IDF and Word2Vec produced flat vectors that ignored word order.
- Sequence-to-Sequence Models
Two RNNs pretending to be a translator. The bottleneck they hit is the reason attention exists. Classification maps a variable-length sequence to a single label.
- Attention Mechanism — The Breakthrough
The decoder stops squinting at a compressed summary and starts looking at the whole source. Everything after this is attention plus engineering. Lesson 09 ended on a measured failure.
- Machine Translation
Translation is the task that paid for NLP research for thirty years and keeps paying now. A model reads a sentence in one language and produces a sentence in another. Length varies. Word order varies.
- Text Summarization
Extractive systems tell you what the document said. Abstractive systems tell you what the author meant. Different tasks, different pitfalls. A 2,000-word news article lands in your feed.
- Question Answering Systems
Three systems shaped modern QA. Extractive found spans. Retrieval-augmented grounded them in documents. Generative produced answers. Every modern AI assistant is a mix of the three.
- Information Retrieval and Search
BM25 is precise but brittle. Dense casts a wide net but misses keywords. Hybrid is the 2026 default. Everything else is tuning.
- Topic Modeling — LDA and BERTopic
LDA: documents are mixtures of topics, topics are distributions over words. BERTopic: documents cluster in embedding space, clusters are topics. Same goal, different decompositions.
- Text Generation Before Transformers — N-gram Language Models
If a word is surprising, the model is bad. Perplexity makes surprise a number. Smoothing keeps it finite. Before transformers, before RNNs, before word embeddings, a language model predicted the…
- Chatbots — Rule-Based to Neural to LLM Agents
ELIZA replied with pattern matches. DialogFlow mapped intents. GPT answered from weights. Claude runs tools and verifies. Each era solved the previous one's worst failure.
- Multilingual NLP
One model, 100+ languages, zero training data for most of them. Cross-lingual transfer is the practical miracle of the 2020s. English has billions of labeled examples. Urdu has thousands.
- Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece
Word tokenizers choke on unseen words. Character tokenizers blow up sequence length. Subword tokenizers split the difference. Every modern LLM ships on one. Your vocabulary has 50,000 words.
- Structured Outputs & Constrained Decoding
Ask an LLM for JSON. Get JSON most of the time. In production, "most" is the problem. Constrained decoding turns "most" into "always" by editing the logits before sampling.
- Natural Language Inference — Textual Entailment
"t entails h" means a human reading t would conclude h is true. NLI is the task of predicting entailment / contradiction / neutral. Boring on the surface, load-bearing in production.
- Embedding Models — The 2026 Deep Dive
Word2Vec gave you a vector per word. Modern embedding models give you a vector per passage, cross-lingual, with sparse, dense, and multi-vector views, sized to fit your index.
- Chunking Strategies for RAG
Chunking configuration influences retrieval quality as much as the choice of embedding model (Vectara NAACL 2025). Get chunking wrong and no amount of reranking saves you.
- Coreference Resolution
"She called him. He did not answer. The doctor was at lunch." Three references to two people and nobody is named. Coreference resolution figures out who is who. Extract every mention of Apple Inc.
- Entity Linking & Disambiguation
NER found "Paris." Entity linking decides: Paris, France? Paris Hilton? Paris, Texas? Paris (the Trojan prince)? Without linking, your knowledge graph stays ambiguous.
- Relation Extraction & Knowledge Graph Construction
NER found the entities. Entity linking anchored them. Relation extraction finds the edges between them. A knowledge graph is the sum of nodes, edges, and their provenance.
- LLM Evaluation — RAGAS, DeepEval, G-Eval
Exact-match and F1 miss semantic equivalence. Human review does not scale. LLM-as-judge is the production answer — with enough calibration to trust the number.
- Long-Context Evaluation — NIAH, RULER, LongBench, MRCR
Gemini 3 Pro advertises 10M tokens of context. At 1M tokens, 8-needle MRCR drops to 26.3%. Advertised ≠ usable. Long-context evaluation tells you the actual capacity of the model you are shipping on.
- Dialogue State Tracking
"I want a cheap restaurant in the north... actually make it moderate... and add Italian." Three turns, three state updates. DST keeps the slot-value dict in sync so the booking works.
Glossary terms in this phase
- AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
- CNN (Convolutional Neural Network)A neural network that uses convolution operations (sliding filters over the input) to detect local patterns.
- DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
- Dense RetrievalFirst-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
- Reciprocal Rank Fusion (RRF)A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
- ROUGEA family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
- Structured OutputModel output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form…
- TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
- Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.
Frequently asked questions
How many lessons are in Phase 05: NLP: Foundations to Advanced?
Phase 05 has 29 lessons: 24 Build lessons and 5 Learn lessons. The lesson code uses Python.
What should I know before I start Phase 05?
The phase guide gives these prerequisites: Phase 2 Lesson 14, Naive Bayes. Python 3.11+ is enough for the first lesson. In the course roadmap, this phase builds on Phase 03: Deep Learning Core.
Is Phase 05 free?
Yes. All 29 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 05 take?
The time estimates of all 29 lessons add up to about 31 hours.
What comes after Phase 05?
Phase 07: Transformers Deep Dive builds on this phase.