Phase 19 · Capstone Projects
AI Engineering Capstone Projects: 85 Free Lessons
Prove everything you learned. Build portfolio-grade systems.
- 85 lessons
- ~627 hours
- Python, YAML
Start Phase 19
Starter project Capstone 01 — Terminal-Native Coding Agent
Run this command from the repository root:
python3 phases/19-capstone-projects/01-terminal-native-coding-agent/code/main.pyKeep the command, exit code, final plan state, budget totals, trace events, and the architecture decision you would change before connecting a real model.
All 85 lessons in Phase 19
- Capstone 01 — Terminal-Native Coding Agent
By 2026 the shape of a coding agent is settled. A TUI harness, a stateful plan, a sandboxed tool surface, a loop that plans, acts, observes, recovers.
- Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search)
Every serious engineering org in 2026 runs an internal code search that understands meaning, not just strings. Sourcegraph Amp, Cursor's codebase answers, Augment's enterprise graph, Aider's…
- Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS)
A voice agent that feels right has end-to-end latency under 800ms, knows when you have stopped talking, handles barge-in, and can call a tool without stalling.
- Capstone 04 — Multimodal Document QA (Vision-First PDF, Tables, Charts)
The 2026 document-QA frontier moved away from OCR-then-text and toward vision-first late interaction. ColPali, ColQwen2.5, and ColQwen3-omni treat each PDF page as an image, embed it with…
- Capstone 05 — Autonomous Research Agent (AI-Scientist Class)
Sakana's AI-Scientist-v2 published full papers. Agent Laboratory ran the experiments. Allen AI shared traces. The 2026 shape is plan-execute-verify tree search over experiments, budgeted cost,…
- Capstone 06 — DevOps Troubleshooting Agent for Kubernetes
AWS's DevOps Agent went GA, Resolve AI published its K8s playbooks, NeuBird demoed semantic monitoring, and Metoro tied AI SRE to per-service SLOs.
- Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve)
An 8B model trained on your own data, DPO-aligned on your own preferences, quantized, speculative-decoded, and served at measurable $/1M tokens.
- Capstone 08 — Production RAG Chatbot for a Regulated Vertical
Harvey, Glean, Mendable, and LlamaCloud all run the same production shape in 2026. Ingest with docling or Unstructured and ColPali for visuals. Hybrid search. Re-rank with bge-reranker-v2-gemma.
- Capstone 09 — Code Migration Agent (Repo-Level Language / Runtime Upgrade)
Amazon's MigrationBench (Java 8 to 17) and Google's App Engine Py2-to-Py3 migrator set the 2026 bar. Moderne's OpenRewrite does deterministic AST rewrites at scale.
- Capstone 10 — Multi-Agent Software Engineering Team
The 2026 shape of a multi-agent engineering team has converged: an architect plans, N coders work in parallel worktrees, a reviewer gates, a tester verifies.
- Capstone 11 — LLM Observability & Eval Dashboard
Langfuse went open-core. Arize Phoenix published the 2026 GenAI semconv mappings. Helicone and Braintrust both doubled down on per-user cost attribution.
- Capstone 12 — Video Understanding Pipeline (Scene, QA, Search)
Twelve Labs productized Marengo + Pegasus. VideoDB shipped the CRUD-for-video API. AI2's Molmo 2 published open VLM checkpoints. Gemini long-context handles hours of video natively.
- Capstone 13: Stateless MCP Server with Registry and Governance
Production MCP is not one server process. It is a chain of contracts: publishable metadata, live discovery, a stateless request envelope, authorization, policy, audit, and deployment evidence.
- Capstone 14 — Speculative-Decoding Inference Server
Speculative decoding — a cheap draft proposes tokens, the target model verifies them in one pass — is now a production-ready optimization, not a research trick.
- Capstone 15 — Constitutional Safety Harness + Red-Team Range
Anthropic's Constitutional Classifiers, Meta's Llama Guard 4, Google's ShieldGemma-2, NVIDIA's Nemotron 3 Content Safety, and X-Guard for multilingual coverage defined the 2026 safety-classifier…
- Capstone 16 — GitHub Issue-to-PR Autonomous Agent
Label an issue, get a PR — the 2026 product shape for autonomous coding agents: run an agent in a cloud sandbox, verify tests pass, and post a review-ready PR with rationale.
- Capstone 17 — Personal AI Tutor (Adaptive, Multimodal, with Memory)
Khanmigo (Khan Academy), Duolingo Max, Google LearnLM / Gemini for Education, Quizlet Q-Chat, and Synthesis Tutor all shipped adaptive multimodal tutoring at scale in 2026.
- Agent Harness Loop Contract
The harness is the agent. The model is a coprocessor. This lesson freezes the loop contract you can wire any model into.
- Tool Registry with Schema Validation
A tool the agent cannot validate is a tool the agent cannot call. Build the registry and the schema checker before you build the tools.
- JSON-RPC 2.0 Over Newline-Delimited Stdio
The transport between a model client and a tool server is JSON-RPC over stdio. Hand-rolling it once teaches you what every framing layer is paying for.
- Function Call Dispatcher
The dispatcher is where the harness pays for every promise the schema made. Timeouts, retries, dedupe, error mapping. All on one seam.
- Plan-Execute Control Flow
A plan that cannot survive a failure is a script. A script that can replan is an agent. Build the replanner first. Represent a plan as an ordered list of typed steps so the executor can reason about…
- Capstone Lesson 25: Verification Gates and the Observation Budget
An agent harness without a verification layer is a wish in a trenchcoat. This lesson builds the deterministic gate chain that decides whether a tool call is allowed to fire, how much of its output…
- Capstone Lesson 26: Sandbox Runner with Denylist and Path Jail
The verification gate decides whether a tool call should run. The sandbox decides what happens when it does. This lesson ships a subprocess runner that refuses dangerous executables, refuses…
- Capstone Lesson 27: Eval Harness with Fixture Tasks
A coding agent is only as good as the suite of tasks you measure it against. This lesson builds an evaluation harness that takes a folder of fixture tasks, runs each through a candidate agent,…
- Capstone Lesson 28: Observability with OTel GenAI Spans and Prometheus Metrics
An agent harness without observability is a black box that costs money. This lesson hand-rolls a span builder that emits records compliant with the OpenTelemetry GenAI semantic conventions, writes…
- Capstone Lesson 29: End-to-End Coding Agent on the Harness
Track A's payoff. This lesson stitches the gate chain, the sandbox, the eval harness, and the OTel spans into one working coding agent that fixes a real (small, fixture-scale) bug in a multi-file…
- BPE Tokenizer From Scratch
Bytes in, ids out, ids back to the same bytes. Build the tokenizer that every modern text model still starts from. Train a Byte-Pair Encoding vocabulary from a raw text corpus by repeatedly merging…
- Tokenized Dataset with Sliding Window
A pretraining run is a function from token ids to gradients. This lesson builds the conveyor that feeds the ids in. Convert a raw corpus into a stream of token ids by calling the tokenizer once.
- Token and Positional Embeddings
Ids are integers. The model wants vectors. Two lookup tables sit between them, and the choice of the positional one shapes what the model can learn.
- Multi-Head Self-Attention
One linear projection, three views, H parallel heads, one mask. The attention block as the model actually uses it. Implement a batched Query/Key/Value projection as a single linear layer split into…
- Transformer Block from Scratch
One block is the unit of every modern decoder LLM. Layer norm, multi head attention, residual, MLP, residual. The pre-LN variant trains stably without warmup.
- GPT Model Assembly
Twelve blocks stacked, a token embedding, a learned position embedding, a final LayerNorm, and a tied language model head. That is the entire 124 million parameter GPT model.
- Training Loop and Evaluation
A loop that does not measure is a loop that lies. This lesson builds the training loop that drives the GPT model: AdamW with weight decay split, a warmup plus cosine learning rate schedule, a…
- Loading Pretrained Weights
Training a 124 million parameter model from scratch is a budget decision; loading a published checkpoint is a Tuesday. This lesson loads pretrained GPT-2 style weights from a safetensors file into…
- Capstone Lesson 38: Classifier Fine-Tuning by Head Swap
Track B's first capstone. A pretrained language model is a stack of self-attention blocks ending in a token-prediction head. When you want spam vs ham, the head is wrong but the body is mostly right.
- Capstone Lesson 39: Instruction Tuning by Supervised Fine-Tuning
A pretrained base model can extend a sequence but cannot follow an instruction. Supervised fine-tuning is the smallest change that fixes this: feed the model paired examples of an instruction and a…
- Capstone Lesson 40: Direct Preference Optimization from Scratch
Reward models and PPO are the classical RLHF stack. DPO collapses that stack into a single supervised loss that fits a policy directly against preference pairs.
- Capstone Lesson 41: Full Evaluation Pipeline
Training is the part you can monitor with loss curves. Evaluation is the part you have to design. This lesson builds a unified eval pipeline that takes any trained language model, runs four…
- Large Corpus Downloader
Training a language model begins long before the first forward pass. The corpus has to land on disk, decompressed, deduplicated, and addressable, with the resume story already worked out before the…
- HDF5 Tokenized Corpus
The downloaded corpus has to land in a layout the trainer can stream from at line speed. JSONL on disk does not survive 16 dataloader workers. HDF5 with a resizable, chunked integer dataset does.
- Cosine LR with Linear Warmup
The learning-rate schedule is the second most important decision after the loss function. AdamW with a cosine decay and a linear warmup is the modern default for language-model training because it…
- Gradient Clipping and Mixed Precision
The optimizer and schedule from the previous lesson assume gradients are sane. They usually are not. A single bad batch can spike the gradient norm by three orders of magnitude.
- Gradient Accumulation
Train at an effective batch you cannot afford, one micro-batch at a time. Scale the loss, hold the optimizer step, and let the gradients pile up.
- Checkpoint Save and Resume
Train interrupts kill runs; checkpoints let them continue. Save model, optimizer, scheduler, loss history, step counter, and RNG state, atomically, so a kill at any moment leaves a valid file on disk.
- Distributed Data Parallel and FSDP from Scratch
Multi-rank training is two collectives and one rule. Broadcast the parameters at startup, average the gradients after backward, never let the ranks disagree about what step they are on.
- Language Model Evaluation Harness
A model that does well on a task you cannot define is a model that does well by accident. The harness is the task definition, the metric, the runner, and the leaderboard, in one short, swappable…
- Hypothesis Generator
A research agent that asks the same question twice is wasting tokens. The trick is forcing each draft to land somewhere new.
- Literature Retrieval
A hypothesis is cheap. Knowing whether someone already proved it is the expensive part. Build the retrieval layer that answers that question before the runner spins up a sandbox.
- Experiment Runner
The loop is only as honest as its measurements. Build the runner that takes a spec, executes it in a sandboxed subprocess, and emits a json metrics blob the evaluator can trust.
- Result Evaluator
The runner produced numbers. The evaluator decides whether those numbers are an improvement, a regression, or noise. Build the verdict path that turns metrics into a one line conclusion.
- Paper Writer
A LaTeX skeleton is a contract between the researcher and the typesetter. If the contract is broken the document does not compile, and the failure is loud. Build the skeleton first, then fill it.
- Critic Loop
A critic that returns "looks good" the first time is broken. A critic that always returns "needs work" is broken. The interesting critic is the one that converges, and you have to engineer…
- Iteration Scheduler
A research loop without a scheduler is a queue with delusions. The scheduler is where the loop decides what to stop exploring, and that decision is the whole game.
- End-to-End Research Demo
A demo is the place where every contract you wrote earlier has to compose. If any one of them leaks, the demo is the lesson that catches it.
- Vision Encoder Patches
A vision model that reads pixels needs a tokenizer for pixels. Patch embedding is that tokenizer. Cut the image into a grid of squares, flatten each square, project it through one linear layer, then…
- Vision Transformer Encoder
Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…
- Projection Layer for Modality Alignment
A vision encoder produces image tokens. A text decoder consumes text tokens. The two live in different vector spaces. A small two-layer MLP projects image tokens into the text embedding space, and a…
- Cross-Attention Fusion
The projection layer aligns one image vector with one caption vector. A real vision-language decoder needs every text token to attend to every patch token, so the model can ground each word in a…
- Vision-Language Pretraining
The encoder, projection, and decoder are wired. Now train them together. Two objectives drive learning: a contrastive image-text loss (InfoNCE) that pulls matching pairs together in the joint…
- Multimodal Evaluation
Training is half the loop. The other half is measurement. This lesson builds three evaluation surfaces from primitives: image-caption retrieval reported as R@1, R@5, R@10; visual question answering…
- Chunking Strategies, Compared
Chunking decides what your retriever can ever surface. Get the boundaries wrong and no embedding model, no reranker, no LLM can repair the damage downstream.
- Hybrid Retrieval with BM25 and Dense Embeddings
Lexical and semantic retrieval fail on opposite query distributions. Hybrid retrieval with reciprocal rank fusion does not interpolate, it votes - and the vote wins on every query class.
- Cross-Encoder Reranker
A bi-encoder embeds query and document independently. A cross-encoder concatenates them and reads both at once. The cross-encoder is the smartest reader and the slowest.
- Query Rewriting: HyDE, Multi-Query, and Decomposition
The query the user types is not the query your retriever wants. Rewriting bridges the gap before retrieval, so the index sees something closer to what the answer looks like.
- RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance
If you cannot grade your retrieval and your answer at the same time, you cannot ship the system. The two are not the same metric and the same prompt fails on different axes.
- End-to-End RAG System
Six lessons of components. One pipeline. One eval loop. One self-terminating demo. This is the system you ship. Compose the chunker, hybrid retriever, query rewriter, cross-encoder reranker, and…
- Task Spec Format
An eval harness is only as good as the contract its tasks honour. Freeze the JSONL shape and the metric vocabulary before you write a single scoring function.
- Classical Metrics
BLEU, ROUGE-L, F1, exact-match, accuracy. Five metrics that still account for most published LLM eval numbers. Implement each from first principles so you know what the number means.
- Code Exec Metric
Generated code is right when it passes the tests. The eval harness has to extract code, run it without crashing the host, and tally pass-rates honestly. This lesson builds that surface.
- Perplexity and Calibration
If your model says 90 percent confident on a thousand answers and gets six hundred right, it is not well calibrated. Calibration is half of trustworthy eval.
- Leaderboard Aggregation
Per-task scores are easy. Per-model rankings across heterogeneous tasks are harder. Statistical significance on a thousand-prediction leaderboard is the part everyone skips.
- End-to-End Eval Runner
Five lessons of plumbing, one lesson to glue them. The runner reads the task spec from lesson 70, calls a model through an adapter, scores with lessons 71 and 72, attaches the calibration report…
- Collective Ops From Scratch
The four collective operations that hold distributed training together are allreduce, broadcast, allgather, and reducescatter.
- Data Parallel DDP From Scratch
DistributedDataParallel is a hook on top of allreduce. Wrap a model, broadcast the initial parameters from rank 0 so every rank starts identical, install a backward hook on every parameter that…
- ZeRO Optimizer State Sharding
Adam stores two moment estimates per parameter, both in float32. A 7B-parameter model carries 56 GB of optimiser state. ZeRO stage 1 shards that across N ranks; each rank owns 1/N of the optimiser.
- Pipeline Parallel and Bubble Analysis
Tensor parallelism splits the matrix multiply across ranks. Pipeline parallelism splits the model across ranks, one stage per rank. Microbatches flow through the pipeline.
- Sharded Checkpoint and Atomic Resume
A 70B-parameter training job is paused by a node failure every few hours. The checkpoint format decides whether you lose 30 minutes or 30 hours.
- End-to-End Distributed Training
Lessons 76 through 80 each built one piece. This is the assembly: a tiny GPT trained across 4 simulated ranks with DDP for gradient sync, ZeRO-1 for optimiser-state sharding, and a sharded…
- Capstone 82 — Jailbreak Taxonomy
A safety harness without a taxonomy is a coin flip. Name the attack before you defend it. A model deployed without an attack model is a model defended against nothing in particular.
- Capstone 83 — Prompt Injection Detector
A detector is a function from prompt to confidence and category. Anything else is a vibe. A team reads about a jailbreak on social media, writes a single regex like r"ignore (all )?previous", ships…
- Capstone 84 — Refusal Evaluation
Helpfulness on benign prompts and refusal on harmful prompts are two metrics, not one. Measure both. A safety pass on an assistant goes wrong in two opposite ways.
- Capstone 85 — Content Classifier Integration
Classifiers on the output side answer a different question than rules on the input side. Both need a policy router. Inputs are not the only attack surface.
- Capstone 86 — Constitutional Rules Engine
A rule is a name, a predicate, and an explanation. Anything missing one of those three is a vibe, not a rule. Classifiers cover the recognizable failures. Rules engines cover the contractual ones.
- Capstone 87 — End-to-End Safety Gate
Pre-gen, during-gen, post-gen. Three checkpoints, one verdict, an audit trail per request. Lessons 82-86 in this track each shipped a single piece: a taxonomy, an input detector, an evaluation…
Glossary terms in this phase
- AdamWAn Adam variant that decouples weight decay from the gradient-based parameter update.
- AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
- Agent HarnessThe runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the…
- AlignmentThe effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected…
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
- BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
- Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
- ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
- Coding AgentAn agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance…
- Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
- Dense RetrievalFirst-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
- DPO (Direct Preference Optimization)A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
- EpochOne traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- Exact Match (EM)A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
- GradientA vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent)…
- Gradient AccumulationSumming or averaging gradients from several microbatches before performing one optimizer update.
- Gradient ClippingLimiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- Hybrid RetrievalRetrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- JailbreakAn adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are…
- Late FusionProcessing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- MCP (Model Context Protocol)An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request,…
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
- Modality AlignmentLearning or establishing correspondences between representations from different modalities so semantically or temporally related items can…
- Multimodal FusionCombining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated…
- NaN (Not a Number)A floating-point value representing an undefined or unrepresentable numerical result.
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- OptimizerAn algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and…
- PatchA reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
- PerplexityThe exponentiated average negative log-likelihood under a stated tokenization and logarithm convention.
- Prompt InjectionAn attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or…
- RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
- Recall@KFor one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`.
- Reciprocal Rank Fusion (RRF)A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
- RerankerA second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
- SandboxAn isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host…
- Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
- Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
- SFT (Supervised Fine-Tuning)Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training…
- Stateless MCPThe MCP 2026-07-28 request model in which every request carries the protocol version and client capabilities in `params._meta`, while…
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Trust BoundaryAn interface where data, instructions, identity, or authority crosses between components or principals that operate under different trust…
- Verification GateA control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
- Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
- Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
- WarmupAn initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
- WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
Frequently asked questions
How many lessons are in Phase 19: Capstone Projects?
Phase 19 has 85 lessons, all Capstone lessons. The lesson code uses Python and YAML.
What should I know before I start Phase 19?
The phase guide gives these prerequisites: Choose a project whose listed phases you have completed. The Terminal-Native Coding Agent expects Phases 11, 13, 14, 15, and 17. In the course roadmap, this phase builds on Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production and Phase 18: Ethics, Safety & Alignment.
Is Phase 19 free?
Yes. All 85 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 19 take?
The time estimates of all 85 lessons add up to about 627 hours.