Phase 11 · LLM Engineering

Learn LLM Engineering: 17 Free Lessons

Put LLMs to work in production applications.

  • 17 lessons
  • 16 build
  • 1 learn
  • ~21 hours
  • Python

Start Phase 11

First lesson Prompt Engineering: Techniques & Patterns

Run this command from the repository root:

python3 phases/11-llm-engineering/01-prompt-engineering/code/prompt_engineering.py

Keep the command, exit code, generated prompt metadata, test results, and one prompt change with the output difference it caused. The demo uses simulated model responses and needs no API key.

All 17 lessons in Phase 11

  1. Prompt Engineering: Techniques & Patterns

    Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.

    Build · Python · ~90 min

  2. Few-Shot, Chain-of-Thought, Tree-of-Thought

    Telling a model what to do is prompting. Showing it how to think is engineering. The gap between 78% and 91% accuracy on the same model, same task, same data is not a better model.

    Build · Python · ~45 min

  3. Structured Outputs: JSON, Schema Validation, Constrained Decoding

    Your LLM returns a string. Your application needs JSON. That gap has crashed more production systems than any model hallucination.

    Build · Python · ~90 min

  4. Embeddings & Vector Representations

    Text is discrete. Math is continuous. Every time you ask an LLM to find "similar" documents, compare meanings, or search beyond keywords, you're relying on a bridge between these two worlds.

    Build · Python · ~75 min

  5. Context Engineering: Windows, Budgets, Memory, and Retrieval

    Prompt engineering is a subset. Context engineering is the whole game. A prompt is a string you type. Context is everything that goes into the model's window: system instructions, retrieved…

    Build · Python · ~90 min

  6. RAG (Retrieval-Augmented Generation)

    Your LLM knows everything up to its training cutoff. It knows nothing about your company's docs, your codebase, or last week's meeting notes.

    Build · Python · ~90 min

  7. Advanced RAG (Chunking, Reranking, Hybrid Search)

    Basic RAG retrieves the top-k most similar chunks. That works for simple questions. It falls apart for multi-hop reasoning, ambiguous queries, and large corpora.

    Build · Python · ~90 min

  8. Fine-Tuning with LoRA & QLoRA

    Full fine-tuning a 7B model requires 56GB of VRAM. You don't have that. Neither do most companies. LoRA lets you fine-tune the same model in 6GB by training less than 1% of the parameters.

    Build · Python · ~75 min

  9. Function Calling & Tool Use

    LLMs cannot do anything. They generate text. That is the entire capability. They cannot check the weather, query a database, send an email, run code, or read a file.

    Build · Python · ~75 min

  10. Evaluation & Testing LLM Applications

    You would never deploy a web app without tests. You would never ship a database migration without a rollback plan. But right now, most teams ship LLM applications by reading 10 outputs and saying…

    Build · Python · ~45 min

  11. Caching, Rate Limiting & Cost Optimization

    Most AI startups do not die from bad models. They die from bad unit economics. A single GPT-4o call costs fractions of a cent.

    Build · Python · ~45 min

  12. Guardrails, Safety & Content Filtering

    Your LLM application will be attacked. Not might. Will. The first prompt injection attempt against your production system will come within 48 hours of launch.

    Build · Python · ~45 min

  13. Building a Production LLM Application

    You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.

    Build · Python · ~120 min

  14. Model Context Protocol (MCP)

    MCP gives an AI host one protocol for discovering and invoking tools, resources, and prompts. The 2026-07-28 revision makes that protocol stateless: capability and version context travels with every…

    Build · Python · ~75 min

  15. Prompt Caching and Context Caching

    Your system prompt is 4,000 tokens. Your RAG context is 20,000 tokens. You send both with every request. You also pay for both — every time.

    Build · Python · ~60 min

  16. Agent State Machines — Graphs, Nodes, Checkpoints

    A ReAct loop written by hand is a while True. The same loop written as an explicit graph is something you can checkpoint, interrupt, branch, and time-travel through. The agent hasn't changed.

    Build · Python · ~75 min

  17. Agent Framework Tradeoffs — Graph, Role, and Actor Orchestration

    Every framework sells the same demo (research agent builds a report) and hides the same bug (state schema fights with the orchestration layer).

    Learn · Python · ~45 min

Glossary terms in this phase

  • AgentA software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration…
  • Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
  • BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
  • Chain of Thought (CoT)Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems…
  • CheckpointA durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references.
  • ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
  • Circuit BreakerA reliability control that temporarily stops calls to a dependency after failures cross a threshold, then probes whether the dependency…
  • Context CompressionReducing the token footprint of source material while attempting to preserve the information required for a later model decision.
  • Context EngineeringDesigning the full information environment supplied to a model at each step, including instructions, selected files, retrieved evidence,…
  • Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • Function CallingA provider or application interface through which a model emits a structured request naming a tool and its arguments.
  • Graceful DegradationPreserving a bounded core service when capacity or dependencies are impaired by reducing optional quality, features, freshness, or…
  • GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • HNSWAn approximate-nearest-neighbor index that organizes vectors in layered proximity graphs and searches from coarse upper layers toward…
  • Human-in-the-Loop (HITL)A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
  • Hybrid RetrievalRetrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • LoRA (Low-Rank Adaptation)A method that keeps base weights frozen and learns low-rank update matrices for selected layers.
  • MCP (Model Context Protocol)An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request,…
  • Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • OrchestrationThe control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
  • ParameterA value learned during training, commonly a weight, bias, embedding element, or normalization parameter.
  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
  • Prompt EngineeringDesigning model-facing instructions, examples, constraints, and output requirements to improve behavior on a defined task.
  • QLoRAA parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training…
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
  • ReActAn agent pattern that interleaves task reasoning, a concrete action, and an observation returned by the environment before deciding the…
  • Readiness ProbeA diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
  • Reciprocal Rank Fusion (RRF)A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
  • RerankerA second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
  • Semantic CacheA cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
  • Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • Structured OutputModel output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form…
  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
  • Vector DatabaseA storage and indexing system that supports nearest-neighbor queries over vector representations, often with metadata filtering,…
  • WeightA trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the…
  • Zero-ShotPerforming a task from instructions or task framing without including task-specific demonstrations in the immediate input.

Frequently asked questions

How many lessons are in Phase 11: LLM Engineering?

Phase 11 has 17 lessons: 16 Build lessons and 1 Learn lesson. The lesson code uses Python.

What should I know before I start Phase 11?

The phase guide gives these prerequisites: Phase 10 Lessons 01 through 05, or equivalent knowledge of tokenization, data pipelines, pretraining, and scaling. In the course roadmap, this phase builds on Phase 10: LLMs from Scratch.

Is Phase 11 free?

Yes. All 17 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 11 take?

The time estimates of all 17 lessons add up to about 21 hours.

What comes after Phase 11?

Phase 13: Tools & Protocols builds on this phase.