Phase 17 · Infrastructure & Production

Learn LLM Serving and AI Infrastructure: 28 Free Lessons

Ship AI to the real world. Scale, monitor, optimize.

  • 28 lessons
  • 27 learn
  • 1 build
  • ~29 hours
  • Python

Start Phase 17

First lesson Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI

Run this command from the repository root:

python3 phases/17-infrastructure-and-production/01-managed-llm-platforms/code/main.py

Keep the command, exit code, cost and latency comparison, redundancy uplift, and a short decision record naming the workload assumptions behind your choice.

All 28 lessons in Phase 17

  1. Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI

    Three hyperscalers, three distinct strategies. AWS Bedrock is a model marketplace — Claude, Llama, Titan, Stability, Cohere behind one API.

    Learn · Python · ~60 min

  2. Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale

    The 2026 inference market is no longer GPU time rental. It bifurcates into custom silicon (Groq, Cerebras, SambaNova), GPU platforms (Baseten, Together, Fireworks, Modal), and API-first marketplaces…

    Learn · Python · ~60 min

  3. GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling

    Three layers, not one. Karpenter provisions nodes dynamically (under one minute, 40% faster than Cluster Autoscaler). KAI Scheduler handles gang scheduling, topology awareness, and hierarchical…

    Learn · Python · ~75 min

  4. Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill

    Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…

    Learn · Python · ~75 min

  5. EAGLE-3 Speculative Decoding in Production

    Speculative decoding pairs a fast draft model with the target model. The draft proposes K tokens; the target verifies in a single forward; accepted tokens are free.

    Learn · Python · ~60 min

  6. Prefix-Cache Serving — RadixAttention and KV Reuse

    Treat the KV cache as a first-class, reusable resource stored in a radix tree, and scheduling changes with it: instead of FCFS (first-come, first-served) as vLLM schedules, a cache-aware scheduler…

    Learn · Python · ~75 min

  7. Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell

    Hardware-specialized inference compilation trades portability for throughput, and TensorRT-LLM — NVIDIA-only, tuned for Blackwell — is the clearest example of the trade paying off.

    Learn · Python · ~75 min

  8. Inference Metrics — TTFT, TPOT, ITL, Goodput, P99

    Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.

    Learn · Python · ~60 min

  9. Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4

    Quantization format is not a universal choice — it is a function of hardware, serving engine, and workload. GGUF Q4KM or Q5KM owns CPU and edge, delivered through llama.cpp and Ollama.

    Learn · Python · ~75 min

  10. Cold Start Mitigation for Serverless LLMs

    A 20 GB model image takes 5-10 minutes (7B) to 20+ minutes (70B) to go from cold to serving. In a true serverless world, that is not a warm-up — it is an outage.

    Learn · Python · ~60 min

  11. Multi-Region LLM Serving and KV Cache Locality

    Round-robin load balancing is actively harmful for cached LLM inference. A request that does not land on the node holding its prefix pays full prefill cost — roughly 800 ms at P50 on a long prompt…

    Learn · Python · ~60 min

  12. Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson

    The core edge constraint is memory bandwidth, not compute. Mobile DRAM sits at 50-90 GB/s; datacenter HBM3 clears 2-3 TB/s — a 30-50x gap. Decode is memory-bound so the gap is decisive.

    Learn · Python · ~60 min

  13. LLM Observability Stack Selection

    The 2026 observability market splits into two categories. Development platforms (LangSmith, Langfuse, Comet Opik) bundle monitoring with evals, prompt management, session replays.

    Learn · Python · ~60 min

  14. Prompt Caching and Semantic Caching Economics

    Pricing snapshot dated 2026-04. Numeric claims below reflect vendor rate cards captured at this lesson's publication; verify against the linked docs before quoting them downstream.

    Learn · Python · ~60 min

  15. Batch APIs — the 50% Discount as Industry Standard

    Every major provider ships an async batch API with a 50% discount and 24-hour turnaround. OpenAI, Anthropic, Google, and most of the inference platforms (Fireworks batch tier, Together batch)…

    Learn · Python · ~45 min

  16. Model Routing as a Cost-Reduction Primitive

    A dynamic broker evaluates every request (task type, token length, embedding similarity, confidence) and sends simple queries to a cheap model, escalating complex ones to a frontier model.

    Learn · Python · ~60 min

  17. Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d

    Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…

    Learn · Python · ~75 min

  18. Production Serving Stack — KV Offloading and Cache-Aware Routing

    A production serving stack wires router, engines, and observability into one Kubernetes deployment — and treats KV cache as a resource that can leave the GPU.

    Learn · Python · ~60 min

  19. AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost

    A gateway sits between your apps and model providers. Core features are provider routing, fallback, retries, rate limiting, secret references, observability, guardrails.

    Learn · Python · ~60 min

  20. Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs

    LLM rollouts combine the hardest parts of software deployment: no unit tests, diffuse failure modes, delayed signals. The sequence is (1) shadow mode — duplicate prod requests to candidate model,…

    Learn · Python · ~60 min

  21. A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem

    Traditional A/B testing was not built for non-deterministic LLMs. The critical distinction: evals answer "can the model do the job?" A/B tests answer "do users care?" Both are required; shipping on…

    Learn · Python · ~60 min

  22. Load Testing LLM APIs — Why k6 and Locust Lie

    Traditional load testers were not designed for streaming responses, variable output lengths, token-level metrics, or GPU saturation. Two traps bite most teams.

    Build · Python · ~75 min

  23. SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection

    AI SRE uses LLMs grounded in infrastructure data (logs, runbooks, service topology) via RAG to automate investigation, documentation, and coordination phases.

    Learn · Python · ~60 min

  24. Chaos Engineering for LLM Production

    Chaos engineering for LLMs is its own discipline in 2026. Prerequisites before running experiments in production: defined SLI/SLO, trace+metric+log observability, automated rollback, runbooks,…

    Learn · Python · ~60 min

  25. Security — Secrets, API Key Rotation, Audit Logs, Guardrails

    Eliminate secret sprawl via centralized vaults (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault). Never store credentials in config files, env files in VCS, spreadsheets.

    Learn · Python · ~60 min

  26. Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001

    Multi-framework coverage is table stakes for 2026 enterprise deals. EU AI Act: in force since August 1, 2024. Most high-risk requirements enforce August 2, 2026.

    Learn · Python · ~60 min

  27. FinOps for LLMs — Unit Economics and Multi-Tenant Attribution

    Traditional FinOps breaks on LLM spend. Costs are token-transactions, not resource-uptime. Tags don't map — an API call is a transaction, not an asset.

    Learn · Python · ~60 min

  28. Self-Hosted Serving Selection — Matching Engine to Hardware and Scale

    Engine selection is a function of hardware, scale, and ecosystem — not a leaderboard read. Four engines dominate self-hosted inference in 2026: llama.cpp, Ollama, vLLM, SGLang, with TGI trailing in…

    Learn · Python · ~45 min

Glossary terms in this phase

  • Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
  • AutoscalingA control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within…
  • CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
  • Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
  • Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
  • Dynamic BatchingA runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed…
  • Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
  • GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
  • Incident ResponseThe coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
  • MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…
  • PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
  • Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
  • QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
  • RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
  • Service Level Objective (SLO)A target range or threshold for a service-level indicator over a stated population and measurement window.
  • Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
  • Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…
  • Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
  • Zero TrustA security model that grants no implicit trust from network location or asset ownership and instead evaluates each access request against…

Frequently asked questions

How many lessons are in Phase 17: Infrastructure & Production?

Phase 17 has 28 lessons: 27 Learn lessons and 1 Build lesson. The lesson code uses Python.

What should I know before I start Phase 17?

The phase guide gives these prerequisites: Phase 11 LLM Engineering and Phase 13 Tools and Protocols. In the course roadmap, this phase builds on Phase 14: Agent Engineering.

Is Phase 17 free?

Yes. All 28 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.

How long does Phase 17 take?

The time estimates of all 28 lessons add up to about 29 hours.

What comes after Phase 17?

Phase 19: Capstone Projects builds on this phase.