Phase 17 · Infrastructure & Production
Learn LLM Serving and AI Infrastructure: 28 Free Lessons
Ship AI to the real world. Scale, monitor, optimize.
- 28 lessons
- 27 learn
- 1 build
- ~29 hours
- Python
Start Phase 17
First lesson Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI
Run this command from the repository root:
python3 phases/17-infrastructure-and-production/01-managed-llm-platforms/code/main.pyKeep the command, exit code, cost and latency comparison, redundancy uplift, and a short decision record naming the workload assumptions behind your choice.
All 28 lessons in Phase 17
- Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI
Three hyperscalers, three distinct strategies. AWS Bedrock is a model marketplace — Claude, Llama, Titan, Stability, Cohere behind one API.
- Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale
The 2026 inference market is no longer GPU time rental. It bifurcates into custom silicon (Groq, Cerebras, SambaNova), GPU platforms (Baseten, Together, Fireworks, Modal), and API-first marketplaces…
- GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling
Three layers, not one. Karpenter provisions nodes dynamically (under one minute, 40% faster than Cluster Autoscaler). KAI Scheduler handles gang scheduling, topology awareness, and hierarchical…
- Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…
- EAGLE-3 Speculative Decoding in Production
Speculative decoding pairs a fast draft model with the target model. The draft proposes K tokens; the target verifies in a single forward; accepted tokens are free.
- Prefix-Cache Serving — RadixAttention and KV Reuse
Treat the KV cache as a first-class, reusable resource stored in a radix tree, and scheduling changes with it: instead of FCFS (first-come, first-served) as vLLM schedules, a cache-aware scheduler…
- Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell
Hardware-specialized inference compilation trades portability for throughput, and TensorRT-LLM — NVIDIA-only, tuned for Blackwell — is the clearest example of the trade paying off.
- Inference Metrics — TTFT, TPOT, ITL, Goodput, P99
Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.
- Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4
Quantization format is not a universal choice — it is a function of hardware, serving engine, and workload. GGUF Q4KM or Q5KM owns CPU and edge, delivered through llama.cpp and Ollama.
- Cold Start Mitigation for Serverless LLMs
A 20 GB model image takes 5-10 minutes (7B) to 20+ minutes (70B) to go from cold to serving. In a true serverless world, that is not a warm-up — it is an outage.
- Multi-Region LLM Serving and KV Cache Locality
Round-robin load balancing is actively harmful for cached LLM inference. A request that does not land on the node holding its prefix pays full prefill cost — roughly 800 ms at P50 on a long prompt…
- Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson
The core edge constraint is memory bandwidth, not compute. Mobile DRAM sits at 50-90 GB/s; datacenter HBM3 clears 2-3 TB/s — a 30-50x gap. Decode is memory-bound so the gap is decisive.
- LLM Observability Stack Selection
The 2026 observability market splits into two categories. Development platforms (LangSmith, Langfuse, Comet Opik) bundle monitoring with evals, prompt management, session replays.
- Prompt Caching and Semantic Caching Economics
Pricing snapshot dated 2026-04. Numeric claims below reflect vendor rate cards captured at this lesson's publication; verify against the linked docs before quoting them downstream.
- Batch APIs — the 50% Discount as Industry Standard
Every major provider ships an async batch API with a 50% discount and 24-hour turnaround. OpenAI, Anthropic, Google, and most of the inference platforms (Fireworks batch tier, Together batch)…
- Model Routing as a Cost-Reduction Primitive
A dynamic broker evaluates every request (task type, token length, embedding similarity, confidence) and sends simple queries to a cheap model, escalating complex ones to a frontier model.
- Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d
Prefill is compute-bound; decode is memory-bound. Running both on the same GPU wastes one resource. Disaggregation splits them onto separate pools and transfers KV cache between them over NIXL…
- Production Serving Stack — KV Offloading and Cache-Aware Routing
A production serving stack wires router, engines, and observability into one Kubernetes deployment — and treats KV cache as a resource that can leave the GPU.
- AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost
A gateway sits between your apps and model providers. Core features are provider routing, fallback, retries, rate limiting, secret references, observability, guardrails.
- Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs
LLM rollouts combine the hardest parts of software deployment: no unit tests, diffuse failure modes, delayed signals. The sequence is (1) shadow mode — duplicate prod requests to candidate model,…
- A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem
Traditional A/B testing was not built for non-deterministic LLMs. The critical distinction: evals answer "can the model do the job?" A/B tests answer "do users care?" Both are required; shipping on…
- Load Testing LLM APIs — Why k6 and Locust Lie
Traditional load testers were not designed for streaming responses, variable output lengths, token-level metrics, or GPU saturation. Two traps bite most teams.
- SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection
AI SRE uses LLMs grounded in infrastructure data (logs, runbooks, service topology) via RAG to automate investigation, documentation, and coordination phases.
- Chaos Engineering for LLM Production
Chaos engineering for LLMs is its own discipline in 2026. Prerequisites before running experiments in production: defined SLI/SLO, trace+metric+log observability, automated rollback, runbooks,…
- Security — Secrets, API Key Rotation, Audit Logs, Guardrails
Eliminate secret sprawl via centralized vaults (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault). Never store credentials in config files, env files in VCS, spreadsheets.
- Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001
Multi-framework coverage is table stakes for 2026 enterprise deals. EU AI Act: in force since August 1, 2024. Most high-risk requirements enforce August 2, 2026.
- FinOps for LLMs — Unit Economics and Multi-Tenant Attribution
Traditional FinOps breaks on LLM spend. Costs are token-transactions, not resource-uptime. Tags don't map — an API call is a transaction, not an asset.
- Self-Hosted Serving Selection — Matching Engine to Hardware and Scale
Engine selection is a function of hardware, scale, and ecosystem — not a leaderboard read. Four engines dominate self-hosted inference in 2026: llama.cpp, Ollama, vLLM, SGLang, with TGI trailing in…
Glossary terms in this phase
- Audit LogA durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it…
- AutoscalingA control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within…
- CalibrationThe agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- Chunked PrefillA serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with…
- Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
- Dynamic BatchingA runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed…
- Evaluation (Eval)A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and…
- FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
- GoodputThe rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency…
- GuardrailsSystem controls that constrain inputs, tool use, outputs, permissions, and escalation.
- Incident ResponseThe coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- Inter-Token Latency (ITL)The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token…
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
- MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…
- PrefillThe initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for…
- Prefix CachingReusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix…
- QuantizationRepresenting weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost.
- RollbackRestoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
- Service Level Objective (SLO)A target range or threshold for a service-level indicator over a stated population and measurement window.
- Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
- Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- Tail LatencyThe latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time…
- Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- TraceA correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
- Zero TrustA security model that grants no implicit trust from network location or asset ownership and instead evaluates each access request against…
Frequently asked questions
How many lessons are in Phase 17: Infrastructure & Production?
Phase 17 has 28 lessons: 27 Learn lessons and 1 Build lesson. The lesson code uses Python.
What should I know before I start Phase 17?
The phase guide gives these prerequisites: Phase 11 LLM Engineering and Phase 13 Tools and Protocols. In the course roadmap, this phase builds on Phase 14: Agent Engineering.
Is Phase 17 free?
Yes. All 28 lessons are free to read on this site, and you do not need an account. The lesson code is open source under the MIT license.
How long does Phase 17 take?
The time estimates of all 28 lessons add up to about 29 hours.
What comes after Phase 17?
Phase 19: Capstone Projects builds on this phase.