Infrastructure & serving · Glossary term
What is Model Serving?
The runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and returns results under an operational contract.
Why does Model Serving matter?
A capable model can still produce an unreliable product when queueing, batching, placement, versioning, cancellation, and response boundaries are not engineered explicitly.
Model Serving in practice
Pin model and tokenizer versions, validate request limits, expose readiness and latency signals, control concurrency, and test rollback before routing production traffic.
What is the common confusion about Model Serving?
Model serving is broader than calling inference once and narrower than the complete application, which may also include retrieval, tools, policy, and user state.
Learn Model Serving in the course
Start with
- Self-Hosted Serving Selection — Matching Engine to Hardware and Scale
Engine selection is a function of hardware, scale, and ecosystem — not a leaderboard read. Four engines dominate self-hosted inference in 2026: llama.cpp, Ollama, vLLM, SGLang, with TGI trailing in…
Taught in Phase 17: Infrastructure & Production.
Related terms
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
- AutoscalingA control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within…
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- Disaggregated ServingA serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state…
- Expert ParallelismDistributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected…
- Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…
- Pipeline ParallelismPartitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
- Readiness ProbeA diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
- Shadow TrafficA copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user…
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.