Infrastructure & serving · Glossary term
What is Autoscaling?
A control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within configured bounds.
Why does Autoscaling matter?
AI workloads can change faster than manual provisioning, but scaling decisions must account for model-load time, accelerator availability, queueing, and request cost.
Autoscaling in practice
Scale from a demand signal tied to useful work, set minimum warm capacity, bound scale-down churn, and verify that new replicas pass readiness checks before receiving traffic.
What is the common confusion about Autoscaling?
Autoscaling adds or removes capacity. It does not make an overloaded dependency faster or guarantee that enough hardware can be acquired in time.
Learn Autoscaling in the course
Start with
- GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling
Three layers, not one. Karpenter provisions nodes dynamically (under one minute, 40% faster than Cluster Autoscaler). KAI Scheduler handles gang scheduling, topology awareness, and hierarchical…
Taught in Phase 17: Infrastructure & Production.
Related terms
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
- SaturationThe degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
- Readiness ProbeA diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
- BackpressureA flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.