Infrastructure & serving · Glossary term

What is Autoscaling?

A control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within configured bounds.

Why does Autoscaling matter?

AI workloads can change faster than manual provisioning, but scaling decisions must account for model-load time, accelerator availability, queueing, and request cost.

Autoscaling in practice

Scale from a demand signal tied to useful work, set minimum warm capacity, bound scale-down churn, and verify that new replicas pass readiness checks before receiving traffic.

What is the common confusion about Autoscaling?

Autoscaling adds or removes capacity. It does not make an overloaded dependency faster or guarantee that enough hardware can be acquired in time.

Learn Autoscaling in the course

Start with

Taught in Phase 17: Infrastructure & Production.

  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
  • SaturationThe degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
  • Readiness ProbeA diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
  • BackpressureA flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.

Sources

More terms in Infrastructure & serving

Open the Infrastructure & serving list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.