AI-native development · Glossary term

What is Rate Limit?

A policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.

Why does Rate Limit matter?

It protects providers and your own system from overload, uncontrolled spend, and unfair resource use.

Rate Limit in practice

Enforce per-tenant token and concurrency limits, read provider retry metadata, and queue or reject excess work predictably.

What is the common confusion about Rate Limit?

A rate limit controls allowed usage. Backpressure propagates downstream capacity constraints through a system.

Learn Rate Limit in the course

Lessons that name Rate Limit in a title or section

  • Stateless MCP Gateways and Registry Admission

    A gateway should make every route explicit. The 2026-07-28 protocol gives it method, name, version, capability, identity, cache, and trace boundaries without a transport session.

    Phase 13: Tools & Protocols

  • LLM Routing Layer — LiteLLM, OpenRouter, Portkey

    Provider lock-in is expensive. Different tool-calling workloads suit different models. Routing gateways give one API surface, retries, failover, cost tracking, and guardrails.

    Phase 13: Tools & Protocols

Covered in Phase 13: Tools & Protocols.

  • BackpressureA flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.
  • Retry with BackoffRepeating a failed transient operation after progressively longer delays, usually with randomized jitter and a strict retry limit.
  • Circuit BreakerA reliability control that temporarily stops calls to a dependency after failures cross a threshold, then probes whether the dependency…
  • Admission ControlA pre-acceptance gate that decides whether a request may enter a bounded queue or service under the system's current capacity, priority,…
  • Continuous BatchingA serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed…
  • Load SheddingDeliberately rejecting, dropping, or cancelling selected work at one or more overload boundaries when demand exceeds the capacity…
  • Model RouterA component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and…
  • Retry BudgetA bound on retry traffic, usually expressed relative to original requests or over a time window, that prevents retries from consuming…

More terms in AI-native development

Open the AI-native development list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.