Infrastructure & serving · Glossary term
What is Continuous Batching?
A serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed batch to finish.
Why does Continuous Batching matter?
Autoregressive requests produce different output lengths, so continuous batching can keep accelerators utilized without forcing short requests to wait for the longest one.
Continuous Batching in practice
Admit new requests when capacity becomes available, track per-request latency, and apply backpressure when the live batch or KV-cache budget is full.
What is the common confusion about Continuous Batching?
Continuous batching is an inference scheduling policy, not gradient accumulation or a training batch-size technique.
Learn Continuous Batching in the course
Lessons that name Continuous Batching in a title or section
- Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill
Modern serving-engine throughput rests on three compounding defaults, not a single trick. PagedAttention is always on. Continuous batching injects new requests into the active batch between decode…
- KV Cache, Flash Attention & Inference Optimization
Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…
- Inference Optimization
Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.
Covered in Phase 07: Transformers Deep Dive, Phase 10: LLMs from Scratch and Phase 17: Infrastructure & Production.
Related terms
- Dynamic BatchingA runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed…
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- BackpressureA flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.
- Rate LimitA policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.
Sources
More terms in Infrastructure & serving
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.