Continuous batching schedules active sequences at each decoding iteration, allowing new requests to join and completed requests to leave without waiting for a fixed batch.
A family may send a voice request, document question, and coding prompt to one local model within seconds. Their prompts and output lengths differ, so a fixed batch wastes slots while shorter responses wait for the longest. Continuous batching keeps rebuilding useful work around the model, but it matters only when requests overlap and the server has enough memory and scheduling headroom.
Iteration-Level Scheduling Is the Defining Idea
Autoregressive decoding advances each active sequence by roughly one token per model iteration. A continuous scheduler selects the runnable sequences for the next iteration, admits new requests when capacity opens, and removes sequences immediately after completion.
The Orca paper introduced iteration-level scheduling at the granularity of model iterations rather than whole requests. Selective batching then groups compatible operations while leaving request-specific work separate. This distinction remains visible during later household testing.
This is not the same as streaming tokens to a user. Streaming changes when output is delivered; continuous batching changes how several requests share model execution internally. The intermediate result must remain inspectable before automation follows.
It Differs From Static and Arrival-Window Batching
Static batching locks a fixed group together and often pads shorter sequences until the longest finishes. Arrival-window or dynamic batching waits briefly to collect requests, but may still execute the resulting group as one unit. Continuous batching revisits membership every iteration.
The vLLM paper pairs iteration-level scheduling with paged KV-cache management so changing sequence sets do not require rigid contiguous reservations. Scheduling and memory management are complementary rather than interchangeable features. That boundary should be measured separately under realistic operating conditions.
More active sequences amortize weight reads and can improve throughput, yet each request competes for KV memory and compute. A larger live batch is not automatically better for latency or fairness. The practical consequence appears when several sources compete for limited context.
Concurrency, Not Model Size Alone, Creates the Benefit
A single interactive user may see little gain because no second request is available to fill unused capacity. Benefits emerge with overlapping household users, agent branches, background summaries, or several applications sharing one resident model.
Sarathi-Serve analyzes how prefill interference can disrupt decode latency and uses chunked prefills to make mixed scheduling more predictable. The result shows that admission policy matters alongside the continuous-batching label. This dependency should remain explicit in the final interface.
The failure boundary is memory pressure or aggressive admission that inflates time per output token and tail latency. When KV cache fills, preemption, swapping, or recomputation can erase throughput gains and make interactive service unstable.
Determine Whether Concurrent Demand Justifies It
Replay one, two, four, and eight overlapping requests with realistic prompt and output lengths. Record throughput, time to first token, time per output token, p95 completion time, KV utilization, preemptions, and fairness by request class.
Compare behavior with continuous batching gaps. Repeat with continuous batching disabled or a fixed-batch baseline while holding model, quantization, context limits, and hardware constant. The result must therefore be checked against the original evidence.
Use continuous batching when overlap produces material throughput or capacity gains without violating interactive tail latency. If requests rarely overlap, prioritize model residency and startup latency before adding scheduler complexity. This distinction remains visible during later household testing.
Tech & AI HUB
More to Read

What Is Embedding Drift, and When Does a Private Search Index Need Rebuilding?
Decode model, preprocessing, corpus, and query drift; distinguish monitoring from incompatibility; and decide when a private index needs rebuilding.

What Is Tokenizer Compatibility, and Why Can It Break Model Switching?
Decode vocabulary identity, special-token semantics, chat templates, cached tokens, adapters, and compatibility checks for local model switching.

What Is Model Residency, and When Should a Local AI Service Keep Weights Loaded?
Decode weight residency, cache levels, cold starts, eviction, multiplexing, memory pressure, and when a home AI service should stay warm.

