Back to archive
#ai#llm#glossary#aigen

Continuous Batching

A server generates two responses at once. One finishes quickly; the other is long. If the next user must wait for the entire pair to finish, the freed slot is wasted for many steps.

Continuous Batching allows the group of requests being served to change between successive model iterations. A completed response leaves the batch — a group processed together — and a new request can take its place.

The long response continues to grow while the short one is replaced by the next. A scheduler, a program that plans the work, selects the participants. There is no need to wait for the slowest member of the entire previous group.

Using slots this way can increase throughput, but does not guarantee a shorter wait for every user. Request lengths, memory limits, and scheduling rules matter. With similar lengths, the advantage over a fixed group may be small.

Mechanism source: Orca, §3, “Iteration-level scheduling”.

Mechanism and details

A press with two tracks: a long turquoise strip continues, while an orange strip takes the place of a short ochre one.

This is especially useful in Decode: one person needs a short response, another several hundred tokens. Merely gathering several requests before running the model does not ensure that an active batch's membership can change. A new request also needs Prefill; how this stage is handled alongside generation depends on the engine's implementation.

Does C have to wait for A?

An original experiment: two slots are available. A and B are ready at time 0 and need six and two rounds respectively. C becomes ready at time 1 and needs two rounds. Both variants perform the same computations, but in a fixed group, B's slot does not accept C until A finishes.

With Continuous Batching, C starts at time 2 and finishes at 4; with a fixed group, it starts at 6 and finishes at 8. Switch B to six rounds: in this case, early release of a slot disappears, and both strategies produce the same schedule.

Rounds here have an illustrative, equal duration. Prefill and transfers are outside the experiment; “ready” means ready for the displayed rounds. On a GPU, iteration time depends on its workload, so the number of occupied slots is not the device-utilization percentage.

The ability to join a request does not guarantee an immediate start: memory limits and scheduler policy apply (Orca, §4.2). PagedAttention helps manage memory but does not itself determine serving order.

I use AI-generated content as part of my daily learning process.