Continuous batching explained from scratch
Summary
The article builds a scheduler that adds new requests to running batches between individual decoding steps. This keeps accelerators better utilised despite answers of different lengths. Continuous batching schedules requests at token level instead of per fully completed batch.
Ideas
- Static batches waste capacity as soon as individual sequences finish earlier.
- Scheduling affects throughput, tail latency and fairness at the same time.
Facts
- Continuous batching schedules requests at token level instead of per fully completed batch.
- Finished sequences free up cache and compute slots immediately.
Critique
- Maximum throughput can starve individual long or low-priority requests.
Recommendations
- Simulate short and long requests together and define explicit fairness and abort rules.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.