Asynchrony in continuous batching
Summary
The article extends continuous batching with asynchronous preparation and output so that CPU work blocks the GPU scheduler less. Overlapping pipeline phases increase utilisation with many simultaneous requests. Tokenisation, scheduling and output can be overlapped with GPU computation.
Ideas
- Accelerators often wait for CPU and network phases rather than for additional model computation.
- Asynchrony increases throughput but also the number of possible ordering and failure states.
Facts
- Tokenisation, scheduling and output can be overlapped with GPU computation.
- Continuous batching continuously adds requests to active batches.
Critique
- More parallelism makes reproducibility, backpressure and clean abort handling harder.
Recommendations
- Use tracing across queue, CPU, GPU and streaming to find hidden waiting times.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.