How BLOOM inference was optimised
Summary
The report follows the optimisation of a 176-billion-parameter model from memory planning to distributed execution. It shows that production inference is a systems problem of parallelisation, kernels and communication. BLOOM is too large for the memory of a single common GPU.
Ideas
- With very large models, the distribution of weights and cache determines the possible parallelism.
- Optimisation needs reliable profiles instead of guesses about the most expensive part.
Facts
- The solution distributes the model computation across several accelerators.
Critique
- The architecture shown is optimised for BLOOM and the hardware used, not universally.
Recommendations
- Profile preprocessing, prefill, decoding, communication and memory separately.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.