bk99.de entertain the web since 1997

How BLOOM inference was optimised

Summary

The report follows the optimisation of a 176-billion-parameter model from memory planning to distributed execution. It shows that production inference is a systems problem of parallelisation, kernels and communication. BLOOM is too large for the memory of a single common GPU.

Ideas

  • With very large models, the distribution of weights and cache determines the possible parallelism.
  • Optimisation needs reliable profiles instead of guesses about the most expensive part.

Facts

  • The solution distributes the model computation across several accelerators.

Critique

  • The architecture shown is optimised for BLOOM and the hardware used, not universally.

Recommendations

  • Profile preprocessing, prefill, decoding, communication and memory separately.

References

Read the original article on Hugging Face

Search the Web Archive