Loading very large models with Accelerate
Summary
Accelerate distributes model parts across GPU, CPU and disk instead of first building the whole model in memory. Empty initialisation and automatic device maps avoid unnecessary peaks. init_empty_weights first creates the model structure without parameter data.
Ideas
- Memory peaks during loading can be larger than the later inference requirement.
- Hierarchical offloading trades capacity for transfer time and complexity.
Facts
- device_map can distribute layers across GPU, CPU and disk.
Critique
- A model that loads technically can become too slow for interactive use because of offloading.
Recommendations
- Plan the memory budget and transfer paths per layer and measure the resulting token latency.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.