8-bit matrix multiplication for large transformers
Summary
The bitsandbytes integration loads large models with 8-bit weights and handles sensitive outliers separately. This lowers memory requirements while important numerical parts keep higher precision. LLM.int8 combines 8-bit matrix multiplication with higher-precision handling of outliers.
Ideas
- Mixed precision works better when only the outliers, rather than all values, are processed expensively.
- Quantisation shifts the limit of what can be used experimentally on existing hardware.
Insights
- Lower precision is a system decision between quality, memory, kernels and target hardware.
Facts
- Transformers and Accelerate integrate the loading of quantised models.
Critique
- The benefits depend on supported GPUs, operators and the model’s distribution of outliers.
Recommendations
- Compare memory, throughput and task quality against a higher-precision reference.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.