Making LLMs lighter with AutoGPTQ
Summary
The AutoGPTQ integration quantises already trained language models to a few bits and integrates them directly into Transformers. This considerably lowers the memory requirement for local inference. GPTQ typically quantises weights to four bits after training.
Ideas
- Post-training quantisation makes large models accessible without retraining them completely.
- Calibration data influences which weights are preserved at low precision.
Insights
- Lower precision is a system decision between quality, memory, kernels and target hardware.
Facts
- Quantised models can be loaded and run through the Transformers API.
Critique
- A low bit width can degrade certain abilities unevenly and needs suitable kernels.
Recommendations
- Test several calibration samples and measure the quality loss on your own tasks.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.