bk99.de entertain the web since 1997

Making LLMs lighter with AutoGPTQ

Summary

The AutoGPTQ integration quantises already trained language models to a few bits and integrates them directly into Transformers. This considerably lowers the memory requirement for local inference. GPTQ typically quantises weights to four bits after training.

Ideas

  • Post-training quantisation makes large models accessible without retraining them completely.
  • Calibration data influences which weights are preserved at low precision.

Insights

  • Lower precision is a system decision between quality, memory, kernels and target hardware.

Facts

  • Quantised models can be loaded and run through the Transformers API.

Critique

  • A low bit width can degrade certain abilities unevenly and needs suitable kernels.

Recommendations

  • Test several calibration samples and measure the quality loss on your own tasks.

References

Read the original article on Hugging Face

Search the Web Archive