bk99.de entertain the web since 1997

An overview of quantisation methods in Transformers

Summary

The article compares integrated methods such as bitsandbytes, GPTQ and AWQ in terms of memory, hardware and deployment path. It helps to understand quantisation as a trade-off rather than a single switch. Transformers supports several quantisation backends through their own configurations.

Ideas

  • Quantisation methods differ in calibration, bit width, kernel support and changeability.
  • The smallest model format is not automatically the fastest on a given hardware.

Insights

  • Lower precision is a system decision between quality, memory, kernels and target hardware.

Facts

  • Some methods target inference, others also allow adapter training.

Critique

  • Comparison figures age quickly because kernels and hardware support change constantly.

Recommendations

  • Create a decision matrix from quality, memory, latency, platform and maintainability.

References

Read the original article on Hugging Face

Search the Web Archive