Quanto as a PyTorch backend for quantisation
Summary
Quanto brings weight and activation quantisation into a PyTorch-centred workflow. Models therefore remain easier to examine and do not have to be exported to a foreign runtime straight away. Quanto is a quantisation backend for Hugging Face Optimum.
Ideas
- Quantisation within PyTorch shortens the path between experiment, measurement and adjustment.
- Weights and activations often need different precision strategies.
Insights
- Lower precision is a system decision between quality, memory, kernels and target hardware.
Facts
- It supports several number formats for weights and activations.
Critique
- Backend support and real acceleration differ depending on device and operator.
Recommendations
- Record the error and runtime per layer before you settle on a global quantisation configuration.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.