bk99.de entertain the web since 1997

8-bit matrix multiplication for large transformers

Summary

The bitsandbytes integration loads large models with 8-bit weights and handles sensitive outliers separately. This lowers memory requirements while important numerical parts keep higher precision. LLM.int8 combines 8-bit matrix multiplication with higher-precision handling of outliers.

Ideas

  • Mixed precision works better when only the outliers, rather than all values, are processed expensively.
  • Quantisation shifts the limit of what can be used experimentally on existing hardware.

Insights

  • Lower precision is a system decision between quality, memory, kernels and target hardware.

Facts

  • Transformers and Accelerate integrate the loading of quantised models.

Critique

  • The benefits depend on supported GPUs, operators and the model’s distribution of outliers.

Recommendations

  • Compare memory, throughput and task quality against a higher-precision reference.

References

Read the original article on Hugging Face

Search the Web Archive