bk99.de entertain the web since 1997

Quantisation backends for diffusion models

Summary

The article compares several quantisation routes for Diffusers and their effect on memory, speed and image quality. Because of their various components, diffusion pipelines have different requirements from pure LLMs. Diffusers supports several backends via quantised model components.

Ideas

  • Text encoder, transformer and VAE do not necessarily need the same quantisation strategy.
  • Memory savings can be paid for with dequantisation or slow kernels.

Insights

  • Lower precision is a system decision between quality, memory, kernels and target hardware.

Facts

  • The comparisons look at memory consumption, runtime and output quality.

Critique

  • Visual spot checks can easily miss rare quality losses or prompt dependencies.

Recommendations

  • Quantise pipeline components individually and check image artefacts with fixed seeds.

References

Read the original article on Hugging Face

Search the Web Archive