Accelerated pipelines with Optimum
Summary
Optimum integrates optimised runtimes into familiar Transformers pipelines. This allows models to be accelerated with ONNX Runtime without rewriting the whole application. Optimum offers ORTModel classes for ONNX Runtime.
Ideas
- Compatible APIs lower the cost of trying out an optimised runtime in practice.
- Export and optimisation must remain reproducible as a build step of their own.
Facts
- Optimised models can be used through the Transformers pipeline interface.
Critique
- API compatibility guarantees neither identical results nor acceleration for every model.
Recommendations
- Compare outputs and numerical tolerances before and after export and quantisation.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.