Millisecond inference on modern CPUs
Summary
The case study optimises transformer inference on CPUs with ONNX Runtime, quantisation and dynamic batching. It shows when existing server hardware can be a practical alternative to GPUs. The platform runs optimised transformer models on modern CPUs.
Ideas
- Small models often benefit more from low system latency than from maximum accelerator performance.
- Quantisation, batching and runtime have to be tuned together for the service goal.
Facts
- The case study measures latencies in the millisecond range for selected models.
Critique
- The advertised latency only applies to the models, CPUs and load conditions shown.
Recommendations
- Test CPU and GPU variants with your real request size, parallelism and power consumption.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.