bk99.de entertain the web since 1997

Millisecond inference on modern CPUs

Summary

The case study optimises transformer inference on CPUs with ONNX Runtime, quantisation and dynamic batching. It shows when existing server hardware can be a practical alternative to GPUs. The platform runs optimised transformer models on modern CPUs.

Ideas

  • Small models often benefit more from low system latency than from maximum accelerator performance.
  • Quantisation, batching and runtime have to be tuned together for the service goal.

Facts

  • The case study measures latencies in the millisecond range for selected models.

Critique

  • The advertised latency only applies to the models, CPUs and load conditions shown.

Recommendations

  • Test CPU and GPU variants with your real request size, parallelism and power consumption.

References

Read the original article on Hugging Face

Search the Web Archive