epoll or io_uring? Asynchronous I/O compared
Summary
The author of TinyGate compares epoll and io_uring using a reverse proxy with many simultaneous connections. epoll appeared in the Linux kernel in 2002. One io_uring_enter() call can batch several submissions and completions.
Ideas
- epoll reports readiness; applications then carry out read and write operations themselves.
- io_uring reports completed operations via shared ring buffers.
- Batching spreads the cost of a system call over several I/O operations.
- SQPOLL reduces system calls but permanently ties up the computing time of a kernel thread.
- Switching to io_uring can force a complete change of architecture.
- A realistic proxy benchmark reveals limits better than isolated micro-benchmarks.
Insights
- Fewer context switches only pay off with sufficiently high concurrency.
- Performance gains often shift complexity from system calls into state and error management.
- The best I/O interface depends more on the load profile than on age.
Facts
- io_uring arrived in 2019 with Linux 5.1.
- Submission and completion queues reside in memory shared between application and kernel.
Recommendations
- Measure both models with your real connection and payload profile.
- Only enable SQPOLL after measuring the additional CPU consumption.
- Plan migrations as an architecture project rather than swapping individual system calls.
References
Links to the original source and the Web Archive open in a new tab.