Linux Epoll, Edge-Triggered & Event Loop Architecture Studio
Architect, simulate, and benchmark high-throughput event loops under Linux I/O multiplexing. Compare Level-Triggered (LT) versus Edge-Triggered (ET) notification semantics, trace the kernel Red-Black Tree and interrupt-driven ready-list data structures, simulate non-blocking EAGAIN drain loops, analyze EPOLLEXCLUSIVE thundering herd protection, and synthesize production C, Rust, and Go netpoller blueprints.
Interactive Epoll Reactor & Notification Simulator
Simulate incoming byte arrivals on active TCP sockets. Toggle between Level-Triggered (LT) and Edge-Triggered (ET) modes to observe how event wakeups, partial reads, and EAGAIN drain loops operate in production event loops.
SIMULATED MONITORED TCP SOCKETS (EPOLL_CTL_ADD)
EVENT LOOP & SYSCALL EXECUTION TRACE
Linux Kernel Epoll Data Structures & Lifecycle Architecture
Deep architectural dissection of struct eventpoll, the Red-Black Tree interest list, and interrupt callback readiness dispatching.
1. The Interest List: Red-Black Tree (struct eventpoll.rbr)
Every file descriptor registered via epoll_ctl(epfd, EPOLL_CTL_ADD, fd, &ev) is stored as an epitem node in a Red-Black Tree keyed by (struct file *, int fd).
- Search, insert, and remove complexity:
O(log N). - Allows registering 1,000,000+ sockets without degrading lookup speed.
- Persistent kernel structure: fds remain registered until explicitly removed with
EPOLL_CTL_DELor closed.
2. The Ready List: Doubly-Linked List (struct eventpoll.rdllist)
When an I/O event occurs (e.g. NIC receives packet, TCP stack processes payload), the driver invokes the kernel callback ep_poll_callback().
- The callback adds the corresponding
epitemto the doubly-linked ready list inO(1)time. - Wakes up threads sleeping in
epoll_wait()on the wait queue (wq). epoll_wait()copies only items fromrdllistdirectly to user space inO(k)time!
Linux Kernel fs/eventpoll.c Core Structures
The Multi-Threaded Thundering Herd Problem & EPOLLEXCLUSIVE
Understand how concurrent worker processes monitoring a shared listening socket trigger wake-up storms, and how EPOLLEXCLUSIVE and SO_REUSEPORT solve it.
Without EPOLLEXCLUSIVE (Legacy Flaw)
16 worker threads call epoll_wait() on the shared listen socket. When 1 client connects (SYN):
2. Thread 1 calls accept() → Returns client fd.
3. Remaining 15 threads call accept() → Return -1 (EAGAIN).
4. Waste: 15 context switches, L1 cache thrashing, CPU lock spin.
With EPOLLEXCLUSIVE (Linux 4.5+)
Each worker registers the listen socket with the EPOLLEXCLUSIVE flag set:
2. That thread calls accept() → Returns client fd.
3. Remaining 15 threads stay asleep in epoll_wait().
4. Benefit: Zero wasted CPU cycles, linear multicore scaling.
Essential Linux Kernel Epoll Tuning Parameters (/proc/sys)
| Kernel Parameter | Default | Recommended Production Value | Rationale |
|---|---|---|---|
| /proc/sys/fs/epoll/max_user_watches | ~1,000,000 | 16,777,216 | Limits max file descriptors any single user ID can monitor with epoll. Prevents ENOSPC errors on high-scale gateways. |
| /proc/sys/net/core/somaxconn | 4096 | 65535 | Maximum listen socket queue length. Essential to prevent TCP SYN drops under connection bursts. |
| /proc/sys/fs/file-max | System RAM dependent | 2,097,152 | System-wide limit on total open file descriptors across all processes. |
I/O Multiplexing Architecture Comparison: epoll vs io_uring vs select
Comprehensive architectural matrix contrasting Linux I/O models across system call overhead, complexity, memory transfer, and kernel requirements.
| Dimension | select() / poll() | Linux epoll (Reactor) | Linux io_uring (Proactor) |
|---|---|---|---|
| Time Complexity | O(N) Linear Scan | O(k) Active Events Only | O(k) Ring Buffer Completions |
| Kernel State Persistence | Stateless (resends full FD array every loop) | Stateful (Red-Black tree maintained in kernel) | Stateful (Memory-mapped ring buffers) |
| Syscall Overhead per Event | High (1 syscall to poll + 1 per active read) | Medium (1 epoll_wait + 1 read/write syscall) | Zero (with IORING_SETUP_SQPOLL kernel thread) |
| I/O Paradigm | Readiness Notification | Readiness Notification (Reactor Pattern) | Completion-based Asynchronous (Proactor) |
| Zero-Copy Support | None | None (Requires user/kernel buffer copy in read) | Full (Fixed buffers / Registered buffers) |
| Kernel Version Requirement | Ancient (Linux 1.0+) | Linux 2.5.44+ (Universal) | Linux 5.1+ (Production 5.15+ recommended) |
| Standard Industry Adopters | Legacy embedded tools | NGINX, Redis, Node.js (libuv), Netty, Envoy | High-performance databases, Ceph, modern Rust async |
Production Linux Epoll Reactor Blueprints
Battle-tested, memory-safe implementations of Edge-Triggered epoll servers in C, Rust (mio), and Go.
Critical Epoll Architecture Pitfalls & Starvation Bugs
read() will get stuck processing that single socket forever! The remaining 50,000 connections in the event loop starve. Solution: Cap the drain loop to a maximum batch quota (e.g. 64 KiB or 16 iterations), re-queue the socket on a ready deque, and yield control to epoll_wait().
EPOLLET but fail to set fcntl(fd, F_SETFL, flags | O_NONBLOCK), the mandatory drain loop will eventually call read() when the kernel buffer is empty. Instead of returning EAGAIN, the call blocks the entire operating system thread indefinitely! The event loop freezes.
struct file are closed, if a descriptor was duplicated via dup() or inherited by a fork() child, closing the user-space fd does NOT remove the watch from epoll! Events continue firing on closed descriptors. Solution: Always execute epoll_ctl(epfd, EPOLL_CTL_DEL, fd, NULL) before calling close(fd).