systems.architecture

Low-Level Linux Kernel Tuning for Ultra-Low-Latency Network Applications

Tail latency is an engineering budget, not a configuration flag. Measure first, then remove the largest source of variance.

Applications with microsecond-scale latency budgets — trading systems, telemetry ingestion, real-time control planes — are dominated by variance rather than average throughput. Reducing the ninety-ninth percentile means eliminating sources of jitter across the network interface, the interrupt path, the scheduler, and the memory subsystem. The work is empirical: nothing here should be applied without before-and-after measurement on the actual workload.

Measure the right thing first

Hardware timestamping at the network interface gives a reference point independent of the application clock. Combined with high-resolution histograms, it separates network delay from stack delay from application delay.

Track the distribution, not the mean. A change that improves average latency while widening the tail is a regression for this class of system.

Control interrupts and CPU placement

Pin interrupt handling for the receiving queue to a specific core and pin the consuming thread to a nearby core on the same NUMA node. Disable the interrupt balancing daemon so the assignment holds under load.

Isolate those cores from the general scheduler and move housekeeping work — timers, kernel threads, background daemons — elsewhere. A single migration event costs more than most micro-optimizations save.

Choose a receive strategy deliberately

Interrupt coalescing improves throughput and hurts latency; for latency-sensitive paths, reduce or disable it and accept the CPU cost. Busy polling removes wakeup delay entirely at the price of a fully consumed core.

When the in-kernel path cannot meet the budget, kernel bypass frameworks or an in-kernel express data path move processing closer to the wire. Both add operational complexity and should be adopted only when measurement shows the stack itself is the bottleneck.

Remove memory and power variance

Huge pages reduce translation lookaside buffer misses on large working sets. Pre-faulting and locking critical memory avoids page faults on the hot path. Allocate buffers on the NUMA node local to the interface.

Disable deep C-states and frequency scaling on the isolated cores. Waking a core from a deep sleep state can cost more than the entire packet-processing budget. Record every setting in configuration management so a rebuilt host does not silently lose the tuning.

key takeaways

  • Use hardware timestamps and full histograms, never averages alone.
  • Pin interrupts and consumer threads to isolated, NUMA-local cores.
  • Trade coalescing and polling settings against CPU cost explicitly.
  • Lock memory, use huge pages, and allocate on the local node.
  • Disable deep C-states and codify every setting in configuration management.