Tuning the kernel, after you have measured
Measure, change one thing, verify.
Every earlier lesson in this course measured something. This one changes something, and only after a measurement has named the limit. You will work through one tuning case from symptom to verified fix, make the change persistent without breaking something else, read the memory and block settings that are not sysctls, raise a service's file-descriptor limit where it actually applies, and see what TuneD, RHEL's tuning daemon, does to a host. A setting copied from an old guide can make a current kernel slower, so the method matters more than any value.
The loop: measure, change one thing, verify
The prediction step separates tuning from guessing: if you cannot say what a setting should do to the number you measured, you do not know that it is the limit. One change at a time keeps the cause visible, and the record lets the next engineer remove a setting that a kernel upgrade made unnecessary.
A measured case: an upload over a long path
The symptom: backups uploaded to a remote site over a 1 Gbit/s link run at about a quarter of the link's speed, with no errors and idle CPUs. To reproduce it safely, build the path between two network namespaces: a veth pair, and netem (the kernel's network emulator queueing discipline) adding 50 ms of delay in each direction and limiting the sender to 1 Gbit/s. Nothing outside the namespaces changes. iperf3 runs as a server in kt-b.
The round trip takes at least 100 ms (the first packet is slower while the neighbour entry is resolved). Now measure the upload from kt-a. -O 2 leaves the first two seconds of slow start out of the average.
About 246 Mbit/s, steady, and no retransmissions (Retr 0), so nothing is being lost. TCP can have at most one window of unacknowledged data in flight per round trip, so throughput is roughly bytes in flight divided by the round-trip time. Filling this path needs its bandwidth-delay product in flight: 1 Gbit/s × 0.1 s = 12.5 MB. Look at the connection while it runs; ss -tin prints TCP's internal state for each socket.
The first socket is iperf3's control connection (left out). For the data connection, three numbers decide the limit. cwnd:4818 segments of 1448 bytes is about 7 MB that congestion control would allow; snd_wnd:23005184 is the 21.9 MiB window the receiver offers; but unacked:2398 segments, about 3.5 MB, are actually in flight, and Send-Q holds 3471640 bytes. 3.5 MB per round trip of 0.10 to 0.127 s is roughly 27 to 35 MB/s, in line with the 246 Mbit/s (31 MB/s) iperf3 measured. Neither the network nor the receiver holds the sender back: its own send buffer does. (busy and rwnd_limited count time over the connection's whole life, start included, so the current amounts are better evidence.) For TCP, that buffer is sized automatically up to the third field of net.ipv4.tcp_wmem.
The sender may grow its buffer to 4 MiB; after the kernel's own overhead per packet, that leaves about 3.5 MB for data in flight, well short of 12.5 MB. The receiver's maximum, tcp_rmem, is already about 29 MiB. (These TCP settings are per network namespace, and a new namespace starts with the host's values, so changing them inside kt-a leaves the host alone.) The prediction: with a 16 MiB maximum the sender can keep the whole bandwidth-delay product in flight, and throughput should approach the link rate.
869 Mbit/s, close to what 1 Gbit/s carries after packet headers. unacked is now 6364 segments, about 9.2 MB, practically equal to cwnd (6383), and another 2.5 MB wait in the socket unsent (notsent): the sender's buffer is no longer the limit; congestion control on the network path is, which is where the limit should be. One setting, one prediction, confirmed.
net.core.wmem_max and net.core.rmem_max look like the obvious settings, and many guides raise them. The kernel's ip-sysctl documentation says they cap only buffers that an application requests explicitly with SO_SNDBUF or SO_RCVBUF; a socket that does so switches off automatic tuning. Automatically tuned TCP buffers are capped by the third field of tcp_wmem and tcp_rmem, which "does not override" the core values. The memory cost is also real: every connection that fills its buffer can use up to the maximum, and net.ipv4.tcp_mem is the global ceiling. Raise the maximum for hosts that move bulk data over long paths, not by habit.Check that cost before you keep the change. net.ipv4.tcp_mem holds three values in pages for all TCP sockets of the host together: below the first TCP does not limit its memory, above the second it enters memory pressure and holds back every socket's buffers, and the third is a hard limit. The kernel sizes them from the machine's memory at boot. The mem field of /proc/net/sockstat is the current use, also in pages.
All TCP sockets together held 3820 pages of 4 KiB, about 15.6 MB, nearly all of it this one upload, against a pressure threshold of 60037 pages (about 234 MiB) on this small VM. About fifteen such uploads at once would push every TCP socket on the host into memory pressure. On a real host, read them under the real load, change one host first and compare, and only then roll the file out to the fleet.
Making it persistent without side effects
A value written with sysctl -w lasts until reboot. At boot, systemd-sysctl reads .conf files from /usr/lib/sysctl.d (the distribution's) and /etc/sysctl.d (yours); all files are sorted by name across the directories and a later file wins, which is why yours should start with 60 to 90 as sysctl.d(5) recommends. The /proc and /sys lesson listed these directories and the RHEL difference, and the hardening course's network sysctl lesson (optional) covers the mechanics in full. The measured value goes in a file of its own, with the reason on comment lines (sysctl files do not accept a comment after a value):
# Uploads to the backup site (1 Gbit/s, 100 ms round trip) were limited by# the TCP send buffer: 3 MB in flight, less than cwnd and the receive window.net.ipv4.tcp_wmem = 4096 16384 16777216
Apply only that file with sysctl -p. Many guides say to run sysctl --system instead, which re-applies every file, and on Ubuntu that has a side effect worth seeing.
The file order is the precedence order, with yours after the vendor's (99-lima.conf belongs to the lab's VM tool, not to Ubuntu). But 10-coredump-debian.conf sets kernel.core_pattern = core, which apport replaces with its own pipe when apport.service starts after systemd-sysctl. sysctl --system puts the file's value back, and until apport is restarted, crashes are written as plain core files in the crashing process's directory. (On RHEL, 50-coredump.conf itself holds systemd-coredump's pipe, so re-applying files is harmless there.) Restart apport, then remove the lab's file and put the default back:
Memory and block settings that are not sysctls
Transparent huge pages (THP) let the kernel back anonymous memory with 2 MiB pages instead of 4 KiB ones, which cuts page-table work and TLB misses for large heaps. The cost is latency when the kernel must compact memory to find a free 2 MiB block, which is why some database vendors ask for THP to be limited. The I/O scheduler (see the block layer lesson) and read-ahead are per-device block settings. All of these live in /sys, not /proc/sys, so sysctl.d cannot set them.
Ubuntu uses THP only where a program asks for it with madvise(); RHEL uses it for all anonymous memory ([always]). Both reclaim or compact synchronously only for madvise regions (defrag). AnonHugePages shows how much memory huge pages back now; on this Ubuntu VM it is not zero because on AArch64 glibc 2.43's malloc asks for huge pages by default (its NEWS file says so). On x86_64 only programs that call madvise() themselves, or set the glibc.malloc.hugetlb tunable, get them. The read-ahead values differ (8192 against 128 KiB) because the two kernels size the default differently for the same kind of virtual disk, not because of a distribution rule; check read_ahead_kb on your own devices, since 8 MiB of read-ahead on a volume that serves random reads wastes I/O. To change such settings persistently, use the kernel command line (transparent_hugepage=, see the boot lesson), a udev rule for a block device, or, on a host that runs TuneD, a TuneD profile.
TuneD applies a named profile of sysctls, /sys settings, CPU governors and more, and can reverse it. It is part of a standard RHEL 10 installation: the Base package group lists it as mandatory, and the systemd preset enables it, so a RHEL server installed from the ISO runs a profile from its first boot (virtual-guest on most VMs, throughput-performance otherwise). Run tuned-adm active before you read any value on a RHEL host as a default. This Rocky cloud image is the exception: it does not include the package, although its preset would enable the service.
TuneD chose throughput-performance. Its documentation says it picks virtual-guest for virtual machines, but it detects them with virt-what, which prints nothing for this hypervisor (Apple's virtualization framework, under the lab's VM tool): a reminder to check what an automatic choice was based on. The profile:
One profile changed swappiness from 60 to 10, the dirty limit from 20% to 40% and read-ahead from 128 to 4096 KiB, and asked for performance CPU settings: several changes at once, none measured on this workload. (> in a value means "raise it to at least this", so somaxconn keeps its default of 4096.) tuned-adm verify reports a difference, and its log names it: the virtual CPUs have no frequency boost control, so boost=1 cannot be applied. That is what TuneD offers over a pasted list of settings: it knows what it set, can check it, and can undo it.
On a host where TuneD runs, make your own persistent changes as a custom profile, a tuned.conf in its own directory under /etc/tuned/profiles whose [main] section has include= naming the active profile, rather than as a udev rule or a bare /sys write: TuneD's [disk] and [vm] plugins write read-ahead and THP too, and with two writers of one file the value depends on which ran last. Sysctls are the exception, as TuneD's own configuration says:
So a file in /etc/sysctl.d still wins over the profile. tuned-adm off puts every value back:
Resource limits: the service's, not your shell's
The limits that most often need raising are not sysctls but per-process resource limits, and the usual one is RLIMIT_NOFILE, the number of open file descriptors (the lesson on files and descriptors explains the limits and the EMFILE error). A small program stands in for a proxy that needs 3000 descriptors:
# fds.py: opens 3000 descriptors, as a busy proxy would, and says how far it gotimport osimport resourcesoft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)fds = []try:for _ in range(3000):fds.append(os.open("/dev/null", os.O_RDONLY))print(f"limit {soft}/{hard}: opened {len(fds)}")except OSError as e:print(f"limit {soft}/{hard}: failed after {len(fds)}: {e}")raise SystemExit(1)
At the default soft limit of 1024 it fails after 1021 (descriptors 0 to 2 are already open). In the shell, ulimit -n 4096 raises the limit and the program succeeds; without -S it sets the hard limit too, so the shell cannot go back above 4096. Now run the same program as a service, from that same shell:
It fails again. (--wait --pipe connect a transient service to your terminal; --collect removes the unit even when it fails.) A service is started by systemd, not by your shell, so it gets systemd's defaults (soft 1024, hard 524288), and neither your ulimit nor /etc/security/limits.conf, which pam_limits applies only to login sessions, reaches it. The fix belongs to the unit: LimitNOFILE=, here as a property of a transient service; for an installed service, put LimitNOFILE=4096 under [Service] in a drop-in, run systemctl daemon-reload, restart it, and read /proc/PID/limits of the new process.
The systemd.exec manual also suggests the cleaner fix where you control the code: a program that does not use select() can raise its own soft limit up to the hard limit of 524288 at start-up, and needs no unit change at all.
Where to go from here
This course followed Linux from boot to the scheduler, memory, storage and network, and then the tools that measure each layer. Two courses build on it: Advanced container security builds a container by hand from the namespaces, cgroups and capabilities you used here, then hardens and attacks that boundary, and Runtime & eBPF security takes the eBPF tools into Falco, Tetragon and Cilium for runtime detection and enforcement. The hardening course, optional for this one, supplies the LSM and sysctl background they also use.
Try this
Keep the path from the measured case: the namespaces and the iperf3 server are still there, and the sender's maximum is still 16 MiB. Double the delay in both directions with sudo ip netns exec kt-a tc qdisc change dev kt-va root netem delay 100ms rate 1gbit limit 40000 and the matching command for kt-vb (without the rate). Before running iperf3 again, predict the result: the bandwidth-delay product is now 25 MB, more than the 16 MiB buffer, so throughput should fall to roughly half the link. Measure, then raise the maximum to 32 MiB, which covers the 25 MB, and predict again: most of the link. The lab run measured 440 then 764 Mbit/s. The VM's two CPUs also run the emulator, so repeated runs differ noticeably: compare runs made one after the other, and repeat before you trust a small difference. When you are done, stop the iperf3 server (ip netns pids lists the processes inside a namespace) and delete both namespaces:
Takeaway
Change a kernel setting only when a measurement shows that it is the limit and you can predict what the change will do; apply it with a file of its own and sysctl -p, record why, and check the effect with the same measurement.