Tuning the kernel, after you have measured

Measure, change one thing, verify.

Advanced14 min · lesson 21 of 21

Every earlier lesson in this course measured something. This one changes something, and only after a measurement has named the limit. You will work through one tuning case from symptom to verified fix, make the change persistent without breaking something else, read the memory and block settings that are not sysctls, raise a service's file-descriptor limit where it actually applies, and see what TuneD, RHEL's tuning daemon, does to a host. A setting copied from an old guide can make a current kernel slower, so the method matters more than any value.

The loop: measure, change one thing, verify

The tuning loop
1Measure the symptom
throughput, latency, errors, with numbers
2Find the limiting resource
which queue, buffer or limit is full
3Predict the effect
what should change, and by how much
4Change one setting
sysctl, TuneD profile or unit property
5Measure again
same workload, same tool
6Keep and record, or revert
a comment that says why
One change per pass, or you cannot tell which change did what.

The prediction step separates tuning from guessing: if you cannot say what a setting should do to the number you measured, you do not know that it is the limit. One change at a time keeps the cause visible, and the record lets the next engineer remove a setting that a kernel upgrade made unnecessary.

A measured case: an upload over a long path

The symptom: backups uploaded to a remote site over a 1 Gbit/s link run at about a quarter of the link's speed, with no errors and idle CPUs. To reproduce it safely, build the path between two network namespaces: a veth pair, and netem (the kernel's network emulator queueing discipline) adding 50 ms of delay in each direction and limiting the sender to 1 Gbit/s. Nothing outside the namespaces changes. iperf3 runs as a server in kt-b.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns add kt-a sudo ip netns add kt-b sudo ip link add kt-va netns kt-a type veth peer name kt-vb netns kt-b sudo ip -n kt-a addr add 198.51.100.1/24 dev kt-va sudo ip -n kt-b addr add 198.51.100.2/24 dev kt-vb sudo ip -n kt-a link set kt-va up sudo ip -n kt-b link set kt-vb up sudo ip netns exec kt-a tc qdisc add dev kt-va root netem delay 50ms rate 1gbit limit 20000 sudo ip netns exec kt-b tc qdisc add dev kt-vb root netem delay 50ms limit 20000 sudo ip netns exec kt-b iperf3 -s -D
$ sudo ip netns exec kt-a ping -c 3 -q 198.51.100.2
… rtt min/avg/max/mdev = 107.658/145.452/216.141/50.024 ms

The round trip takes at least 100 ms (the first packet is slower while the neighbour entry is resolved). Now measure the upload from kt-a. -O 2 leaves the first two seconds of slow start out of the average.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec kt-a iperf3 -c 198.51.100.2 -t 10 -O 2
Connecting to host 198.51.100.2, port 5201 [ 5] local 198.51.100.1 port 47658 connected to 198.51.100.2 port 5201 [ ID] Interval Transfer Bitrate Retr Cwnd … [ 5] 0.00-1.01 sec 26.5 MBytes 220 Mbits/sec 0 3.91 MBytes [ 5] 1.01-2.01 sec 36.4 MBytes 305 Mbits/sec 0 5.70 MBytes … - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-10.01 sec 294 MBytes 246 Mbits/sec 0 sender [ 5] 0.00-10.11 sec 297 MBytes 246 Mbits/sec receiver …

About 246 Mbit/s, steady, and no retransmissions (Retr 0), so nothing is being lost. TCP can have at most one window of unacknowledged data in flight per round trip, so throughput is roughly bytes in flight divided by the round-trip time. Filling this path needs its bandwidth-delay product in flight: 1 Gbit/s × 0.1 s = 12.5 MB. Look at the connection while it runs; ss -tin prints TCP's internal state for each socket.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec kt-a iperf3 -c 198.51.100.2 -t 12 > /dev/null & sleep 6 sudo ip netns exec kt-a ss -tin dst 198.51.100.2
State Recv-Q Send-Q Local Address:Port Peer Address:Port … ESTAB 0 3471640 198.51.100.1:36154 198.51.100.2:5201 cubic wscale:9,9 rto:327 rtt:126.693/4.264 mss:1448 pmtu:1500 rcvmss:536 advmss:1448 cwnd:4818 ssthresh:1603 bytes_sent:130938789 bytes_acked:127467150 segs_out:90449 segs_in:2723 data_segs_out:90447 send 440527196bps lastsnd:49 lastrcv:5628 lastack:43 pacing_rate 528631064bps delivery_rate 396053544bps delivered:88050 app_limited busy:5627ms rwnd_limited:1045ms(18.6%) unacked:2398 rcv_space:14480 rcv_ssthresh:64088 minrtt:100.671 snd_wnd:23005184 rcv_wnd:64512

The first socket is iperf3's control connection (left out). For the data connection, three numbers decide the limit. cwnd:4818 segments of 1448 bytes is about 7 MB that congestion control would allow; snd_wnd:23005184 is the 21.9 MiB window the receiver offers; but unacked:2398 segments, about 3.5 MB, are actually in flight, and Send-Q holds 3471640 bytes. 3.5 MB per round trip of 0.10 to 0.127 s is roughly 27 to 35 MB/s, in line with the 246 Mbit/s (31 MB/s) iperf3 measured. Neither the network nor the receiver holds the sender back: its own send buffer does. (busy and rwnd_limited count time over the connection's whole life, start included, so the current amounts are better evidence.) For TCP, that buffer is sized automatically up to the third field of net.ipv4.tcp_wmem.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec kt-a sysctl net.ipv4.tcp_wmem sudo ip netns exec kt-b sysctl net.ipv4.tcp_rmem
net.ipv4.tcp_wmem = 4096 16384 4194304 net.ipv4.tcp_rmem = 4096 131072 30739360

The sender may grow its buffer to 4 MiB; after the kernel's own overhead per packet, that leaves about 3.5 MB for data in flight, well short of 12.5 MB. The receiver's maximum, tcp_rmem, is already about 29 MiB. (These TCP settings are per network namespace, and a new namespace starts with the host's values, so changing them inside kt-a leaves the host alone.) The prediction: with a 16 MiB maximum the sender can keep the whole bandwidth-delay product in flight, and throughput should approach the link rate.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec kt-a sysctl -w net.ipv4.tcp_wmem="4096 16384 16777216"
net.ipv4.tcp_wmem = 4096 16384 16777216
$ sudo ip netns exec kt-a iperf3 -c 198.51.100.2 -t 10 -O 2
… [ ID] Interval Transfer Bitrate Retr Cwnd … [ 5] 0.00-1.00 sec 97.1 MBytes 815 Mbits/sec 0 16.3 MBytes [ 5] 1.00-2.00 sec 108 MBytes 901 Mbits/sec 0 16.3 MBytes … - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-10.00 sec 1.01 GBytes 869 Mbits/sec 0 sender [ 5] 0.00-10.10 sec 1.02 GBytes 869 Mbits/sec receiver …
$ sudo ip netns exec kt-a iperf3 -c 198.51.100.2 -t 12 > /dev/null & sleep 6 sudo ip netns exec kt-a ss -tin dst 198.51.100.2
State Recv-Q Send-Q Local Address:Port Peer Address:Port … ESTAB 0 11727560 198.51.100.1:35572 198.51.100.2:5201 cubic wscale:9,9 rto:302 rtt:101.302/0.277 mss:1448 pmtu:1500 rcvmss:536 advmss:1448 cwnd:6383 ssthresh:954 bytes_sent:168190821 bytes_acked:158975750 segs_out:116160 segs_in:3471 data_segs_out:116158 send 729903378bps lastrcv:5646 pacing_rate 875880808bps delivery_rate 693563744bps delivered:109795 busy:5644ms rwnd_limited:872ms(15.5%) unacked:6364 rcv_space:14480 rcv_ssthresh:64088 notsent:2512488 minrtt:100.675 snd_wnd:19276288 rcv_wnd:64512

869 Mbit/s, close to what 1 Gbit/s carries after packet headers. unacked is now 6364 segments, about 9.2 MB, practically equal to cwnd (6383), and another 2.5 MB wait in the socket unsent (notsent): the sender's buffer is no longer the limit; congestion control on the network path is, which is where the limit should be. One setting, one prediction, confirmed.

The wrong knob for this case
net.core.wmem_max and net.core.rmem_max look like the obvious settings, and many guides raise them. The kernel's ip-sysctl documentation says they cap only buffers that an application requests explicitly with SO_SNDBUF or SO_RCVBUF; a socket that does so switches off automatic tuning. Automatically tuned TCP buffers are capped by the third field of tcp_wmem and tcp_rmem, which "does not override" the core values. The memory cost is also real: every connection that fills its buffer can use up to the maximum, and net.ipv4.tcp_mem is the global ceiling. Raise the maximum for hosts that move bulk data over long paths, not by habit.

Check that cost before you keep the change. net.ipv4.tcp_mem holds three values in pages for all TCP sockets of the host together: below the first TCP does not limit its memory, above the second it enters memory pressure and holds back every socket's buffers, and the third is a hard limit. The kernel sizes them from the machine's memory at boot. The mem field of /proc/net/sockstat is the current use, also in pages.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec kt-a iperf3 -c 198.51.100.2 -t 10 > /dev/null & sleep 6 grep TCP: /proc/net/sockstat sysctl net.ipv4.tcp_mem getconf PAGESIZE wait
TCP: inuse 3 orphan 0 tw 0 alloc 9 mem 3820 net.ipv4.tcp_mem = 45027 60037 90054 4096

All TCP sockets together held 3820 pages of 4 KiB, about 15.6 MB, nearly all of it this one upload, against a pressure threshold of 60037 pages (about 234 MiB) on this small VM. About fifteen such uploads at once would push every TCP socket on the host into memory pressure. On a real host, read them under the real load, change one host first and compare, and only then roll the file out to the fleet.

Making it persistent without side effects

A value written with sysctl -w lasts until reboot. At boot, systemd-sysctl reads .conf files from /usr/lib/sysctl.d (the distribution's) and /etc/sysctl.d (yours); all files are sorted by name across the directories and a later file wins, which is why yours should start with 60 to 90 as sysctl.d(5) recommends. The /proc and /sys lesson listed these directories and the RHEL difference, and the hardening course's network sysctl lesson (optional) covers the mechanics in full. The measured value goes in a file of its own, with the reason on comment lines (sysctl files do not accept a comment after a value):

/etc/sysctl.d/60-wan-sndbuf.conf
# Uploads to the backup site (1 Gbit/s, 100 ms round trip) were limited by
# the TCP send buffer: 3 MB in flight, less than cwnd and the receive window.
net.ipv4.tcp_wmem = 4096 16384 16777216

Apply only that file with sysctl -p. Many guides say to run sysctl --system instead, which re-applies every file, and on Ubuntu that has a side effect worth seeing.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo sysctl -p /etc/sysctl.d/60-wan-sndbuf.conf
net.ipv4.tcp_wmem = 4096 16384 16777216
$ cat /proc/sys/kernel/core_pattern
|/usr/share/apport/apport -p%p -s%s -c%c -d%d -P%P -u%u -g%g -F%F -- %E
$ sudo sysctl --system | grep -E "^\* Applying|core_pattern|tcp_wmem"
* Applying /usr/lib/sysctl.d/10-apparmor.conf ... * Applying /usr/lib/sysctl.d/10-coredump-debian.conf ... * Applying /usr/lib/sysctl.d/50-default.conf ... * Applying /usr/lib/sysctl.d/50-pid-max.conf ... * Applying /usr/lib/sysctl.d/55-bufferbloat.conf ... * Applying /usr/lib/sysctl.d/55-console-messages.conf ... * Applying /usr/lib/sysctl.d/55-ipv6-privacy.conf ... * Applying /usr/lib/sysctl.d/55-kernel-hardening.conf ... * Applying /usr/lib/sysctl.d/55-magic-sysrq.conf ... * Applying /usr/lib/sysctl.d/55-map-count.conf ... * Applying /usr/lib/sysctl.d/55-network-security.conf ... * Applying /usr/lib/sysctl.d/55-ptrace.conf ... * Applying /usr/lib/sysctl.d/55-zeropage.conf ... * Applying /etc/sysctl.d/60-wan-sndbuf.conf ... * Applying /etc/sysctl.d/99-cloudimg-ipv6.conf ... * Applying /etc/sysctl.d/99-lima.conf ... kernel.core_pattern = core net.ipv4.tcp_wmem = 4096 16384 16777216
$ cat /proc/sys/kernel/core_pattern
core

The file order is the precedence order, with yours after the vendor's (99-lima.conf belongs to the lab's VM tool, not to Ubuntu). But 10-coredump-debian.conf sets kernel.core_pattern = core, which apport replaces with its own pipe when apport.service starts after systemd-sysctl. sysctl --system puts the file's value back, and until apport is restarted, crashes are written as plain core files in the crashing process's directory. (On RHEL, 50-coredump.conf itself holds systemd-coredump's pipe, so re-applying files is harmless there.) Restart apport, then remove the lab's file and put the default back:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl restart apport cat /proc/sys/kernel/core_pattern
|/usr/share/apport/apport -p%p -s%s -c%c -d%d -P%P -u%u -g%g -F%F -- %E
$ sudo rm /etc/sysctl.d/60-wan-sndbuf.conf sudo sysctl -w net.ipv4.tcp_wmem="4096 16384 4194304"
net.ipv4.tcp_wmem = 4096 16384 4194304

Memory and block settings that are not sysctls

Transparent huge pages (THP) let the kernel back anonymous memory with 2 MiB pages instead of 4 KiB ones, which cuts page-table work and TLB misses for large heaps. The cost is latency when the kernel must compact memory to find a free 2 MiB block, which is why some database vendors ask for THP to be limited. The I/O scheduler (see the block layer lesson) and read-ahead are per-device block settings. All of these live in /sys, not /proc/sys, so sysctl.d cannot set them.

deploy@web01 · Ubuntu 26.04 LTS
$ grep . /sys/kernel/mm/transparent_hugepage/enabled /sys/kernel/mm/transparent_hugepage/defrag grep AnonHugePages /proc/meminfo grep . /sys/block/vda/queue/scheduler /sys/block/vda/queue/read_ahead_kb
/sys/kernel/mm/transparent_hugepage/enabled:always [madvise] never /sys/kernel/mm/transparent_hugepage/defrag:always defer defer+madvise [madvise] never AnonHugePages: 81920 kB /sys/block/vda/queue/scheduler:none [mq-deadline] /sys/block/vda/queue/read_ahead_kb:8192
deploy@rocky10 · Rocky Linux 10.2
$ grep . /sys/kernel/mm/transparent_hugepage/enabled /sys/kernel/mm/transparent_hugepage/defrag grep . /sys/block/vda/queue/scheduler /sys/block/vda/queue/read_ahead_kb
/sys/kernel/mm/transparent_hugepage/enabled:[always] madvise never /sys/kernel/mm/transparent_hugepage/defrag:always defer defer+madvise [madvise] never /sys/block/vda/queue/scheduler:none [mq-deadline] kyber bfq /sys/block/vda/queue/read_ahead_kb:128

Ubuntu uses THP only where a program asks for it with madvise(); RHEL uses it for all anonymous memory ([always]). Both reclaim or compact synchronously only for madvise regions (defrag). AnonHugePages shows how much memory huge pages back now; on this Ubuntu VM it is not zero because on AArch64 glibc 2.43's malloc asks for huge pages by default (its NEWS file says so). On x86_64 only programs that call madvise() themselves, or set the glibc.malloc.hugetlb tunable, get them. The read-ahead values differ (8192 against 128 KiB) because the two kernels size the default differently for the same kind of virtual disk, not because of a distribution rule; check read_ahead_kb on your own devices, since 8 MiB of read-ahead on a volume that serves random reads wastes I/O. To change such settings persistently, use the kernel command line (transparent_hugepage=, see the boot lesson), a udev rule for a block device, or, on a host that runs TuneD, a TuneD profile.

TuneD applies a named profile of sysctls, /sys settings, CPU governors and more, and can reverse it. It is part of a standard RHEL 10 installation: the Base package group lists it as mandatory, and the systemd preset enables it, so a RHEL server installed from the ISO runs a profile from its first boot (virtual-guest on most VMs, throughput-performance otherwise). Run tuned-adm active before you read any value on a RHEL host as a default. This Rocky cloud image is the exception: it does not include the package, although its preset would enable the service.

deploy@rocky10 · Rocky Linux 10.2
$ rpm -q tuned
package tuned is not installed
$ grep -x "enable tuned.service" /usr/lib/systemd/system-preset/90-default.preset
enable tuned.service
$ sysctl vm.swappiness vm.dirty_ratio vm.dirty_background_ratio
vm.swappiness = 60 vm.dirty_ratio = 20 vm.dirty_background_ratio = 10
$ sudo dnf install -y tuned
… Installing: tuned noarch 2.27.0-2.el10_2 baseos 458 k Installing dependencies: hdparm aarch64 9.65-6.el10 baseos 93 k python3-inotify noarch 0.9.6-36.el10 baseos 62 k python3-linux-procfs noarch 0.7.4-1.el10 baseos 40 k python3-pyudev noarch 0.24.1-10.el10 baseos 101 k virt-what aarch64 1.27-3.el10 baseos 39 k which aarch64 2.21-44.el10_0 baseos 41 k Installing weak dependencies: python3-perf aarch64 6.12.0-211.60.1.el10_2 appstream 2.8 M … Complete!
$ sudo systemctl enable --now tuned
$ tuned-adm recommend tuned-adm active
throughput-performance Current active profile: throughput-performance
$ sudo virt-what; echo "virt-what exit $?"
virt-what exit 0

TuneD chose throughput-performance. Its documentation says it picks virtual-guest for virtual machines, but it detects them with virt-what, which prints nothing for this hypervisor (Apple's virtualization framework, under the lab's VM tool): a reminder to check what an automatic choice was based on. The profile:

deploy@rocky10 · Rocky Linux 10.2
$ grep -v "^#" /usr/lib/tuned/profiles/throughput-performance/tuned.conf | grep .
[main] summary=Broadly applicable tuning that provides excellent performance across a variety of common server workloads … [cpu] governor=performance … boost=1 … [vm] dirty_bytes = 40% dirty_background_bytes = 10% … [disk] readahead=>4096 [sysctl] vm.swappiness=10 net.core.somaxconn=>2048 …
$ sysctl vm.swappiness vm.dirty_ratio vm.dirty_background_ratio grep . /sys/kernel/mm/transparent_hugepage/enabled /sys/block/vda/queue/read_ahead_kb
vm.swappiness = 10 vm.dirty_ratio = 40 vm.dirty_background_ratio = 10 /sys/kernel/mm/transparent_hugepage/enabled:[always] madvise never /sys/block/vda/queue/read_ahead_kb:4096
$ sudo tuned-adm verify
Verification failed, current system settings differ from the preset profile. You can mostly fix this by restarting the TuneD daemon, e.g.: systemctl restart tuned or service tuned restart Sometimes (if some plugins like bootloader are used) a reboot may be required. See TuneD log file ('/var/log/tuned/tuned.log') for details.
$ sudo grep -E "ERROR|WARNING" /var/log/tuned/tuned.log | tail -n 6
… 2026-09-27 09:26:16,981 ERROR tuned.plugins.base: verify: failed: device cpu0: 'boost' = 'None', expected '1' 2026-09-27 09:26:16,981 ERROR tuned.plugins.base: verify: failed: device cpu1: 'boost' = 'None', expected '1'

One profile changed swappiness from 60 to 10, the dirty limit from 20% to 40% and read-ahead from 128 to 4096 KiB, and asked for performance CPU settings: several changes at once, none measured on this workload. (> in a value means "raise it to at least this", so somaxconn keeps its default of 4096.) tuned-adm verify reports a difference, and its log names it: the virtual CPUs have no frequency boost control, so boost=1 cannot be applied. That is what TuneD offers over a pasted list of settings: it knows what it set, can check it, and can undo it.

On a host where TuneD runs, make your own persistent changes as a custom profile, a tuned.conf in its own directory under /etc/tuned/profiles whose [main] section has include= naming the active profile, rather than as a udev rule or a bare /sys write: TuneD's [disk] and [vm] plugins write read-ahead and THP too, and with two writers of one file the value depends on which ran last. Sysctls are the exception, as TuneD's own configuration says:

deploy@rocky10 · Rocky Linux 10.2
$ grep -B3 "^reapply_sysctl" /etc/tuned/tuned-main.conf
# /etc/sysctl.conf. If enabled, these sysctls will be reappliead # after TuneD sysctls are applied, i.e. TuneD sysctls will not # override user-provided system sysctls. reapply_sysctl = 1

So a file in /etc/sysctl.d still wins over the profile. tuned-adm off puts every value back:

deploy@rocky10 · Rocky Linux 10.2
$ sudo tuned-adm off tuned-adm active sysctl vm.swappiness vm.dirty_ratio vm.dirty_background_ratio cat /sys/block/vda/queue/read_ahead_kb
No current active profile. vm.swappiness = 60 vm.dirty_ratio = 20 vm.dirty_background_ratio = 10 128
$ sudo systemctl disable --now tuned sudo dnf remove -y tuned
Removed '/etc/systemd/system/multi-user.target.wants/tuned.service'. … Complete!

Resource limits: the service's, not your shell's

The limits that most often need raising are not sysctls but per-process resource limits, and the usual one is RLIMIT_NOFILE, the number of open file descriptors (the lesson on files and descriptors explains the limits and the EMFILE error). A small program stands in for a proxy that needs 3000 descriptors:

~/kt/fds.py
# fds.py: opens 3000 descriptors, as a busy proxy would, and says how far it got
import os
import resource
soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)
fds = []
try:
for _ in range(3000):
fds.append(os.open("/dev/null", os.O_RDONLY))
print(f"limit {soft}/{hard}: opened {len(fds)}")
except OSError as e:
print(f"limit {soft}/{hard}: failed after {len(fds)}: {e}")
raise SystemExit(1)
deploy@web01 · Ubuntu 26.04 LTS
$ ulimit -n ulimit -Hn python3 ~/kt/fds.py; ulimit -n 4096; python3 ~/kt/fds.py
1024 524288 limit 1024/524288: failed after 1021: [Errno 24] Too many open files: '/dev/null' limit 4096/4096: opened 3000

At the default soft limit of 1024 it fails after 1021 (descriptors 0 to 2 are already open). In the shell, ulimit -n 4096 raises the limit and the program succeeds; without -S it sets the hard limit too, so the shell cannot go back above 4096. Now run the same program as a service, from that same shell:

deploy@web01 · Ubuntu 26.04 LTS
$ ulimit -n 4096 sudo systemd-run --quiet --wait --pipe --collect --uid=$USER python3 ~/kt/fds.py
limit 1024/524288: failed after 1021: [Errno 24] Too many open files: '/dev/null'
$ systemctl show -p DefaultLimitNOFILE -p DefaultLimitNOFILESoft
DefaultLimitNOFILE=524288 DefaultLimitNOFILESoft=1024

It fails again. (--wait --pipe connect a transient service to your terminal; --collect removes the unit even when it fails.) A service is started by systemd, not by your shell, so it gets systemd's defaults (soft 1024, hard 524288), and neither your ulimit nor /etc/security/limits.conf, which pam_limits applies only to login sessions, reaches it. The fix belongs to the unit: LimitNOFILE=, here as a property of a transient service; for an installed service, put LimitNOFILE=4096 under [Service] in a drop-in, run systemctl daemon-reload, restart it, and read /proc/PID/limits of the new process.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --quiet --wait --pipe --collect --uid=$USER -p LimitNOFILE=4096 python3 ~/kt/fds.py
limit 4096/4096: opened 3000

The systemd.exec manual also suggests the cleaner fix where you control the code: a program that does not use select() can raise its own soft limit up to the hard limit of 524288 at start-up, and needs no unit change at all.

Where to go from here

This course followed Linux from boot to the scheduler, memory, storage and network, and then the tools that measure each layer. Two courses build on it: Advanced container security builds a container by hand from the namespaces, cgroups and capabilities you used here, then hardens and attacks that boundary, and Runtime & eBPF security takes the eBPF tools into Falco, Tetragon and Cilium for runtime detection and enforcement. The hardening course, optional for this one, supplies the LSM and sysctl background they also use.

Try this

Keep the path from the measured case: the namespaces and the iperf3 server are still there, and the sender's maximum is still 16 MiB. Double the delay in both directions with sudo ip netns exec kt-a tc qdisc change dev kt-va root netem delay 100ms rate 1gbit limit 40000 and the matching command for kt-vb (without the rate). Before running iperf3 again, predict the result: the bandwidth-delay product is now 25 MB, more than the 16 MiB buffer, so throughput should fall to roughly half the link. Measure, then raise the maximum to 32 MiB, which covers the 25 MB, and predict again: most of the link. The lab run measured 440 then 764 Mbit/s. The VM's two CPUs also run the emulator, so repeated runs differ noticeably: compare runs made one after the other, and repeat before you trust a small difference. When you are done, stop the iperf3 server (ip netns pids lists the processes inside a namespace) and delete both namespaces:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns pids kt-b | xargs sudo kill sudo ip netns del kt-a sudo ip netns del kt-b

Takeaway

Change a kernel setting only when a measurement shows that it is the limit and you can predict what the change will do; apply it with a file of its own and sysctl -p, record why, and check the effect with the same measurement.

Quick check
01A service on Ubuntu fails with "Too many open files" at about 1000 connections. A colleague adds "* soft nofile 65536" to /etc/security/limits.conf, logs in again, sees ulimit -n print 65536, and restarts the service. It still fails at the same point. Why?
Incorrect — pam_limits applies the file at each login, which is why the new login showed 65536; a reboot changes nothing for a service.
Correct — systemd starts services with its own defaults, soft 1024, and LimitNOFILE= in the unit is where the service's limit is set.
Incorrect — fs.file-max is a system-wide count, which systemd raises to its maximum at boot; it does not cap a per-process limit.
Incorrect — ulimit -n without -H prints the soft limit; the shell's limits simply never reach a process that systemd starts.
02Uploads over a 100 ms, 1 Gbit/s path run at 240 Mbit/s with no retransmissions. ss -tin on the sender shows cwnd about 8 MB, snd_wnd 22 MB and about 3 MB unacknowledged. Which change do you expect to help?
Incorrect — That value caps only buffers requested with SO_SNDBUF; automatically tuned TCP buffers are capped by tcp_wmem.
Incorrect — The receiver already offers 22 MB, far more than the 3 MB in flight, so its window is not the limit here.
Correct — In flight is below both cwnd and the receiver's window, so the sender's autotuned buffer is the limit.
Incorrect — cwnd is already larger than what is in flight, and there is no loss; congestion control is not holding the sender back.
03After you run sudo sysctl --system on an Ubuntu 26.04 server to load a new file, a crashing service no longer produces anything in /var/lib/apport/coredump or /var/crash. What happened?
Incorrect — Ubuntu's files do not set suid_dumpable, and apport's value for it was left alone; the dump went elsewhere.
Incorrect — sysctl does not touch apport's own settings; apport never saw this crash at all.
Incorrect — ptrace_scope controls debugger attachment and has no effect on whether a crash is dumped.
Correct — apport sets its pipe at run time; re-applying all files resets it until apport restarts. Apply one file with sysctl -p.

Related