Kernel sysctls: ptrace, pointers, BPF and perf

Measure defaults, change what differs.

Intermediate14 min · lesson 19 of 24

A handful of kernel sysctls decide how much one foothold on a server is worth: whether a compromised process can read the memory of its neighbours, whether it can learn where the kernel sits in memory, and whether it can reach the kernel's most complex interfaces (BPF and perf) without privilege. Ubuntu 26.04 and RHEL 10 already set most of these well, but not identically, and copying an old hardening file can loosen them. In this lesson you measure each value and where it comes from, see what each control refuses and what it breaks for your own tools, recognise the settings that cannot be undone without a reboot, and write a short file that pins the good defaults and changes only what differs.

Measure first: values and where they come from

The network sysctls lesson covered how /proc/sys, sysctl -w and the sysctl.d directories work; the same rules apply here. Start by reading the running values on Ubuntu.

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl kernel.yama.ptrace_scope kernel.kptr_restrict kernel.dmesg_restrict kernel.unprivileged_bpf_disabled kernel.perf_event_paranoid kernel.randomize_va_space fs.protected_symlinks fs.protected_hardlinks fs.protected_fifos fs.protected_regular fs.suid_dumpable kernel.apparmor_restrict_unprivileged_userns
kernel.yama.ptrace_scope = 1 kernel.kptr_restrict = 1 kernel.dmesg_restrict = 1 kernel.unprivileged_bpf_disabled = 2 kernel.perf_event_paranoid = 4 kernel.randomize_va_space = 2 fs.protected_symlinks = 1 fs.protected_hardlinks = 1 fs.protected_fifos = 1 fs.protected_regular = 2 fs.suid_dumpable = 2 kernel.apparmor_restrict_unprivileged_userns = 1

Then find which of them a file sets. Only ptrace_scope (55-ptrace.conf), kptr_restrict (55-kernel-hardening.conf) and the fs.protected_* keys (systemd's 50-default.conf) come from files. The 99-lima.conf line belongs to the lab VM's tooling. The rest are compiled into Ubuntu's kernel as build options, so no file on disk explains them:

deploy@web01 · Ubuntu 26.04 LTS
$ grep -rE '^[^#]*(ptrace|kptr|dmesg|bpf|perf_event|randomize|protected_|suid_dumpable|userns)' /usr/lib/sysctl.d/ /etc/sysctl.d/
/usr/lib/sysctl.d/55-ptrace.conf:kernel.yama.ptrace_scope = 1 /usr/lib/sysctl.d/55-kernel-hardening.conf:kernel.kptr_restrict = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_hardlinks = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_symlinks = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_regular = 2 /usr/lib/sysctl.d/50-default.conf:fs.protected_fifos = 1 /usr/lib/sysctl.d/10-apparmor.conf:kernel.apparmor_restrict_unprivileged_userns = 1 /etc/sysctl.d/99-lima.conf:kernel.unprivileged_userns_clone=1
$ grep -E 'CONFIG_(SECURITY_DMESG_RESTRICT|BPF_UNPRIV_DEFAULT_OFF|SECURITY_PERF_EVENTS_RESTRICT)=' /boot/config-$(uname -r)
CONFIG_BPF_UNPRIV_DEFAULT_OFF=y CONFIG_SECURITY_DMESG_RESTRICT=y CONFIG_SECURITY_PERF_EVENTS_RESTRICT=y

CONFIG_SECURITY_DMESG_RESTRICT makes dmesg_restrict start at 1, CONFIG_BPF_UNPRIV_DEFAULT_OFF starts unprivileged_bpf_disabled at 2, and CONFIG_SECURITY_PERF_EVENTS_RESTRICT is a Debian and Ubuntu kernel option behind perf_event_paranoid = 4, a value above the upstream kernel's documented maximum of 2. RHEL 10 differs in three places: ptrace_scope is 0 (set by 10-default-yama-scope.conf), perf_event_paranoid is 2 (the kernel has no Ubuntu restrict option), and fs.protected_regular is 1 instead of 2.

deploy@rocky10 · Rocky Linux 10.2
$ sysctl kernel.yama.ptrace_scope kernel.kptr_restrict kernel.dmesg_restrict kernel.unprivileged_bpf_disabled kernel.perf_event_paranoid kernel.randomize_va_space fs.protected_symlinks fs.protected_hardlinks fs.protected_fifos fs.protected_regular fs.suid_dumpable
kernel.yama.ptrace_scope = 0 kernel.kptr_restrict = 1 kernel.dmesg_restrict = 1 kernel.unprivileged_bpf_disabled = 2 kernel.perf_event_paranoid = 2 kernel.randomize_va_space = 2 fs.protected_symlinks = 1 fs.protected_hardlinks = 1 fs.protected_fifos = 1 fs.protected_regular = 1 fs.suid_dumpable = 2
$ grep -rE '^[^#]*(ptrace|kptr|dmesg|bpf|perf_event|randomize|protected_|suid_dumpable)' /usr/lib/sysctl.d/ /etc/sysctl.d/
/usr/lib/sysctl.d/10-default-yama-scope.conf:kernel.yama.ptrace_scope = 0 /usr/lib/sysctl.d/50-coredump.conf:fs.suid_dumpable=2 /usr/lib/sysctl.d/50-default.conf:fs.protected_hardlinks = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_symlinks = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_regular = 1 /usr/lib/sysctl.d/50-default.conf:fs.protected_fifos = 1 /usr/lib/sysctl.d/50-redhat.conf:kernel.kptr_restrict = 1
$ grep -E 'CONFIG_(SECURITY_DMESG_RESTRICT|BPF_UNPRIV_DEFAULT_OFF|SECURITY_PERF_EVENTS_RESTRICT)=' /boot/config-$(uname -r)
CONFIG_BPF_UNPRIV_DEFAULT_OFF=y CONFIG_SECURITY_DMESG_RESTRICT=y

So the policy is short. On Ubuntu almost everything is already where CIS wants it; on RHEL, ptrace_scope is the notable difference. fs.suid_dumpable is 2 on both and CIS asks for 0. Everything else is a default worth pinning, and one value, perf_event_paranoid, is one you must not copy between distributions.

ptrace_scope: who may inspect another process

ptrace is the system call debuggers and tracers use to attach to a running process, read and write its memory and registers, and stop it. The classic Unix rule lets any process attach to any other process of the same user. That is the threat: a service usually runs several workers under one account, so an attacker who takes over one worker can read session keys and decrypted secrets out of the others. Yama, a small security module in the kernel, tightens the rule through kernel.yama.ptrace_scope: 0 is the classic rule, 1 allows attaching only to your own descendants (processes you started, and theirs), 2 allows attaching only with the CAP_SYS_PTRACE capability, and 3 turns attaching off entirely.

On Ubuntu (scope 1), tracing a program you start still works, because it is your child:

deploy@web01 · Ubuntu 26.04 LTS
$ strace -e trace=execve true
execve("/usr/bin/true", ["true"], 0xffffdcfef500 /* 12 vars */) = 0 +++ exited with 0 +++

Attaching to a sibling does not. Here the shell starts sleep in the background and then runs strace -p on it; strace and sleep are both children of the shell, so neither is the other's ancestor. That is exactly the position of one compromised worker next to another.

deploy@web01 · Ubuntu 26.04 LTS
$ sleep 600 & pid=$! timeout 2 strace -p $pid kill $pid
strace: attach: ptrace(PTRACE_SEIZE, 374597): Operation not permitted
$ sleep 600 & pid=$! gdb -q -batch -p $pid kill $pid
Could not attach to process. If your uid matches the uid of the target process, check the setting of /proc/sys/kernel/yama/ptrace_scope, or try again as the root user. For more details, see /etc/sysctl.d/10-ptrace.conf ptrace: Inappropriate ioctl for device.

gdb (not on a default Ubuntu Server: sudo apt install gdb; strace is) fails the same way. Its hint points at /etc/sysctl.d/10-ptrace.conf, a file name from older Ubuntu releases; on 26.04 the setting is in /usr/lib/sysctl.d/55-ptrace.conf. The "Inappropriate ioctl for device" line is gdb's garbled report of the same refusal. An administrator who needs to attach uses sudo, which carries CAP_SYS_PTRACE:

deploy@web01 · Ubuntu 26.04 LTS
$ sleep 600 & pid=$! sudo timeout 2 strace -p $pid kill $pid
strace: Process 374652 attached restart_syscall(<... resuming interrupted clock_nanosleep ...>strace: Process 374652 detached <detached ...>

On RHEL 10 (scope 0) the same sibling attach succeeds. Attaching to another user's process, here chronyd running as chrony, is refused at every scope; that is the ordinary same-user rule, not Yama. Setting scope 1 on RHEL closes the same-user gap, and because 1 is not a latch, writing 0 rolls it back:

deploy@rocky10 · Rocky Linux 10.2
$ sleep 600 & pid=$! timeout 2 strace -p $pid kill $pid
strace: Process 41535 attached restart_syscall(<... resuming interrupted clock_nanosleep ...>strace: Process 41535 detached <detached ...>
$ ps -o user=,pid=,comm= -C chronyd strace -p $(pgrep -x chronyd)
chrony 757 chronyd strace: attach: ptrace(PTRACE_SEIZE, 757): Operation not permitted
$ sudo sysctl -w kernel.yama.ptrace_scope=1 sleep 600 & pid=$! timeout 2 strace -p $pid kill $pid
kernel.yama.ptrace_scope = 1 strace: attach: ptrace(PTRACE_SEIZE, 41611): Operation not permitted
$ sudo sysctl -w kernel.yama.ptrace_scope=0
kernel.yama.ptrace_scope = 0

ptrace_scope = 1 is Ubuntu's default and a CIS recommendation (the CIS RHEL 10 Level 1 profile in scap-security-guide checks it), so on RHEL it is a change you make. The impact lands on people who attach to running processes: gdb -p, strace -p, crash reporters, and some APM (application performance monitoring) and profiling agents, which attach to application processes to sample what they are doing. Give them sudo or CAP_SYS_PTRACE, or have the application name its debugger with prctl(PR_SET_PTRACER), instead of lowering the value host-wide. Scope 3 is a one-way latch: once written, the kernel refuses any further change until reboot, so it suits appliances where nobody will ever debug in production.

Kernel addresses and the kernel log

Many kernel exploits need to know where kernel code and data sit in memory, and the kernel usually randomises that at boot, so they have to find out. /proc/kallsyms lists every kernel symbol with its address; kernel.kptr_restrict decides who sees the real addresses. At 1, a process without CAP_SYSLOG sees zeros and root sees the addresses; at 2, everyone sees zeros.

deploy@web01 · Ubuntu 26.04 LTS
$ grep -w do_nanosleep /proc/kallsyms sudo grep -w do_nanosleep /proc/kallsyms
0000000000000000 t do_nanosleep ffffb7b7aaeb9328 t do_nanosleep

Level 2 sounds stricter and therefore better, but look at what it costs. Kernel tracing tools use the same symbol table to turn addresses into function names. With 1, a bpftrace kernel stack is readable; with 2, root gets raw hex. (The frames are aarch64 functions from the lab CPU; an x86_64 server shows __x64_sys_clock_nanosleep and its own call chain.)

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -e 'kprobe:do_nanosleep { printf("%s", kstack(4)); exit(); }' -c 'sleep 0.2'
Attached 1 probe do_nanosleep+0 common_nsleep+84 __arm64_sys_clock_nanosleep+240 invoke_syscall.constprop.0+96
$ sudo sysctl -w kernel.kptr_restrict=2 sudo grep -w do_nanosleep /proc/kallsyms
kernel.kptr_restrict = 2 0000000000000000 t do_nanosleep
$ sudo bpftrace -e 'kprobe:do_nanosleep { printf("%s", kstack(4)); exit(); }' -c 'sleep 0.2'
Attached 1 probe 0xffffb7b7aaeb9328 0xffffb7b7a965563c 0xffffb7b7a9657678 0xffffb7b7a9419cc8
$ sudo sysctl -w kernel.kptr_restrict=1
kernel.kptr_restrict = 1

Keep 1, the Ubuntu and RHEL default and the value the CIS profile uses, on any host where someone may profile the kernel with perf, bpftrace or the bcc tools (a collection of ready-made BPF tracing programs). Use 2 only where nobody will, and treat it as costing you those tools.

The kernel ring buffer (what dmesg reads) carries driver messages, addresses in warnings and stack traces, and firewall log lines. dmesg_restrict = 1 requires CAP_SYSLOG to read it:

deploy@web01 · Ubuntu 26.04 LTS
$ dmesg | tail -n 1
dmesg: read kernel buffer failed: Operation not permitted

That is not the whole story on a systemd host. journald copies kernel messages into the journal, and members of the adm group (which Ubuntu gives the first administrator) can read them with journalctl -k. The protection is only as good as your group membership: review who is in adm and systemd-journal, and read the logging lesson for the file permissions.

deploy@web01 · Ubuntu 26.04 LTS
$ id -nG journalctl -k -b -o cat | head -n 1
deploy adm sudo Booting Linux on physical CPU 0x0000000000 [0x610f0000]

BPF, perf and the one-way latches

BPF lets programs load small verified programs into the kernel for tracing and networking, and perf gives access to CPU and kernel performance counters. Both are large, complex interfaces with a history of local privilege-escalation bugs, so both are closed to ordinary users by default on these distributions. kernel.unprivileged_bpf_disabled is 2 on both platforms, from the kernel build: unprivileged bpf() calls are refused, and root can change the value. bpftool reports it the same way:

deploy@web01 · Ubuntu 26.04 LTS
$ bpftool feature probe kernel unprivileged | grep 'bpf() syscall'
bpf() syscall restricted to privileged users (admin can change) bpf() syscall is available

Value 1 refuses the same calls and cannot be changed back until reboot. Old guides recommend writing 1; on a current kernel that adds no protection over 2 and only removes your ability to undo it. Pin 2.

perf_event_paranoid is where copying between distributions goes wrong. Ubuntu's 4 refuses unprivileged perf entirely, while root still works:

deploy@web01 · Ubuntu 26.04 LTS
$ perf stat -e task-clock true
Error: No supported events found. … perf_event_paranoid setting is 4: … >= 2: Disallow kernel profiling To make the adjusted perf_event_paranoid setting permanent preserve it in /etc/sysctl.conf (e.g. kernel.perf_event_paranoid = <setting>)
$ sudo perf stat -e task-clock true
Performance counter stats for 'true': 0.19 msec task-clock 0.001436761 seconds time elapsed 0.000000000 seconds user 0.000732000 seconds sys

perf's own advice to use /etc/sysctl.conf is out of date too: that file does not exist on Ubuntu 26.04, and a setting belongs in /etc/sysctl.d/. Many hardening lists set 3, the old Debian value. On Ubuntu 26.04 that loosens the kernel: at 3 an unprivileged user can count their own processes again.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo sysctl -w kernel.perf_event_paranoid=3 perf stat -e task-clock true
kernel.perf_event_paranoid = 3 Performance counter stats for 'true': 0.14 msec task-clock:u 0.000395306 seconds time elapsed 0.000000000 seconds user 0.000437000 seconds sys

On RHEL 10 the default is 2, the upstream maximum: users may measure their own processes but not profile the kernel. The rule is to leave this key at the distribution value and never copy a number from one distribution to another. If developers need unprivileged perf on a build host, grant it there and document the exception.

deploy@rocky10 · Rocky Linux 10.2
$ perf stat -e task-clock true
Performance counter stats for 'true': 213,833 task-clock:u # 0.160 CPUs utilized 0.001338286 seconds time elapsed 0.001375000 seconds user 0.000000000 seconds sys

Four settings are one-way latches: once set, the kernel refuses to change them until the next boot. kernel.modules_disabled = 1 stops all module loading and unloading (the modules lesson covers it), kernel.kexec_load_disabled = 1 stops kexec (booting straight into another kernel from the running one, without the firmware; kdump crash capture uses it), kernel.yama.ptrace_scope = 3 stops all attaching, and kernel.unprivileged_bpf_disabled = 1 locks unprivileged BPF off. They are all off here. Their strength is that an attacker with root cannot undo them either; their cost is that you cannot either, so a mistake needs a reboot, and a latch in a sysctl.d file applies at every boot. This lesson only reads them; set one only on a host whose console you can reach and after testing the boot. One Ubuntu-only value from the first listing, kernel.apparmor_restrict_unprivileged_userns = 1, limits unprivileged user namespaces through AppArmor; the AppArmor lesson explains it.

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl kernel.modules_disabled kernel.kexec_load_disabled
kernel.modules_disabled = 0 kernel.kexec_load_disabled = 0

Link protections, core dumps and the baseline file

fs.protected_symlinks = 1 defends a classic race: a user plants a symbolic link in a shared, sticky directory such as /tmp, and a root process that opens the expected path follows the link to a file the user could not touch. With the protection on, a process may follow a link in a world-writable sticky directory only if it owns the link or the directory's owner does. Even root is refused:

deploy@web01 · Ubuntu 26.04 LTS
$ ln -s /etc/hostname /tmp/hard-sysctlkern-link sudo cat /tmp/hard-sysctlkern-link
cat: /tmp/hard-sysctlkern-link: Permission denied
$ cat /tmp/hard-sysctlkern-link rm /tmp/hard-sysctlkern-link
web01

fs.protected_hardlinks = 1 stops users hard-linking files they cannot read or write, and fs.protected_fifos and fs.protected_regular stop a privileged process from opening someone else's FIFO or file in a sticky directory when it meant to create its own. All are on by default; Ubuntu sets protected_regular to 2, which also covers group-writable sticky directories. kernel.randomize_va_space = 2 is full address-space layout randomisation, the upstream default; a value below 2 on a server is drift to investigate.

fs.suid_dumpable decides whether a process that changed privileges (a SUID program, or a daemon that dropped from root to its own user) may leave a core dump. Both distributions ship 2, which allows it only through a crash handler that is a pipe or an absolute path (apport on Ubuntu, systemd-coredump on RHEL). CIS asks for 0. The trade is real: a core dump of such a process can contain the secrets it held, but with 0 a crash of sshd-session or chronyd leaves no core to debug.

The file below records all of this: it pins the good defaults, changes suid_dumpable (and ptrace_scope on RHEL), and says in a comment why perf_event_paranoid is absent. The 60- prefix sorts after the distribution files, as the network lesson explained.

/etc/sysctl.d/60-secopslog-kernel.conf
# SecOpsLog kernel baseline. CHANGES marks a key that differs from the
# Ubuntu 26.04 or RHEL 10 default; PINS keeps a default so drift is undone.
# Comments stay on their own lines: sysctl.d has no trailing comments.
# PINS on Ubuntu, CHANGES on RHEL (0): ptrace only your own descendants.
kernel.yama.ptrace_scope = 1
# PINS: kernel pointers and the kernel log need CAP_SYSLOG.
kernel.kptr_restrict = 1
kernel.dmesg_restrict = 1
# PINS: unprivileged bpf() off, in the mode root can still change.
kernel.unprivileged_bpf_disabled = 2
# PINS: full address-space randomisation and the link protections.
kernel.randomize_va_space = 2
fs.protected_symlinks = 1
fs.protected_hardlinks = 1
# CHANGES (2 on both): no core dumps from SUID or privilege-dropping
# programs. Their crashes then leave no core for apport or systemd-coredump.
fs.suid_dumpable = 0
# kernel.perf_event_paranoid is left at the distribution value on purpose
# (Ubuntu 4, RHEL 2); a copied value can only loosen Ubuntu's.

sysctl -p FILE applies one file and prints what it wrote, which is the verification. Rollback is to delete the file and write back the one value it changed, because none of these keys return by themselves until reboot.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo sysctl -p /etc/sysctl.d/60-secopslog-kernel.conf
kernel.yama.ptrace_scope = 1 kernel.kptr_restrict = 1 kernel.dmesg_restrict = 1 kernel.unprivileged_bpf_disabled = 2 kernel.randomize_va_space = 2 fs.protected_symlinks = 1 fs.protected_hardlinks = 1 fs.suid_dumpable = 0
$ sudo rm /etc/sysctl.d/60-secopslog-kernel.conf sudo sysctl -w fs.suid_dumpable=2
fs.suid_dumpable = 2

Try this

On a lab machine where you have sudo, start sleep 600 & in one shell and note its PID. From a second shell as the same user, try strace -p PID and read the refusal; then try sudo strace -p PID and stop it with Ctrl-C. Next, run sudo bpftrace -e 'kprobe:do_nanosleep { printf("%s", kstack(4)); exit(); }' -c 'sleep 0.2' and note the function names. Set sudo sysctl -w kernel.kptr_restrict=2, run the same command, and see hex addresses instead. Put it back with sudo sysctl -w kernel.kptr_restrict=1, confirm with sysctl kernel.kptr_restrict, and kill the sleep.

Takeaway

Read each kernel value and its source before you write one, pin the defaults that are already right, and never copy a number across distributions: perf_event_paranoid = 3 loosens Ubuntu, and kptr_restrict = 2 blinds your own tracing tools. Treat the one-way latches as boot-time decisions, not settings you try out.

Quick check
01An APM agent that attaches to application workers running under the same user works on your RHEL 10 servers but fails with "Operation not permitted" on Ubuntu 26.04. What is the least disruptive fix?
Incorrect — That reopens same-user memory reading for every process on the host to fix one agent, and CIS asks for 1.
Correct — Scope 1 still lets a process with CAP_SYS_PTRACE attach, and PR_SET_PTRACER lets a process name its tracer, so only the agent gets the exception.
Incorrect — A different user is refused by the ordinary same-user rule at every scope; that makes the failure permanent.
Incorrect — Scope 2 requires CAP_SYS_PTRACE for everyone; the agent still fails unless it gets the capability, which works at scope 1 too.
02After you set kernel.kptr_restrict = 2 on a server, sudo bpftrace prints kernel stacks as hex addresses instead of function names. Why?
Incorrect — kptr_restrict changes no files; it changes what reads of /proc/kallsyms return.
Incorrect — The probe attached and collected the stack; only the address-to-name step failed.
Correct — Level 1 hides addresses only from users without CAP_SYSLOG; level 2 hides them from root too, and tracing tools use that table.
Incorrect — The two keys are independent, and bpftrace resolves names from the symbol table, not the log.
03Your team's hardening file from an older guide sets kernel.perf_event_paranoid = 3. What does applying it to a default Ubuntu 26.04 server do?
Correct — The lab shows perf stat failing for a user at 4 and succeeding at 3. The file weakens Ubuntu's default.
Incorrect — The lab shows otherwise: at 3 an unprivileged perf stat on the user's own process succeeds.
Incorrect — 3 is higher than RHEL's 2 but lower than Ubuntu's shipped 4; what matters is the value you replace.
Incorrect — Ubuntu's kernel accepts both 3 and 4 (a Debian and Ubuntu option); the write succeeds and changes behaviour.

Related