eBPF for security visibility

Probes, the verifier and bpftrace on real events.

Advanced16 min · lesson 9 of 15

eBPF lets you load a small program into the running kernel and attach it to an event: a process starting, a file opening, a TCP connection leaving. The program sees the event as it happens, from inside the kernel, which is why modern runtime sensors such as Falco are built on it. This lesson explains what eBPF is and why the kernel agrees to run your code, uses bpftrace one-liners to watch program execution, file opens and outbound connections made by an ordinary account, lists the programs and maps the kernel holds with bpftool, and covers who may load BPF and what eBPF cannot see. It is written from the defender's side and stands on its own; linux-perf/bpftools goes deeper into the tooling.

What eBPF is, and why the kernel trusts it

An eBPF program is bytecode for a small virtual instruction set, loaded with the bpf() system call. Before the kernel accepts it, the verifier (a checker inside the kernel) walks every path through the program and rejects it unless it can prove the program ends, reads only memory it is allowed to, and reaches kernel data only through a fixed set of helper functions. An accepted program is JIT-compiled (translated to native machine code) and attached to a hook. Maps are kernel-side key/value stores and queues that let a program keep state between events and pass results to a user-space tool.

BTF (BPF Type Format) describes the running kernel's data structures, so one compiled program can adapt to many kernel versions (CO-RE, compile once, run everywhere).

From one-liner to live probe
1bpftrace script
what to watch and what to print
2Compiled to BPF bytecode
with BTF describing the kernel
3bpf() system call
needs CAP_BPF plus CAP_PERFMON, or root
4Verifier
proves it terminates and stays in bounds
5JIT and attach
native code on a tracepoint or kprobe
6Maps and ring buffer
state kept in the kernel, events out
The verifier makes a crash far less likely; a kernel module gets no such check.

That is the difference from a kernel module, which runs with no such proof and can crash or subvert the kernel. The verifier is not a guarantee, though: bugs in it have been a recurring source of kernel privilege-escalation vulnerabilities, one reason both lab platforms keep unprivileged BPF switched off. And it proves memory safety, not good intent or low cost. A program that passes can still read what it is allowed to read, including other people's data; a probe on a busy code path adds work to every call; and some program types exist to change behaviour (deny an action, drop a packet).

Program types differ by hook. Tracing programs, the only kind this lesson loads, attach to tracepoints, kprobes (almost any kernel function), fentry/fexit (function entry and exit, cheaper than kprobes) and uprobes (functions inside a user-space program). LSM programs attach to the kernel's security hooks and can deny an action. XDP and tc programs process network packets, and cgroup programs filter what a group of processes may do. The lab host has what tracing needs.

deploy@web01 · Ubuntu 26.04 LTS
$ uname -r; bpftrace --version; sysctl kernel.unprivileged_bpf_disabled
7.0.0-34-generic bpftrace v0.25.0 kernel.unprivileged_bpf_disabled = 2
$ ls -l /sys/kernel/btf/vmlinux
-r--r--r-- 1 root root 8132176 Sep 27 08:26 /sys/kernel/btf/vmlinux

Kernel 7.0 exposes its BTF at /sys/kernel/btf/vmlinux, and bpftrace 0.25 is installed (on the Rocky lab, bpftrace 0.24.2). kernel.unprivileged_bpf_disabled = 2 means accounts without privilege cannot use bpf() at all, a setting an administrator may change back at run time; 1 would lock it off until reboot, and 0 would allow some unprivileged program types. Both lab platforms default to 2. So the ordinary account's attempt to list programs fails.

deploy@web01 · Ubuntu 26.04 LTS
$ bpftool prog list
Error: can't get next program: Operation not permitted

Listing and loading programs both take privilege, as the Operation not permitted shows. For tracing, the capabilities are CAP_BPF for BPF operations plus CAP_PERFMON for performance and tracing hooks, split out of CAP_SYS_ADMIN in Linux 5.8 so a sensor need not run as full root. Anyone holding them can watch the whole machine. Walking the list of every loaded program, which bpftool prog list does, still needs CAP_SYS_ADMIN, so in practice the inventory below runs as root.

Attach points: tracepoints and their arguments

A tracepoint is a hook the kernel developers placed on purpose and keep stable across versions, such as the entry and exit of every system call. A kprobe can attach to almost any kernel function by name, but that name is not a promise. If a later kernel renames or removes the function, the probe fails to attach with an error when it loads. The dangerous case is quieter: the function still exists, but the compiler has inlined it at some call sites or the code path you cared about no longer calls it, so the probe attaches and simply stops firing for those events. For detections that must survive kernel upgrades, prefer tracepoints (or fentry with BTF). bpftrace -l lists attach points and -lv shows the arguments a tracepoint gives you.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -l "tracepoint:syscalls:sys_enter_exec*"
tracepoint:syscalls:sys_enter_execve tracepoint:syscalls:sys_enter_execveat
$ sudo bpftrace -lv tracepoint:syscalls:sys_enter_openat
tracepoint:syscalls:sys_enter_openat int __syscall_nr int dfd const char * filename int flags umode_t mode __data_loc char[] __filename_val

sys_enter_openat hands a program the directory descriptor, a pointer to the filename, the flags and the mode, exactly the arguments of openat(2). The one-liners below read these as args.filename and so on.

Watching ordinary activity with bpftrace

A throwaway account, eb-dev, will do ordinary things while bpftrace watches only its UID. Each one-liner starts the account's commands in the background with a short delay, attaches, and ends itself after eight seconds with an interval probe; at a terminal you would leave that probe out and stop the tracer with Ctrl-C. The UID reaches the filter as bpftrace's first positional parameter: $1 in the script is the first argument after it, here $(id -u eb-dev), so no number is hard-coded and the same line works on your host.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo useradd -m -s /bin/bash eb-dev id -u eb-dev
1002
$ sudo -u eb-dev sh -c "sleep 4; id -un; uname -r" >/dev/null & sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve /uid == $1/ { printf("%-7d %-7d %-6s %s\n", curtask->real_parent->tgid, pid, comm, str(args.filename)); } interval:s:8 { exit(); }' $(id -u eb-dev)
Attached 2 probes 81525 81536 sh /usr/bin/id 81525 81537 sh /usr/bin/uname

Each line is one execve call: the parent's PID (the shell; curtask->real_parent->tgid reads it from the kernel's own task structure), the PID of the child the shell forked to run the command, comm (the name of the process making the call, still sh because the new program has not replaced it yet) and the program being run. The shell ran id and uname. This is the raw material of every "unexpected child process" detection, such as a web server whose child is a shell. Next, file opens and their results. The enter probe stores the filename pointer in a map keyed by thread ID, and the exit probe reads it back together with the return value, then deletes the entry (_ = discards the value delete() returns; bpftrace 0.25 prints a warning when that value is silently dropped).

deploy@web01 · Ubuntu 26.04 LTS
$ sudo -u eb-dev sh -c "sleep 4; cat /etc/shadow; cat /etc/hostname" >/dev/null 2>&1 & sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat /uid == $1/ { @path[tid] = args.filename; } tracepoint:syscalls:sys_exit_openat /@path[tid]/ { $p = str(@path[tid]); _ = delete(@path, tid); printf("%-6s %-34s %d\n", comm, $p, args.ret); } interval:s:8 { exit(); }' $(id -u eb-dev)
Attached 3 probes cat /etc/ld.so.cache 3 … cat /usr/share/coreutils/locales/cat/en-US.ftl -2 cat /etc/shadow -13 … cat /etc/hostname 3

Every cat first opens the loader cache and its shared libraries (most lines, elided here), and -2 is ENOENT, a translation file that does not exist. The line that matters is /etc/shadow -13: -13 is EACCES, an unprivileged account trying and failing to read the password hashes, while /etc/hostname 3 succeeded with file descriptor 3. Failed attempts on sensitive files are exactly what application logs never show, because the program wrote nothing; an audit rule on that file would record it (linux-det/auditpipe), but only for files you chose in advance. Last, outbound connections. The sock:inet_sock_set_state tracepoint fires on every TCP state change; state 2 is SYN_SENT, the moment a connection attempt leaves. A probe takes one predicate, so the state test and the UID test are joined with &&.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo -u eb-dev sh -c "sleep 4; curl -s --connect-timeout 2 http://192.0.2.10/" >/dev/null 2>&1 & sudo bpftrace -e 'tracepoint:sock:inet_sock_set_state /args.newstate == 2 && uid == $1/ { printf("%-6s pid=%-7d uid=%-5d -> %s:%d\n", comm, pid, uid, ntop(args.daddr), args.dport); } interval:s:8 { exit(); }' $(id -u eb-dev)
Attached 2 probes curl pid=81614 uid=1002 -> 192.0.2.10:80

curl, running as eb-dev (UID 1002), tried to reach 192.0.2.10 on port 80. That address is from TEST-NET-1, reserved for documentation, so nothing answered; the kernel still reported the attempt, with the process and user behind it. A beacon to an attacker's server looks the same, which is why "who opened a connection to where" is a staple sensor signal.

Inventory: which programs is the kernel running?

Every loaded program and map is visible to root through bpftool. Knowing what is normally there is the baseline that makes an unexpected program stand out.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftool prog list | grep -E "^[0-9]+:" | awk "{print \$2}" | sort | uniq -c
6 cgroup_device 6 cgroup_skb 1 cgroup_sysctl 1 tracepoint
$ sudo bpftool prog list | grep -A2 "name lima"
66: tracepoint name lima_ticker tag c317a0de92703b42 loaded_at 2026-09-27T08:26:09+0000 uid 0 xlated 104B jited 168B memlock 4096B map_ids 29

On this host systemd has loaded cgroup_device, cgroup_skb and cgroup_sysctl programs to enforce unit settings such as device access and IP filtering. The one tracepoint program, lima_ticker, belongs to the lab VM tool's guest agent, not to Ubuntu. Now start a tracer in the background, as a sensor would run, and look again while it is attached.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -e 'tracepoint:syscalls:sys_enter_connect { @connects[comm] = count(); }' >/tmp/eb-connects.txt 2>&1 &
$ sudo bpftool prog list | grep -A3 "name tracepoint_syscalls"
999: tracepoint name tracepoint_syscalls_sys_enter_connect_1 tag 9cd6cbbd503054eb gpl loaded_at 2026-09-27T08:39:39+0000 uid 0 xlated 736B jited 600B memlock 4096B map_ids 61,63,64,60 btf_id 160
$ sudo bpftool map list | grep -A1 "name AT_connects"
60: percpu_hash name AT_connects flags 0x0 key 16B value 8B max_entries 4096 memlock 426880B

bpftrace's program appears as a tracepoint named after its probe, loaded by uid 0, with the IDs of the maps it uses. AT_connects is the @connects map, a per-CPU hash (one copy per CPU, so CPUs never contend) with room for 4096 keys. The kernel also audits every program load and unload by itself, with no rule needed, as long as auditing is switched on (here auditd switched it on).

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ausearch --input-logs -m BPF -c bpftrace -ts recent -i | grep -E "op=LOAD|syscall=bpf" | tail -2
type=SYSCALL msg=audit(09/27/26 08:39:39.430:9825) : arch=aarch64 syscall=bpf success=yes exit=14 a0=BPF_PROG_LOAD a1=0xffffde2300b0 a2=0x98 a3=0xc5f355b66cec items=0 ppid=81698 pid=81703 auid=deploy uid=root gid=root euid=root suid=root fsuid=root egid=root sgid=root fsgid=root tty=(none) ses=42 comm=bpftrace exe=/usr/bin/bpftrace subj=unconfined key=(null) type=BPF msg=audit(09/27/26 08:39:39.430:9825) : prog-id=1001 op=LOAD

The BPF record gives a program ID and op=LOAD (bpftrace loads more than one program, and tail kept its last); its SYSCALL record names the loader (comm=bpftrace) and the person behind it (auid=deploy, as linux-det/auditpipe explains). Now stop the tracer the way Ctrl-C would, with SIGINT, and look for the program's ID from the listing above.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo pkill -INT -x bpftrace sleep 1 sudo bpftool prog list | grep "name tracepoint_syscalls" || echo "no bpftrace program loaded"
no bpftrace program loaded
$ sudo ausearch --input-logs -m BPF -ts recent -i | grep 'prog-id=999 '
type=BPF msg=audit(09/27/26 08:39:39.429:9822) : prog-id=999 op=LOAD type=BPF msg=audit(09/27/26 08:39:39.815:9857) : prog-id=999 op=UNLOAD

The program left the kernel with its loader, and the audit log holds both ends of its life: prog-id=999 op=LOAD and op=UNLOAD. Forwarding these records gives you a history of every program ever loaded on the host. They depend on the audit subsystem being enabled, though, and on a default Ubuntu server, with no auditd, it is not.

deploy@web01 · Ubuntu 26.04 LTS (default install)
$ sudo journalctl -b -k -o cat | grep -m1 "audit: initializing"
audit: initializing netlink subsys (disabled)
$ sudo bpftool prog list | grep -cE "^[0-9]+:" sudo journalctl -b -k -o cat | grep "type=1334" | wc -l
17 0

The kernel initialised audit disabled at boot, and the 17 programs loaded since then (systemd's, plus the VM tool's lima_ticker) left no BPF record (type 1334) at all. Install auditd (linux-hard/auditd) before you rely on this history.

A loaded program you did not expect is an incident
The same access that lets a sensor watch the host lets an intruder with root or CAP_BPF hide files from tools, read data from TLS libraries through uprobes, or drop their own traffic before your sensor sees it; public eBPF rootkits do exactly this. Keep kernel.unprivileged_bpf_disabled at 1 or 2, run auditd and forward the kernel's BPF audit records off the host, and compare bpftool prog list against a baseline. A program you cannot attribute to a package or a known agent is a finding, not a curiosity.

What eBPF does not see

Probes attached from inside the kernel are hard to evade from user space, but they are not the whole truth. An attacker with root can unload a sensor's programs, detach them, or load their own that tamper with the view, so a sensor needs its own liveness check and its alerts must leave the host. Events can be dropped: a ring buffer that fills faster than the user-space side reads it loses events, and tools report that as lost or dropped counts. A sys_enter probe reads arguments from user memory before the kernel does, so a determined attacker can change the buffer in between (a time-of-check, time-of-use race), which is why sensors such as Falco attach extra programs to cross-check it. Operations submitted through io_uring(7) do not pass through the per-call system-call tracepoints, so a file opened that way never fires sys_enter_openat; hooks deeper in the kernel (LSM or VFS functions) catch both. A process inside a container reports its host PID. A kprobe on an internal function can go silent on a kernel upgrade, as described above. And tracing has a cost: a probe on every openat runs on every call, whoever makes it, so on a busy production host keep ad-hoc tracing short and filtered inside the probe, as these one-liners are. Finally, tooling can lag the kernel: on Ubuntu 26.04 several bcc 0.35 tools compile their C at run time against the 7.0 kernel headers and currently fail, as execsnoop-bpfcc does here.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo timeout 30 execsnoop-bpfcc
… include/linux/fs.h:2431:15: error: static assertion failed due to requirement 'sizeof(struct filename) % 64 == 0': sizeof(struct filename) % 64 == 0 … 1 warning and 1 error generated. … Exception: Failed to compile BPF module <text>

clang rejected a static assertion in the kernel's own fs.h header, so the tool never loaded; that is a packaging mismatch between bcc and this kernel, not a problem with eBPF. The bpftrace one-liners above, and the ready-made .bt tools the bpftrace package installs (/usr/sbin/execsnoop.bt, opensnoop.bt, tcpconnect.bt on Ubuntu), work. On RHEL the bcc tools live in /usr/share/bcc/tools and the same tool runs on the Rocky lab's 6.12 kernel.

deploy@rocky10 · Rocky Linux 10.2
$ uname -r; bpftrace --version; sysctl kernel.unprivileged_bpf_disabled; ls /usr/share/bcc/tools | wc -l
6.12.0-211.16.1.el10_2.0.1.aarch64 bpftrace v0.24.2 kernel.unprivileged_bpf_disabled = 2 131
$ sudo -u eb-dev sh -c "sleep 15; id -un" >/dev/null & sudo env PYTHONUNBUFFERED=1 timeout 25 /usr/share/bcc/tools/execsnoop -u eb-dev
COMM PID PPID RET ARGS id 5790 5785 0 /bin/id -un
# PYTHONUNBUFFERED=1 only makes the Python tool flush output when it is not writing to a terminal; timeout ended it (exit 124)

Try this

On the Ubuntu lab host, create a throwaway account with sudo useradd -m eb-test. In one terminal, start the connection tracer for it without the interval probe, so it runs until you stop it: sudo bpftrace -e 'tracepoint:sock:inet_sock_set_state /args.newstate == 2 && uid == $1/ { printf("%-6s pid=%-7d uid=%-5d -> %s:%d\n", comm, pid, uid, ntop(args.daddr), args.dport); }' $(id -u eb-test). In a second terminal run sudo -u eb-test curl -s --connect-timeout 2 http://192.0.2.10/ and watch one line appear with curl, the UID and 192.0.2.10:80. While the tracer runs, find its program with sudo bpftool prog list | grep -A3 'name tracepoint_sock' and note the ID before the colon. Stop the tracer with Ctrl-C and confirm the program has gone from the list. Then find that ID twice, as op=LOAD and op=UNLOAD, with sudo ausearch -m BPF -ts recent -i | grep 'prog-id=<ID> '. Remove the account with sudo userdel -r eb-test.

Takeaway

Use tracepoints for anything you depend on, keep a baseline of bpftool prog list and forward the kernel's BPF audit records, and treat an unexplained loaded program as an intrusion until proven otherwise. eBPF shows what happens at the hooks you chose, so know which paths, such as io_uring, go around them.

Quick check
01Your detection uses a kprobe on an internal kernel function. After a routine kernel upgrade it stops alerting, and no error appears anywhere. What is the most likely cause?
Incorrect — The verifier accepts or rejects a whole program with an error message; it never loads part of one, so this would not be silent.
Incorrect — That sysctl only controls whether unprivileged callers may use bpf(); a sensor running as root is unaffected and nothing is detached.
Incorrect — A full ring buffer drops events while it is full and reports losses; it recovers as soon as the reader catches up.
Correct — Internal functions carry no stability promise. A rename would fail loudly at attach time; inlining or a changed call path leaves the probe attached and silent. Tracepoints (or fentry with BTF) survive upgrades.
02A colleague argues that loading an eBPF program is as risky for stability as loading a kernel module, since both run inside the kernel. What is the real difference?
Incorrect — eBPF programs execute in the kernel after JIT compilation; that is what lets them see events as they happen.
Correct — A module is trusted as soon as it loads; an eBPF program is refused unless the verifier can prove it safe, which makes a crash far less likely, although verifier bugs exist and a verified program can still cost CPU.
Incorrect — Accepted programs are JIT-compiled to native code; the protection comes from the proof before loading, not from supervision afterwards.
Incorrect — Signing is not the mechanism described here; the kernel's safety check for eBPF is the verifier, which modules do not go through.
03bpftool prog list on a production host shows a tracepoint program loaded by uid 0 that no package or agent you run accounts for. What is the right reading?
Incorrect — Legitimate loaders exist, which is exactly why you keep a baseline; a program you still cannot attribute after checking it is not noise.
Incorrect — The verifier checks memory safety, not purpose; a verified tracing program can still read process data and hide activity.
Correct — Unattributed BPF programs are a known rootkit technique; the BPF audit records show who loaded it and when.
Incorrect — The list is the kernel's current set of loaded programs; if it is listed, it is loaded now.

Related