eBPF tools: bcc and bpftrace
Ready-made tools and one-liners in production.
Earlier lessons in this course used a few eBPF tools, runqlat and biolatency, as ready-made instruments. This lesson is about the toolkit itself: what an eBPF tool is, which kinds of probe exist and when to use each, how to write a bpftrace one-liner that answers a question no fixed tool answers, what the bcc collection offers on Ubuntu and RHEL today, and what a probe costs the system you are measuring.
What an eBPF tool is, and where it attaches
eBPF lets you run a small program inside the kernel when an event happens, such as a system call starting, a kernel function being entered or a timer firing. Before the kernel accepts the program, its verifier checks it: the program must finish, may only read memory it is allowed to, and cannot crash the kernel. The program's results go into maps, key-value tables in kernel memory that the program updates on every event and that a user-space tool reads when it wants: a count per process, or a histogram of latencies. Because the counting happens in the kernel, only the summary crosses to user space, which is why these tools can watch events that happen millions of times a second.
You rarely write such a program by hand. Two front ends do it for you. bpftrace compiles a short script, often a one-liner, into an eBPF program, loads it, and prints its maps when it ends. bcc is a library plus a collection of ready-made tools, each a Python program with embedded C. Loading tracing programs needs root. How the kernel loads and verifies programs in depth, and how to watch them from the security side, is covered in the advanced security course's eBPF lesson (optional); this lesson concentrates on performance work.
Every tracing program runs when an event fires, and the event source decides what you can see, how stable it is across kernel versions, and how much each hit costs.
Tracepoints are hooks the kernel developers placed on purpose, with named arguments. fentry and fexit attach to the start and end of almost any kernel function through a BPF trampoline, and read its arguments with their C types from BTF, the kernel's built-in type information; bpftrace's documentation says the trampolines let the kernel call a BPF program "with near zero overhead", where a kprobe reaches the same function through the kprobe mechanism (a trap, or an ftrace hook at a function's entry), which costs more per hit than a trampoline. Function names and arguments are not a stable interface, so a probe on an internal function can break on the next kernel. uprobes do the same for functions in user-space binaries and libraries, and USDT probes are markers that a program's authors compiled in. On the Ubuntu lab, bpftrace lists what it can attach to:
bpftrace 0.25, 132 bcc tools with the -bpfcc suffix and 39 bpftrace tools (.bt) are installed; about 74,000 kernel functions can take an fentry probe and 2548 tracepoints exist. The -v option shows what a probe gives you:
The fentry probe on vfs_fsync_range, the kernel function behind fsync() and synchronous writes, sees the function's real arguments, starting with the struct file being synced. The syscall tracepoint only has the system call number and the return value.
One-liners that answer one question
The workload is a writer that behaves like a small database log: dd writing 4 KiB blocks with oflag=dsync, so every write() waits until the data is on the disk. It runs as your account, in a scope with a ten-minute limit.
A bpftrace program is a list of clauses of the form probe /filter/ { action }. The probe names the event (tracepoint:syscalls:sys_enter_write); the optional filter between slashes decides whether this event counts (/comm == "dd"/: only when the current process is called dd); the action runs when it does. Variables whose names start with @ are maps: @start[tid] is a table keyed by thread ID, and @usecs = hist(...) builds a histogram. Built-in variables describe the event: tid the thread, comm the process name, nsecs a timestamp in nanoseconds, and args the probe's arguments. bpftrace prints every map that is left when it exits.
First question: how long does each write() take inside the kernel for this process? iostat sees the disk but not the process, and strace -T would stop the process at every call. The one-liner stores a timestamp per thread when the call enters, and on exit adds the elapsed time to a histogram, all inside the kernel; _ = delete(...) removes the entry (the _ = discards the value delete returns, which bpftrace 0.25 otherwise warns about). interval:s:5 ends it after five seconds; at a terminal you would press Ctrl-C.
Most writes took 32 to 127 µs, and a handful took several milliseconds, one of them 8 to 16 ms: the tail that an average would hide. Only the summary crossed into user space, printed once at the end. Second question: which files is the kernel syncing, and for whom? With fentry, the arguments are typed, so the one-liner can follow the file pointer to the name of the file.
All syncs in those five seconds were dd on bt-sync.dat: one per written block, because O_DSYNC turns every write into a sync of that file's data. On a real server the same line shows which process makes the journal or database files sync, a question that neither iostat nor pidstat can answer.
bcc: the ready-made tools, as they are today
bcc is a collection of more than a hundred tools written in Python with embedded C. Package names and paths differ: Ubuntu's bpfcc-tools installs them in /usr/sbin with a -bpfcc suffix (opensnoop-bpfcc), RHEL's bcc-tools in /usr/share/bcc/tools without a suffix and not on PATH. Each tool compiles its C code with LLVM when it starts, against the running kernel's headers. That design is the weak point: when the kernel's headers change faster than the bcc package, tools stop compiling. Run ten common tools for 15 seconds each on both labs and look only at the exit status:
On Ubuntu 26.04 with bcc 0.35 and kernel 7.0, five of the ten fail at start-up: execsnoop, biolatency, runqlat, tcpconnect and tcplife (execsnoop's error, shown in the optional advanced security course's eBPF lesson, is a static assertion in the kernel's fs.h that clang rejects). The other five work. On Rocky Linux 10.2, the same bcc version compiles all ten against the 6.12 kernel. Check a tool on your own kernel before an incident, not during one. A working bcc tool worth knowing is funclatency: a latency histogram for any kernel function, named on the command line, with no script to write.
Each vfs_fsync_range call took 79 µs on average, most between 32 and 127 µs, over 58995 calls in five seconds, with a few outliers up to 65 to 131 ms: the same shape as the write() histogram, because almost all of a synchronous write's time is the sync. On the Rocky lab the same tool is /usr/share/bcc/tools/funclatency:
When a bcc tool fails, use its bpftrace version: bpftrace reads kernel types from BTF rather than headers, and its tools are in /usr/sbin on Ubuntu (/usr/share/bpftrace/tools on RHEL). The disk I/O and CPU scheduler lessons of this course used biolatency.bt and runqlat.bt this way.
Inside a program: USDT and uprobes
Kernel probes cannot see a garbage collector or an allocator, because their work happens without system calls. A Python service that stalls now and then is the example. This program creates objects that refer to themselves, which only Python's cyclic garbage collector can free, and a 1 MiB buffer each time round. Create a directory for it first (mkdir -p ~/bt).
# app.py: builds objects that point at themselves; only the cyclic GC frees themimport timeclass Node:def __init__(self):self.peer = selfend = time.monotonic() + 600while time.monotonic() < end:batch = [Node() for _ in range(50000)]buf = bytearray(1 << 20)time.sleep(0.05)
Ubuntu's Python 3.14 is built with USDT markers, including gc__start and gc__done around each collection. Each marker has a semaphore, a counter the program checks so that it skips the probe's work when nobody is tracing; bpftrace activates it when it attaches. -p limits the probes to one process, and pgrep -u $USER -x python3 picks your Python and not the sudo process that started it (pgrep -f app.py would match both). How long does each collection pause the program?
bpftrace 0.25 first prints compiler warnings from building its USDT support (left out); they are harmless. At the end it also prints the @start entry of a collection still running when it stopped (left out). The histogram has two populations: 671 collections finished in under a microsecond, and most of the rest took 256 µs to 1 ms, with a few up to 8 ms. Those are the pauses the service's requests would feel, measured without changing a line of the program. Next, a uprobe on the C library's malloc: which allocation sizes reach it? (On x86_64 the library is /usr/lib/x86_64-linux-gnu/libc.so.6.)
77 calls asked for about 1 MiB, one per loop: the bytearray. Another 76 were just over 512 bytes; the Node objects never reach malloc at all, because Python serves allocations up to 512 bytes from its own allocator. Each uprobe hit is a breakpoint trap in the traced process, so a uprobe on a function called millions of times a second slows the program noticeably; the same applies to ltrace, with less reliable results (see the process debugging lesson).
What a probe costs, and the limits
A probe's cost is paid on every event it fires on, whether or not you print anything. Measure it. dd with a block size of one byte makes about four million system calls for 2 MB, which makes the per-call cost visible. Run it bare, then with a probe on every system call entry, with kernel.bpf_stats_enabled switched on so the kernel counts the time spent in each BPF program.
bpftool prog show lists the program bpftrace loaded and, with statistics on, run_cnt (4012028 runs) and run_time_ns (188 ms): about 47 ns per run. dd itself slowed from 1.15 to 1.52 s, about 90 ns per system call, because entering the probe costs more than the program's own instructions; collecting the statistics adds a little too, which is why they are off by default. These are this VM's numbers, on an Apple-silicon host: the cost per event depends on the CPU and on the kernel's mitigations for speculative-execution attacks, and x86_64 servers often pay more per tracepoint, so measure on your own hardware. The probe also fires for every process on the host, not only the one you care about, so estimate from the host's total rate: at about 100 ns, 100,000 system calls a second cost about 1% of a CPU, and a busy host making several million a second pays several tenths of a CPU. After the tracer ends, its map shows who made the calls:
dd made 4,000,223 of them. pgrep and sleep are the lab's own polling while it waited. Three more limits matter in practice. Tracing needs privilege: kernel.unprivileged_bpf_disabled is 2 on both platforms, so these tools run as root (or with CAP_BPF and CAP_PERFMON), and kernel lockdown under Secure Boot can refuse some helpers. Output has a finite buffer between the kernel and bpftrace, so a one-liner that prints on every event can lose events when they arrive faster than bpftrace reads them; aggregate in maps and print summaries, as every example here does. And probes on kernel functions are tied to one kernel's internals: the same one-liner may need changes after an upgrade, where a tracepoint would not.
interval or a timeout, and check bpftool prog show with statistics on if you are unsure what a probe costs.Try this
Start the dd writer again, then answer a new question with fexit, which runs when a function returns and sees its return value as retval along with its arguments: sudo bpftrace -e 'fexit:vfs_fsync_range /comm == "dd"/ { @[retval] = count(); } interval:s:5 { exit(); }'. Predict first: every sync succeeds, so all calls should be counted under 0. Then change @[retval] to @[args.datasync] and explain the result: O_DSYNC asks only for the data and the metadata needed to read it back, so the kernel passes datasync as 1. Stop the writer with sudo systemctl stop bt-sync.scope.
Takeaway
Pick the most stable probe that answers the question, aggregate in the kernel instead of printing every event, and know what the probe costs per event before you put it on a hot path; when a bcc tool does not compile on your kernel, its bpftrace version usually does.