eBPF tools: bcc and bpftrace

Ready-made tools and one-liners in production.

Advanced16 min · lesson 19 of 21

Earlier lessons in this course used a few eBPF tools, runqlat and biolatency, as ready-made instruments. This lesson is about the toolkit itself: what an eBPF tool is, which kinds of probe exist and when to use each, how to write a bpftrace one-liner that answers a question no fixed tool answers, what the bcc collection offers on Ubuntu and RHEL today, and what a probe costs the system you are measuring.

What an eBPF tool is, and where it attaches

eBPF lets you run a small program inside the kernel when an event happens, such as a system call starting, a kernel function being entered or a timer firing. Before the kernel accepts the program, its verifier checks it: the program must finish, may only read memory it is allowed to, and cannot crash the kernel. The program's results go into maps, key-value tables in kernel memory that the program updates on every event and that a user-space tool reads when it wants: a count per process, or a histogram of latencies. Because the counting happens in the kernel, only the summary crosses to user space, which is why these tools can watch events that happen millions of times a second.

You rarely write such a program by hand. Two front ends do it for you. bpftrace compiles a short script, often a one-liner, into an eBPF program, loads it, and prints its maps when it ends. bcc is a library plus a collection of ready-made tools, each a Python program with embedded C. Loading tracing programs needs root. How the kernel loads and verifies programs in depth, and how to watch them from the security side, is covered in the advanced security course's eBPF lesson (optional); this lesson concentrates on performance work.

Every tracing program runs when an event fires, and the event source decides what you can see, how stable it is across kernel versions, and how much each hit costs.

Where tracing programs attach
In the kernel
tracepoint
placed by kernel developers, stable
fentry / fexit
any kernel function, typed via BTF
kprobe / kretprobe
any kernel function, older and slower
In user space
uprobe / uretprobe
any function in a binary or library
USDT
probe points the program authors added
Timers and counters
profile / interval
sample on a clock, or run periodically
hardware / software
CPU and kernel performance counters
Prefer the most stable source that answers the question.

Tracepoints are hooks the kernel developers placed on purpose, with named arguments. fentry and fexit attach to the start and end of almost any kernel function through a BPF trampoline, and read its arguments with their C types from BTF, the kernel's built-in type information; bpftrace's documentation says the trampolines let the kernel call a BPF program "with near zero overhead", where a kprobe reaches the same function through the kprobe mechanism (a trap, or an ftrace hook at a function's entry), which costs more per hit than a trampoline. Function names and arguments are not a stable interface, so a probe on an internal function can break on the next kernel. uprobes do the same for functions in user-space binaries and libraries, and USDT probes are markers that a program's authors compiled in. On the Ubuntu lab, bpftrace lists what it can attach to:

deploy@web01 · Ubuntu 26.04 LTS
$ bpftrace --version ls /usr/sbin/*-bpfcc | wc -l ls /usr/sbin/*.bt | wc -l
bpftrace v0.25.0 132 39
$ sudo bpftrace -l | cut -d: -f1 | sort | uniq -c
73809 fentry 12 hardware 18 iter 76572 kprobe 1898 rawtracepoint 14 software 2548 tracepoint

bpftrace 0.25, 132 bcc tools with the -bpfcc suffix and 39 bpftrace tools (.bt) are installed; about 74,000 kernel functions can take an fentry probe and 2548 tracepoints exist. The -v option shows what a probe gives you:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -lv 'fentry:vmlinux:vfs_fsync_range'
fentry:vmlinux:vfs_fsync_range struct file * file loff_t start loff_t end int datasync int retval
$ sudo bpftrace -lv 'tracepoint:syscalls:sys_exit_write'
tracepoint:syscalls:sys_exit_write int __syscall_nr long ret

The fentry probe on vfs_fsync_range, the kernel function behind fsync() and synchronous writes, sees the function's real arguments, starting with the struct file being synced. The syscall tracepoint only has the system call number and the return value.

One-liners that answer one question

The workload is a writer that behaves like a small database log: dd writing 4 KiB blocks with oflag=dsync, so every write() waits until the data is on the disk. It runs as your account, in a scope with a ten-minute limit.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=bt-sync --uid=$USER -p MemoryMax=128M timeout 10m \ sh -c 'while :; do dd if=/dev/zero of=/var/tmp/bt-sync.dat bs=4k count=5000 oflag=dsync conv=notrunc status=none; done' > /dev/null 2>&1 &

A bpftrace program is a list of clauses of the form probe /filter/ { action }. The probe names the event (tracepoint:syscalls:sys_enter_write); the optional filter between slashes decides whether this event counts (/comm == "dd"/: only when the current process is called dd); the action runs when it does. Variables whose names start with @ are maps: @start[tid] is a table keyed by thread ID, and @usecs = hist(...) builds a histogram. Built-in variables describe the event: tid the thread, comm the process name, nsecs a timestamp in nanoseconds, and args the probe's arguments. bpftrace prints every map that is left when it exits.

First question: how long does each write() take inside the kernel for this process? iostat sees the disk but not the process, and strace -T would stop the process at every call. The one-liner stores a timestamp per thread when the call enters, and on exit adds the elapsed time to a histogram, all inside the kernel; _ = delete(...) removes the entry (the _ = discards the value delete returns, which bpftrace 0.25 otherwise warns about). interval:s:5 ends it after five seconds; at a terminal you would press Ctrl-C.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -e 'tracepoint:syscalls:sys_enter_write /comm == "dd"/ { @start[tid] = nsecs; } tracepoint:syscalls:sys_exit_write /@start[tid]/ { @usecs = hist((nsecs - @start[tid]) / 1000); _ = delete(@start, tid); } interval:s:5 { exit(); }'
Attached 3 probes @usecs: [32, 64) 44285 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [64, 128) 18768 |@@@@@@@@@@@@@@@@@@@@@@ | [128, 256) 2427 |@@ | [256, 512) 1046 |@ | [512, 1K) 115 | | [1K, 2K) 28 | | [2K, 4K) 12 | | [4K, 8K) 1 | | [8K, 16K) 1 | |

Most writes took 32 to 127 µs, and a handful took several milliseconds, one of them 8 to 16 ms: the tail that an average would hide. Only the summary crossed into user space, printed once at the end. Second question: which files is the kernel syncing, and for whom? With fentry, the arguments are typed, so the one-liner can follow the file pointer to the name of the file.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -e 'fentry:vfs_fsync_range { @[comm, str(args.file->f_path.dentry->d_name.name)] = count(); } interval:s:5 { exit(); }'
Attached 2 probes @[dd, bt-sync.dat]: 66817

All syncs in those five seconds were dd on bt-sync.dat: one per written block, because O_DSYNC turns every write into a sync of that file's data. On a real server the same line shows which process makes the journal or database files sync, a question that neither iostat nor pidstat can answer.

bcc: the ready-made tools, as they are today

bcc is a collection of more than a hundred tools written in Python with embedded C. Package names and paths differ: Ubuntu's bpfcc-tools installs them in /usr/sbin with a -bpfcc suffix (opensnoop-bpfcc), RHEL's bcc-tools in /usr/share/bcc/tools without a suffix and not on PATH. Each tool compiles its C code with LLVM when it starts, against the running kernel's headers. That design is the weak point: when the kernel's headers change faster than the bcc package, tools stop compiling. Run ten common tools for 15 seconds each on both labs and look only at the exit status:

deploy@web01 · Ubuntu 26.04 LTS
$ for t in execsnoop opensnoop biolatency runqlat tcpconnect tcplife cpudist runqlen offcputime syncsnoop; do sudo timeout --preserve-status -s INT 15 $t-bpfcc > /dev/null 2>&1 printf "%-11s exit %s\n" $t $? done
execsnoop exit 1 opensnoop exit 0 biolatency exit 1 runqlat exit 1 tcpconnect exit 1 tcplife exit 1 cpudist exit 0 runqlen exit 0 offcputime exit 0 syncsnoop exit 0
deploy@rocky10 · Rocky Linux 10.2
$ bpftrace --version ls /usr/share/bcc/tools | wc -l ls /usr/share/bpftrace/tools/*.bt | wc -l
bpftrace v0.24.2 131 38
$ for t in execsnoop opensnoop biolatency runqlat tcpconnect tcplife cpudist runqlen offcputime syncsnoop; do sudo timeout --preserve-status -s INT 15 /usr/share/bcc/tools/$t > /dev/null 2>&1 printf "%-11s exit %s\n" $t $? done
execsnoop exit 0 opensnoop exit 0 biolatency exit 0 runqlat exit 0 tcpconnect exit 0 tcplife exit 0 cpudist exit 0 runqlen exit 0 offcputime exit 0 syncsnoop exit 0

On Ubuntu 26.04 with bcc 0.35 and kernel 7.0, five of the ten fail at start-up: execsnoop, biolatency, runqlat, tcpconnect and tcplife (execsnoop's error, shown in the optional advanced security course's eBPF lesson, is a static assertion in the kernel's fs.h that clang rejects). The other five work. On Rocky Linux 10.2, the same bcc version compiles all ten against the 6.12 kernel. Check a tool on your own kernel before an incident, not during one. A working bcc tool worth knowing is funclatency: a latency histogram for any kernel function, named on the command line, with no script to write.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo funclatency-bpfcc -u -d 5 vfs_fsync_range
Tracing 1 functions for "vfs_fsync_range"... Hit Ctrl-C to end. usecs : count distribution 0 -> 1 : 0 | | 2 -> 3 : 0 | | 4 -> 7 : 0 | | 8 -> 15 : 0 | | 16 -> 31 : 0 | | 32 -> 63 : 38231 |****************************************| 64 -> 127 : 16888 |***************** | 128 -> 255 : 2456 |** | 256 -> 511 : 1157 |* | 512 -> 1023 : 161 | | 1024 -> 2047 : 54 | | 2048 -> 4095 : 21 | | 4096 -> 8191 : 1 | | 8192 -> 16383 : 1 | | 16384 -> 32767 : 1 | | 32768 -> 65535 : 1 | | 65536 -> 131071 : 1 | | avg = 79 usecs, total: 4685037 usecs, count: 58995 Detaching...

Each vfs_fsync_range call took 79 µs on average, most between 32 and 127 µs, over 58995 calls in five seconds, with a few outliers up to 65 to 131 ms: the same shape as the write() histogram, because almost all of a synchronous write's time is the sync. On the Rocky lab the same tool is /usr/share/bcc/tools/funclatency:

deploy@rocky10 · Rocky Linux 10.2
$ sudo /usr/share/bcc/tools/funclatency -u -d 5 vfs_fsync_range
Tracing 1 functions for "vfs_fsync_range"... Hit Ctrl-C to end. usecs : count distribution 0 -> 1 : 0 | | 2 -> 3 : 0 | | 4 -> 7 : 0 | | 8 -> 15 : 0 | | 16 -> 31 : 0 | | 32 -> 63 : 17950 |******************************* | 64 -> 127 : 22646 |****************************************| 128 -> 255 : 4020 |******* | 256 -> 511 : 1228 |** | 512 -> 1023 : 197 | | 1024 -> 2047 : 56 | | 2048 -> 4095 : 19 | | 4096 -> 8191 : 9 | | 8192 -> 16383 : 2 | | 16384 -> 32767 : 3 | | 32768 -> 65535 : 1 | | 65536 -> 131071 : 2 | | 131072 -> 262143 : 1 | | avg = 102 usecs, total: 4748908 usecs, count: 46151 Detaching...

When a bcc tool fails, use its bpftrace version: bpftrace reads kernel types from BTF rather than headers, and its tools are in /usr/sbin on Ubuntu (/usr/share/bpftrace/tools on RHEL). The disk I/O and CPU scheduler lessons of this course used biolatency.bt and runqlat.bt this way.

deploy@web01 · Ubuntu 26.04 LTS
$ ls /usr/sbin/*.bt | xargs -n 1 basename | column -c 100
bashreadline.bt gethostlatency.bt setuids.bt tcplife.bt biolatency-kp.bt killsnoop.bt ssllatency.bt tcpretrans.bt biolatency.bt loads.bt sslsnoop.bt tcpsynbl.bt biosnoop.bt mdflush.bt statsnoop.bt threadsnoop.bt biostacks.bt naptime.bt swapin.bt undump.bt bitesize.bt oomkill.bt syncsnoop.bt vfscount.bt capable.bt opensnoop.bt syscount.bt vfsstat.bt cpuwalk.bt pidpersec.bt tcpaccept.bt writeback.bt dcsnoop.bt runqlat.bt tcpconnect.bt xfsdist.bt execsnoop.bt runqlen.bt tcpdrop.bt
$ sudo systemctl stop bt-sync.scope

Inside a program: USDT and uprobes

Kernel probes cannot see a garbage collector or an allocator, because their work happens without system calls. A Python service that stalls now and then is the example. This program creates objects that refer to themselves, which only Python's cyclic garbage collector can free, and a 1 MiB buffer each time round. Create a directory for it first (mkdir -p ~/bt).

~/bt/app.py
# app.py: builds objects that point at themselves; only the cyclic GC frees them
import time
class Node:
def __init__(self):
self.peer = self
end = time.monotonic() + 600
while time.monotonic() < end:
batch = [Node() for _ in range(50000)]
buf = bytearray(1 << 20)
time.sleep(0.05)
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=bt-app --uid=$USER -p MemoryMax=256M python3 ~/bt/app.py > /dev/null 2>&1 &
$ sudo bpftrace -l "usdt:/usr/bin/python3.14:*"
usdt:/usr/bin/python3.14:python:audit usdt:/usr/bin/python3.14:python:gc__done usdt:/usr/bin/python3.14:python:gc__start usdt:/usr/bin/python3.14:python:import__find__load__done usdt:/usr/bin/python3.14:python:import__find__load__start

Ubuntu's Python 3.14 is built with USDT markers, including gc__start and gc__done around each collection. Each marker has a semaphore, a counter the program checks so that it skips the probe's work when nobody is tracing; bpftrace activates it when it attaches. -p limits the probes to one process, and pgrep -u $USER -x python3 picks your Python and not the sudo process that started it (pgrep -f app.py would match both). How long does each collection pause the program?

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -p $(pgrep -u $USER -x python3) -e 'usdt:/usr/bin/python3.14:python:gc__start { @start[tid] = nsecs; } usdt:/usr/bin/python3.14:python:gc__done /@start[tid]/ { @gc_usecs = hist((nsecs - @start[tid]) / 1000); _ = delete(@start, tid); } interval:s:5 { exit(); }'
… Attached 5 probes … @gc_usecs: [0] 671 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | [1] 33 |@ | [2, 4) 24 |@ | [4, 8) 1 | | [8, 16) 1 | | [16, 32) 2 | | [32, 64) 3 | | [64, 128) 5 | | [128, 256) 22 |@ | [256, 512) 907 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [512, 1K) 183 |@@@@@@@@@@ | [1K, 2K) 29 |@ | [2K, 4K) 13 | | [4K, 8K) 2 | | …

bpftrace 0.25 first prints compiler warnings from building its USDT support (left out); they are harmless. At the end it also prints the @start entry of a collection still running when it stopped (left out). The histogram has two populations: 671 collections finished in under a microsecond, and most of the rest took 256 µs to 1 ms, with a few up to 8 ms. Those are the pauses the service's requests would feel, measured without changing a line of the program. Next, a uprobe on the C library's malloc: which allocation sizes reach it? (On x86_64 the library is /usr/lib/x86_64-linux-gnu/libc.so.6.)

deploy@web01 · Ubuntu 26.04 LTS
$ sudo bpftrace -p $(pgrep -u $USER -x python3) -e 'uprobe:/usr/lib/aarch64-linux-gnu/libc.so.6:malloc { @bytes = hist(arg0); } interval:s:5 { exit(); }'
Attached 2 probes @bytes: [512, 1K) 76 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | [1K, 2K) 0 | | [2K, 4K) 0 | | [4K, 8K) 0 | | [8K, 16K) 0 | | [16K, 32K) 0 | | [32K, 64K) 0 | | [64K, 128K) 0 | | [128K, 256K) 0 | | [256K, 512K) 0 | | [512K, 1M) 0 | | [1M, 2M) 77 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@|
$ sudo systemctl stop bt-app.scope

77 calls asked for about 1 MiB, one per loop: the bytearray. Another 76 were just over 512 bytes; the Node objects never reach malloc at all, because Python serves allocations up to 512 bytes from its own allocator. Each uprobe hit is a breakpoint trap in the traced process, so a uprobe on a function called millions of times a second slows the program noticeably; the same applies to ltrace, with less reliable results (see the process debugging lesson).

What a probe costs, and the limits

A probe's cost is paid on every event it fires on, whether or not you print anything. Measure it. dd with a block size of one byte makes about four million system calls for 2 MB, which makes the per-call cost visible. Run it bare, then with a probe on every system call entry, with kernel.bpf_stats_enabled switched on so the kernel counts the time spent in each BPF program.

deploy@web01 · Ubuntu 26.04 LTS
$ dd if=/dev/zero of=/dev/null bs=1 count=2000000
2000000+0 records in 2000000+0 records out 2000000 bytes (2.0 MB, 1.9 MiB) copied, 1.15213 s, 1.7 MB/s
$ sudo bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @[comm] = count(); } interval:s:30 { exit(); }' > /var/tmp/bt-count.txt 2>&1 &
$ sudo sysctl kernel.bpf_stats_enabled=1
kernel.bpf_stats_enabled = 1
$ dd if=/dev/zero of=/dev/null bs=1 count=2000000
2000000+0 records in 2000000+0 records out 2000000 bytes (2.0 MB, 1.9 MiB) copied, 1.52287 s, 1.3 MB/s
$ sudo bpftool prog show | grep -A3 "name tracepoint_raw"
9457: tracepoint name tracepoint_raw_syscalls_sys_enter_1 tag 9cd6cbbd503054eb gpl run_time_ns 187918576 run_cnt 4012028 …
$ sudo sysctl kernel.bpf_stats_enabled=0
kernel.bpf_stats_enabled = 0

bpftool prog show lists the program bpftrace loaded and, with statistics on, run_cnt (4012028 runs) and run_time_ns (188 ms): about 47 ns per run. dd itself slowed from 1.15 to 1.52 s, about 90 ns per system call, because entering the probe costs more than the program's own instructions; collecting the statistics adds a little too, which is why they are off by default. These are this VM's numbers, on an Apple-silicon host: the cost per event depends on the CPU and on the kernel's mitigations for speculative-execution attacks, and x86_64 servers often pay more per tracepoint, so measure on your own hardware. The probe also fires for every process on the host, not only the one you care about, so estimate from the host's total rate: at about 100 ns, 100,000 system calls a second cost about 1% of a CPU, and a busy host making several million a second pays several tenths of a CPU. After the tracer ends, its map shows who made the calls:

deploy@web01 · Ubuntu 26.04 LTS
$ sort -t: -k2 -n -r /var/tmp/bt-count.txt | head -n 4
@[dd]: 4000223 @[pgrep]: 236945 @[sleep]: 15267 @[bash]: 13443

dd made 4,000,223 of them. pgrep and sleep are the lab's own polling while it waited. Three more limits matter in practice. Tracing needs privilege: kernel.unprivileged_bpf_disabled is 2 on both platforms, so these tools run as root (or with CAP_BPF and CAP_PERFMON), and kernel lockdown under Secure Boot can refuse some helpers. Output has a finite buffer between the kernel and bpftrace, so a one-liner that prints on every event can lose events when they arrive faster than bpftrace reads them; aggregate in maps and print summaries, as every example here does. And probes on kernel functions are tied to one kernel's internals: the same one-liner may need changes after an upgrade, where a tracepoint would not.

Measure before you attach to hot paths
A probe on the scheduler, the network receive path or a lock taken millions of times a second can cost more than the problem you are chasing. Start with tracepoints and counts, keep the action small, bound it with interval or a timeout, and check bpftool prog show with statistics on if you are unsure what a probe costs.

Try this

Start the dd writer again, then answer a new question with fexit, which runs when a function returns and sees its return value as retval along with its arguments: sudo bpftrace -e 'fexit:vfs_fsync_range /comm == "dd"/ { @[retval] = count(); } interval:s:5 { exit(); }'. Predict first: every sync succeeds, so all calls should be counted under 0. Then change @[retval] to @[args.datasync] and explain the result: O_DSYNC asks only for the data and the metadata needed to read it back, so the kernel passes datasync as 1. Stop the writer with sudo systemctl stop bt-sync.scope.

Takeaway

Pick the most stable probe that answers the question, aggregate in the kernel instead of printing every event, and know what the probe costs per event before you put it on a hot path; when a bcc tool does not compile on your kernel, its bpftrace version usually does.

Quick check
01During an incident on an Ubuntu 26.04 server, tcpconnect-bpfcc stops with a clang error about a static assertion in include/linux/fs.h. What does that tell you, and what do you do?
Incorrect — eBPF works on this kernel: other bcc tools and bpftrace run. The failure is in compiling the tool's C code.
Correct — bcc compiles at start-up against headers; bpftrace's version reads BTF instead and runs on the same kernel.
Incorrect — Lockdown refuses some helpers at load time with a permission error; this failed earlier, while compiling.
Incorrect — The compiler read the headers and rejected what it found in them; installing the same headers again changes nothing.
02A probe on raw_syscalls:sys_enter adds about 40 ns to every system call on your hardware. You want to leave it running on a host whose processes together make 200,000 system calls a second. What is the right expectation?
Correct — The probe fires for every process, so the host's total rate counts: 200,000 × 40 ns. Printing less saves copying, not the per-event cost.
Incorrect — The verifier checks safety; the program still runs on every event, and entering the probe costs time too.
Incorrect — Every event runs the program and updates the map; the print at the end is only the copy to user space.
Incorrect — strace stops the process at each call through ptrace, which costs microseconds, not tens of nanoseconds.
03A Python service pauses for a few milliseconds now and then, with no system calls in the pauses. You suspect the garbage collector. Which probe answers it most directly?
Incorrect — A pause inside one thread's collector need not touch futexes, and a count of futex calls does not time the pause.
Incorrect — The collector runs on the CPU; the scheduler sees nothing unusual, and kernel functions cannot see Python's work.
Correct — The interpreter marks each collection itself; timing the pair gives the pause directly, in the kernel.
Incorrect — Python serves small objects from its own allocator, and freeing memory is not allocating it; malloc misses the pause.

Related