perf: sampling, flame graphs and off-CPU time
Profile CPU use and read a flame graph.
strace shows what a process asks the kernel for; it cannot tell you which of the program's own functions burn the CPU. perf can. This lesson profiles a small program with a known hot path, reads the result as a report and as a flame graph, shows the ways a profile goes wrong (missing frame pointers, missing symbols, too few samples, interpreted code), and then measures the other half of latency, the time a program spends waiting rather than running. The method is the one this course uses throughout: confirm the symptom, profile, read the picture, change one thing, profile again.
Sampling, and who may use perf
perf is a sampling profiler. It asks the kernel to interrupt the program at a fixed rate, and at every interrupt it records where the program was: the instruction address and, with -g, the whole call stack, the chain of functions that led there. After a few thousand samples, the share of samples that land in a function estimates the share of CPU time spent in it. -F 99 samples 99 times a second per CPU; the odd number avoids running in step with timers that fire at round intervals. Sampling at -F 99 with frame-pointer stacks costs little, because nothing happens between samples, which is why perf can be used on busy systems where strace cannot. Other modes cost more: DWARF call graphs (later in this lesson) copy part of the stack with every sample, and a system-wide recording (-a) grows until you stop it, so bound it (-- sleep 30) and write to a filesystem with room. Reports (perf report, perf script) are CPU-heavy too; on a production host, record there and analyse elsewhere.
On Ubuntu 26.04, perf comes from the linux-perf package, installed by default (version 7.0.14). kernel.perf_event_paranoid decides what unprivileged users may measure, and Ubuntu sets it to 4, a level Ubuntu's kernels add to the upstream scale (which ends at 2) and which blocks unprivileged use of perf entirely. So profiling needs sudo on Ubuntu, as the first recording below shows. On RHEL 10 the value is 2 and sudo dnf install perf installs version 6.12; there a user may profile their own processes in user space without sudo.
Profile a program with a known hot path
The workload is a small C program whose answer you know in advance: for every request, handle_requests calls parse_request once and checksum, which walks the buffer three times, so checksum should dominate. noinline keeps the functions separate so they appear in the profile.
/* fg-demo.c: a CPU-bound program with a known hot path.main -> handle_requests -> parse_request, then checksum for every request. */#include <stdio.h>#include <stdint.h>__attribute__((noinline)) static uint64_t checksum(const unsigned char *buf, int n){uint64_t h = 1469598103934665603ULL; /* FNV-1a, three passes */for (int r = 0; r < 3; r++)for (int i = 0; i < n; i++)h = (h ^ buf[i]) * 1099511628211ULL;return h;}__attribute__((noinline)) static uint64_t parse_request(unsigned char *buf, int n, uint64_t seed){for (int i = 0; i < n; i++)buf[i] = (unsigned char)((seed >> (i & 31)) + i);return seed * 6364136223846793005ULL + 1;}__attribute__((noinline)) static uint64_t handle_requests(long count){static unsigned char buf[4096];uint64_t seed = 42, total = 0;for (long k = 0; k < count; k++) {seed = parse_request(buf, sizeof buf, seed);total += checksum(buf, sizeof buf);}return total;}int main(void){uint64_t total = handle_requests(300000);printf("%llu\n", (unsigned long long)total);return 0;}
Compile it with optimisation (-O2), debug information (-g, for names and line numbers) and -fno-omit-frame-pointer, which the section on wrong profiles explains. gcc is not installed on Ubuntu Server by default (sudo apt install gcc). time shows the symptom: about five seconds, all of it user CPU time. Recording it as an ordinary user fails on Ubuntu:
The message's advice to lower the setting in /etc/sysctl.conf is generic upstream text (Ubuntu 26.04 has no such file) and a host-wide change that the hardening course (optional) discusses; for a one-off profile, use sudo.
perf record wrote 468 samples to fg-demo.data (-o names the file; the default is perf.data). The event is task-clock, a software timer, because this virtual machine exposes no hardware counters; on bare metal, and on cloud instances that expose a virtual PMU (performance monitoring unit), perf samples CPU cycles instead. perf report --no-children ranks functions by self time, the samples in which the function itself was running, and the call graph under each entry shows the path that led there, read from the hot function down to _start. checksum has 92% of the samples. The printf (inlined) frame is not a real caller: Ubuntu's gcc enables _FORTIFY_SOURCE, which makes printf an inline wrapper, and perf attributes part of main to it. constprop.0 marks a copy of the function that gcc specialised for constant arguments.
With --children, the first column is the time spent in a function and everything it called, and the second the self time. main and handle_requests cover 100% of the samples but do no work themselves; checksum does its own. That distinction is what a flame graph draws.
From stacks to a flame graph
A flame graph shows every sampled stack at once. Brendan Gregg's FlameGraph scripts do it in two steps: stackcollapse-perf.pl folds each stack into one line of function names joined by semicolons, followed by its count, and flamegraph.pl draws those lines as an SVG. Pin the scripts to a known commit so the output does not change under you. On a production server, run only perf script there and fold and draw on your workstation, instead of cloning tools onto the server.
This program has only two distinct stacks. The count at the end of each folded line is not a number of samples but the sum of their sampling periods, in nanoseconds of CPU time for task-clock, so the widths are in proportion to CPU time either way. checksum accounts for 4,333,333,290 of 4,727,272,680 (91.67%). The <title> elements are the tooltips you see when you point at a box in a browser: checksum has 91.67%, and every frame below the two leaves covers 100%.
Read width first: the widest boxes at the top of the graph are the functions doing the work, and a box's width includes everything stacked above it. The x-axis is sorted alphabetically so that identical stacks merge into one box; it says nothing about when something ran. perf itself can also produce a flame graph with perf script report flamegraph, but not with Ubuntu's build: its linux-perf package ships without perf's report scripts. RHEL's perf has them.
On RHEL an unprivileged user recorded their own program (perf_event_paranoid 2), and perf script -l lists the flamegraph report. Its default output is an interactive flamegraph.html built from the d3-flame-graph template, which it asks to download from cdn.jsdelivr.net unless you pass --allow-download or --template; on a host without internet access the command waits there, so the lab bounded it with timeout and asked for JSON (--format json writes stacks.json). The FlameGraph scripts need nothing but Perl.
When the profile is wrong
A flame graph is only as good as the stacks behind it, and it rarely looks broken when it is wrong. perf walks a user-space stack by following frame pointers: one register, the frame pointer, holds the address of the current function's frame, and each function saves its caller's frame pointer (with the return address) on the stack when it starts. Following those saved values leads from the running function back to _start. Compilers may use the register for other work (-fomit-frame-pointer, which GCC's documentation lists as enabled from -O1), and then the chain breaks. Since Ubuntu 24.04, Ubuntu builds its packages with frame pointers so that profiles of system libraries work; your own builds and third-party binaries decide for themselves. On this aarch64 system plain -O2 still kept the frame records, so the lab asks for the omission explicitly; on x86_64 plain -O2 is enough to lose them.
Without frame pointers the stacks lose frames and still look plausible: main and __libc_start_call_main are gone, and handle_requests appears to be called straight from the C library's startup code. (The third line is a single sample taken while the CPU was handling an interrupt, so kernel functions sit on top of checksum.) --call-graph dwarf is the fix when you cannot rebuild: perf copies a slice of the stack with every sample (8 KiB by default) and unwinds it afterwards with the debug information, which recovered main here, at the cost of a file more than fifty times larger (4.1 MB against 76 KB). It needs unwind information for every frame, and the three C library startup frames it could not unwind show as [unknown]. The copied stack memory can contain secrets, so treat a DWARF profile from production like a core dump.
RHEL 10 makes the other choice: its packages are built without frame pointers, which the build flags recorded in each package show.
-O2 and -fasynchronous-unwind-tables are there; -fno-omit-frame-pointer is not. On x86_64 RHEL, where -O2 drops the frame pointer, perf record -g of a service that spends time in glibc, OpenSSL or the Python library loses frames in exactly the plausible-looking way shown above. The unwind tables are what --call-graph dwarf needs, so use it there, or --call-graph lbr on bare-metal Intel CPUs, which record recent branches in hardware. The lab's aarch64 Rocky VM profiles only the course's own program, so it could not show the broken case.
A binary without a symbol table (built with -s here) leaves perf only raw addresses: 92% of the time in 0x928, which tells you nothing. The second result is a surprise: a stripped copy of the binary already profiled still shows names, because perf keeps a copy of every binary it records in the recording user's ~/.debug (root's, with sudo), indexed by build ID, and strip does not change the build ID. Five samples a second over the whole run gives only 21 samples, so each sample is almost 5% of the profile: parse_request is a single sample, and the split of 95/5 cannot tell a function with 3% of the time from one with 8%. Even hundreds of samples move by a few points between runs on this shared VM (92% and 95% for checksum above, from the same program). Aim for hundreds of samples, by running longer rather than sampling faster, and treat a few points of difference as noise until a repeat confirms it.
Interpreted and JIT-compiled code is the last trap. The same hot path in Python:
# fg-demo.py: the same hot path in Python.def checksum(buf):h = 1469598103934665603for b in buf:h = ((h ^ b) * 1099511628211) & 0xFFFFFFFFFFFFFFFFreturn hdef handle_requests(count):total = 0for k in range(count):total += checksum(bytes((k + i) & 255 for i in range(4096)))return totalprint(handle_requests(6000))
perf sees the Python interpreter's C functions, never the Python functions it is running, so the first report has no py:: line at all. Since Python 3.12, python3 -X perf makes the interpreter publish a small trampoline for every Python function and a map file perf reads, and the functions appear: py::checksum is on the stack in 68% of the samples, and the generator expression that builds each request in 22%. -X perf gives complete stacks only when the interpreter itself keeps frame pointers; for interpreters built without them (RHEL's, per the flags above), Python 3.13 added -X perf_jit, which works with --call-graph dwarf (RHEL 10's default python3 is 3.12). Java and Node.js need the same kind of help (a perf map agent, or --perf-basic-prof).
Missing library symbols have one more remedy. Ubuntu sets DEBUGINFOD_URLS in login shells so that tools such as perf and gdb can download debug information from Ubuntu's debuginfod server, but sudo resets the environment, and Ubuntu's sudo-rs ignores -E. Pass the one variable on instead.
Off-CPU time: where the waiting goes
A CPU profile only sees a program while it runs. The next program is slow for a different reason: it waits 20 ms for a backend before each short piece of work.
/* fg-wait.c: a worker that spends most of its time waiting for a slow backend. */#include <stdio.h>#include <stdint.h>#include <time.h>__attribute__((noinline)) static void wait_for_backend(void){struct timespec ts = { 0, 20 * 1000 * 1000 }; /* the backend answers in 20 ms */nanosleep(&ts, NULL);}__attribute__((noinline)) static uint64_t render(uint64_t x){for (int i = 0; i < 2000000; i++)x = x * 6364136223846793005ULL + 1;return x;}int main(void){uint64_t x = 1;for (int k = 0; k < 300; k++) {wait_for_backend();x = render(x);}printf("%llu\n", (unsigned long long)x);return 0;}
8.2 seconds of wall-clock time, 0.9 seconds of CPU. The CPU profile is correct and useless: 84 samples, almost all in render, because a sleeping thread is not running when the samples are taken. Off-CPU analysis records the opposite: every time the scheduler switches the thread out, how long it stayed off the CPU, and the stack at that moment. offcputime-bpfcc from the bcc tools does this in the kernel with eBPF; -p selects the process, -U keeps user-space stacks only, -f prints folded output for flamegraph.pl, and the last argument is the duration in seconds. Its probe fires on every context switch on the host, and -p filters only after it has fired; the tool's man page warns that scheduler events "can exceed 1 million events per second" and asks you to test before production use. On a busy host keep the duration short, and add -m with a minimum block time in microseconds (-m 1000) to skip short waits.
Of the three seconds measured, 2.64 were spent blocked in nanosleep called from wait_for_backend; the count is microseconds, so --countname=us labels the SVG correctly and --color=io uses the blue palette conventional for off-CPU graphs. The fix for this program is not in the CPU at all but in the backend it waits for, or in waiting for several requests at once. On Ubuntu 26.04's kernel several bcc 0.35 tools fail to compile but offcputime-bpfcc works; RHEL installs it as /usr/share/bcc/tools/offcputime. The eBPF tools lesson covers these tools and their bpftrace equivalents.
Try this
Predict before you measure: change checksum to one pass instead of three, rebuild, and record again. With three passes, checksum had 11 to 19 times the weight of parse_request in the lab's runs (92% and 95%), so one pass should give it a third of that, four to six times, a share between 79% and 86%. The lab created the edited copy with sed, so the original stays unchanged.
The two folded counts give the answer: checksum 85% and parse_request 15%, inside the predicted range, and 167 samples instead of 468 show the program finished in about a third of the time. That is the loop to keep: profile, change one thing, profile again, and compare the widths.
Takeaway
Before you trust a flame graph, check that its stacks are whole and named and that it holds hundreds of samples; then read width, not height, and when a program is slow but not busy, measure off-CPU time instead.