A method for performance problems

USE, workload characterisation and the first minute.

Advanced16 min · lesson 1 of 21

This course explains how Linux works underneath and how to get from a symptom ("the server is slow", "the job will not die") to its cause with measurements instead of guesses. It is not a hardening course. The only prerequisite is the Linux essentials course: processes, systemd, the journal and the basics of /proc. The hardening course is optional; it gives background where sysctl settings and security modules come up, and each lesson explains what it needs from them. The advanced security course is not required. This first lesson gives you the method every later lesson uses and the commands for the first minute on a slow server. By the end you will be able to start a bounded test workload safely, name the resource it saturates and the process behind it, and prove that the machine recovered.

The method: from symptom to next step

Every lesson in this course works through the same six steps. The symptom is what someone noticed: an alert, a timeout, a slow page. The layer is the part of the system that could produce that symptom: a CPU, memory, a disk, the network, a process, or a limit someone configured. The measurement is the number that would confirm or rule that layer out, and the tool is the command that reads it. Interpretation compares the value with what is normal for this machine. The next step is either a fix or a move one layer deeper.

The method every lesson in this course uses
1Symptom
load 2.1 and rising on a 2-CPU server
2Layer
CPU, memory, storage, network, limits
3Measurement
run queue, per-task wait, CPU pressure
4Tool
vmstat, pidstat, /proc/pressure/cpu
5Interpretation
4 runnable tasks for 2 CPUs: saturated
6Next step
stop or cap the job, then measure again
The examples are the ones this lesson measures.

Two checklists stop you from skipping a layer. The USE method, from Brendan Gregg, asks three questions of every resource. Utilisation is the average time the resource was busy. Saturation is extra work it cannot service yet, usually waiting in a queue. Errors are the count of failures. A resource can be fully utilised and healthy; saturation is what users feel. Resources include software ones too: a cgroup CPU quota, a task limit or a thread pool can saturate while the hardware is idle.

USE looks at the supply. Workload characterisation looks at the demand, with four questions: who is causing the load, why it is being done, what it is (operations, bytes, request types), and how it changes over time. A disk saturated by a backup that should have run at 2 a.m. needs a schedule, not tuning.

A workload you control

Learn the method on a lab machine with a load you started yourself, so you know the answer before you measure. stress-ng generates CPU, memory or I/O load on demand (sudo apt install stress-ng on Ubuntu; sudo dnf install stress-ng from AppStream on RHEL 10). Run it inside a transient systemd scope: systemd-run --scope puts the command in a cgroup of its own, here named perf-method-load.scope, so you can apply limits to it, find it by name and stop everything in it with one command.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=perf-method-load --uid=$USER -p MemoryMax=256M -p TasksMax=16 \ stress-ng --cpu 4 --timeout 10m > /var/tmp/perf-method-load.log 2>&1 &
# An interactive shell also prints the job number and PID of the background job here.

--uid=$USER runs the workload as you, not as root, although sudo is needed to create a system scope. MemoryMax=256M and TasksMax=16 are guard rails in case you mistype a stressor, and --timeout 10m ends the run even if you forget it. Four CPU stressors on this two-vCPU machine ask for twice the CPU it has. The trailing & puts the scope in the background, and the redirect sends its messages to a file instead of your terminal. If sudo asks you for a password, run sudo -v first: a background job cannot answer the prompt.

Tools this course uses
A default Ubuntu Server 26.04 install already has sysstat, strace, tcpdump, perf, bpftrace and the bcc tools. The lab machines also have tools that a default install lacks; install them once with sudo apt install stress-ng fio inotify-tools gdb ltrace gcc iperf3 iotop-c. On RHEL 10, sudo dnf install finds the same package names in BaseOS and AppStream, except inotify-tools, which is in neither.

The first sixty seconds

Brendan Gregg and the Netflix performance team published a ten-command checklist for the first minute on a Linux server: uptime, dmesg | tail, vmstat 1, mpstat -P ALL 1, pidstat 1, iostat -xz 1, free -m, sar -n DEV 1, sar -n TCP,ETCP 1 and top. Together they cover utilisation, saturation and errors for CPU, memory, disks and network, and each takes a second to run. mpstat, pidstat, iostat and sar come from the sysstat package, which Ubuntu Server 26.04 installs by default; on RHEL, sudo dnf install sysstat. The lab let the load run for 45 seconds before starting.

deploy@web01 · Ubuntu 26.04 LTS
$ uptime
10:02:53 up 1:36, 1 user, load average: 2.13, 0.66, 0.71

The three load averages cover 1, 5 and 15 minutes. The 1-minute value is well above the 15-minute one, so something started recently and load is rising. Load average counts tasks that are running, waiting for a CPU, or in uninterruptible sleep, so it tells you that there is demand, not which resource it is on; the CPU saturation lesson takes it apart.

deploy@web01 · Ubuntu 26.04 LTS
$ dmesg --level=err,warn
dmesg: read kernel buffer failed: Operation not permitted
$ journalctl -k -p warning --since -1min --no-hostname
-- No entries --

Both Ubuntu 26.04 and RHEL 10 set kernel.dmesg_restrict=1, so dmesg needs sudo. journalctl -k reads the same kernel messages from the journal, and members of adm (Ubuntu) or wheel (RHEL) may read it without sudo. You are looking for I/O errors, out-of-memory kills, hung-task warnings and network link resets. --since -1min limits the search to roughly the time the load has been running, and -p warning to warnings and worse: nothing, so the kernel reported no errors while the machine slowed down. Over a whole boot a virtual machine logs a few harmless lines about its virtual hardware. One to know: hrtimer: interrupt took ... ns in a VM usually means the host delayed the virtual CPU, while on a physical server repeated lines like it are worth correlating with latency spikes.

deploy@web01 · Ubuntu 26.04 LTS
$ vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 5 0 0 2735572 20496 905944 0 37 720 2308 1985 14 11 5 79 5 0 0 4 0 0 2735524 20496 905956 0 0 0 0 2005 996 100 0 0 0 0 0 4 0 0 2735524 20496 905956 0 0 0 0 1705 854 99 1 0 0 0 0 4 0 0 2741252 20496 905956 0 0 0 0 2001 1101 100 1 0 0 0 0 4 0 0 2741252 20496 905956 0 0 0 0 1977 999 100 0 0 0 0 0

As the essentials course showed, the first line averages everything since boot, so read the rest. r (runnable tasks) holds at 4 on two CPUs: at any moment about half of them are waiting for a CPU, which is CPU saturation. b is zero, us at 100 with id at 0 is full utilisation, and si/so show no swapping. wa is a hint, not CPU use; the lesson "CPU saturation: load, run queue and pressure" explains why.

deploy@web01 · Ubuntu 26.04 LTS
$ mpstat -P ALL 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 10:02:57 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle 10:02:58 all 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 10:02:58 0 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 10:02:58 1 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 …

mpstat splits the same picture per CPU. Both are at 100% user time, so the load is spread across CPUs. One CPU at 100% with the others idle would point at a single-threaded bottleneck, which no amount of extra CPUs fixes.

deploy@web01 · Ubuntu 26.04 LTS
$ pidstat 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 10:02:59 UID PID %usr %system %guest %wait %CPU CPU Command 10:03:00 1001 372885 52.43 0.00 0.00 45.63 52.43 0 stress-ng-cpu 10:03:00 1001 372886 47.57 0.00 0.00 51.46 47.57 1 stress-ng-cpu 10:03:00 1001 372887 47.57 0.00 0.00 50.49 47.57 0 stress-ng-cpu 10:03:00 1001 372888 47.57 0.00 0.00 51.46 47.57 1 stress-ng-cpu …

pidstat 1 names the processes using CPU in that second. Four stress-ng-cpu workers owned by UID 1001 each got about half a CPU, and %wait, the share of time a task was ready to run but waiting for a CPU, is about 50% each. That is saturation measured per process. ps cannot give you this: its %CPU column is CPU time divided by the process's whole lifetime, not what it is using now.

deploy@web01 · Ubuntu 26.04 LTS
$ iostat -xz 1 2
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) avg-cpu: %user %nice %system %iowait %steal %idle 10.66 0.32 5.01 4.60 0.00 79.41 … avg-cpu: %user %nice %system %iowait %steal %idle 99.50 0.00 0.50 0.00 0.00 0.00 Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util
# No device rows in the second report: -z hides idle devices.
$ free -m
total used free shared buff/cache available Mem: 3895 512 2669 1 904 3383 Swap: 0 0 0
$ sar -n DEV 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 10:03:01 IFACE rxpck/s txpck/s rxkB/s txkB/s rxcmp/s txcmp/s rxmcst/s %ifutil 10:03:02 lo 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 10:03:02 eth0 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 …

The rest of the sweep rules layers out. iostat -xz prints a report averaged since boot (its device rows are left out above) and then one per interval, and -z hides idle devices: the second report is 100% user time and has no device rows at all. When a disk is involved, aqu-sz (average queue length) is its saturation and r_await/w_await are the latency; %util only means saturation for a device that serves one request at a time, not for SSDs and RAID. free -m shows 3383 MiB available, so memory is fine, and sar -n DEV shows no network traffic. sar -n TCP,ETCP would show retransmissions and resets.

All ten read host-wide counters. On a host that runs services or containers under limits, add an eleventh check for the limits themselves: grep throttled /sys/fs/cgroup/system.slice/*/cpu.stat shows which service has used up its CPU quota, and a unit's memory.events counts how often it reached its memory limit (the resource-control lesson covers both). Inside a container, uptime, vmstat, free and mpstat read the host's /proc, so they report the host's load, memory and CPUs, not the container's limits; read the container's own /sys/fs/cgroup/cpu.stat, memory.current and memory.events instead.

Pressure stall information

Pressure stall information (PSI) measures saturation directly. The essentials course introduced the format: some is the share of time at least one task was stalled on the resource, full the share in which all non-idle tasks were stalled at once (always zero for CPU at system level), with averages over 10, 60 and 300 seconds and a total in microseconds.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/pressure/cpu
some avg10=97.47 avg60=57.94 avg300=17.18 total=401892175 full avg10=0.00 avg60=0.00 avg300=0.00 total=0

some avg10=97.47 says that in the last ten seconds some task was waiting for a CPU almost all the time. What this course adds is the per-cgroup view: every cgroup has its own cpu.pressure, memory.pressure and io.pressure files, and in those the CPU full line does matter, because it says when all of one service's tasks were waiting at once.

The essentials course also showed that RHEL 10 builds PSI into its kernel but turns it off.

deploy@rocky10 · Rocky Linux 10.2
$ cat /proc/pressure/cpu
cat: /proc/pressure/cpu: No such file or directory
$ grep CONFIG_PSI /boot/config-$(uname -r)
CONFIG_PSI=y CONFIG_PSI_DEFAULT_DISABLED=y

CONFIG_PSI_DEFAULT_DISABLED means tracking starts only when the kernel is booted with psi=1. The kernel's own help text says the cost is too small to affect common workloads but shows up in artificial scheduler stress tests. psi=1 is a kernel command-line parameter: one of the options the boot loader passes to the kernel when it starts it, which the boot lesson traces from the boot loader's configuration to /proc/cmdline. On RHEL you add such parameters with grubby, which edits every boot entry; the change takes effect at the next reboot, and --remove-args rolls it back.

deploy@rocky10 · Rocky Linux 10.2
$ sudo grubby --update-kernel=ALL --args=psi=1
$ sudo grubby --info=DEFAULT | grep ^args
args="console=ttyS0,115200n8 no_timer_check crashkernel=2G-4G:256M,4G-64G:320M,64G-:576M psi=1"
$ sudo grubby --update-kernel=ALL --remove-args=psi=1
Until you reboot, nothing changes
The lab added psi=1 and removed it again without rebooting. On a real host, schedule the reboot, then confirm with cat /proc/cmdline and cat /proc/pressure/cpu. Until then, RHEL hosts need vmstat and pidstat for saturation, and an alert rule written against /proc/pressure fails there.

Who, what and since when

The sweep answered which resource. Workload characterisation answers who. Because the load runs in a named unit, systemd can show all of it at once; on a real host, systemctl status PID finds the unit that owns any process.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl status perf-method-load.scope
● perf-method-load.scope - [systemd-run] /usr/bin/stress-ng --cpu 4 --timeout 10m Loaded: loaded (/run/systemd/transient/perf-method-load.scope; transient) Transient: yes Active: active (running) since Sun 2026-09-27 10:02:08 UTC; 54s ago Invocation: bf8f22f3407349f2b1488b073c80275d Tasks: 5 (limit: 16) Memory: 17.2M (max: 256M, available: 238.7M, peak: 17.3M) CPU: 1min 48.988s CGroup: /system.slice/perf-method-load.scope ├─372879 /usr/bin/stress-ng --cpu 4 --timeout 10m ├─372885 stress-ng-cpu "" "" "" "" . ├─372886 stress-ng-cpu "" "" "" "" . ├─372887 stress-ng-cpu "" "" "" "" . └─372888 stress-ng-cpu "" "" "" "" . Sep 27 10:02:08 web01 systemd[1]: Started perf-method-load.scope - [systemd-run] /usr/bin/stress-ng --cpu 4 --timeout 10m. Sep 27 10:02:08 web01 stress-ng[372879]: invoked with '/usr/bin/stress-ng --cpu 4 --timeout 10m' by user 1001 'deploy' Sep 27 10:02:08 web01 stress-ng[372879]: system: 'web01' Linux 7.0.0-34-generic #34-Ubuntu SMP PREEMPT_DYNAMIC Wed Sep 2 14:34:54 UTC 2026 aarch64 Sep 27 10:02:08 web01 stress-ng[372879]: memory (MB): total 3895.85, free 2662.12, shared 1.19, buffer 20.01, swap 0.00, free swap 0.00

Who: a scope started by user 1001 (deploy) running stress-ng --cpu 4, five tasks against a limit of 16. What: 1 minute 49 seconds of CPU time in the 54 seconds since it started, which is two CPUs' worth, and only 17 MiB of memory. Why is the question you answer with the owner of the job; for the code path inside a process, the flame-graph lesson shows where the CPU time goes.

How it changes over time comes from history. sysstat's sysstat-collect.timer records a sample every 10 minutes into a daily file (/var/log/sysstat/saDD on Ubuntu, /var/log/sa/saDD on RHEL), and sar reads it back. Samples come only every 10 minutes, so start one yourself during the load and today's file contains it.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl start sysstat-collect.service
$ sar -q | grep -m1 runq-sz sar -q | tail -4
20:00:06 runq-sz plist-sz ldavg-1 ldavg-5 ldavg-15 blocked 09:50:03 0 166 0.71 1.00 1.20 0 10:00:04 2 159 0.00 0.15 0.63 0 10:03:03 4 165 2.42 0.77 0.75 0 Average: 2 160 1.12 0.95 0.74 0

The first command prints the column header (it carries the time of the file's first sample), the second the last three samples and the day's average. runq-sz is the run queue, plist-sz the number of tasks, ldavg-* the load averages and blocked the tasks in I/O wait. The 10:03 row, the sample you started, has four runnable tasks and a 1-minute load of 2.42, where the earlier rows show loads below 1 (the 2 at 10:00 is a momentary count): the problem started between 10:00 and 10:03. sar -q -f /var/log/sysstat/sa26 reads yesterday's file. On RHEL, installing sysstat enables the timer but does not start it, so collection begins only after a reboot unless you start it yourself.

deploy@rocky10 · Rocky Linux 10.2
$ systemctl is-active sysstat-collect.timer
inactive
$ sudo systemctl enable --now sysstat
$ systemctl is-active sysstat-collect.timer
active

Stop the load and prove the recovery

The next step here is to remove the cause. On a real server it could equally be capping the job, moving it to another time or adding capacity; whichever you choose, measure the same things again afterwards.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-method-load.scope
$ vmstat 1 3
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 0 0 0 2744296 20504 906304 0 37 717 2298 1980 14 11 5 79 5 0 0 0 0 0 2744932 20504 906308 0 0 0 0 111 92 1 0 100 0 0 0 0 0 0 2744932 20504 906308 0 0 0 0 37 34 0 0 100 0 0 0
$ cat /proc/pressure/cpu
some avg10=20.10 avg60=45.63 avg300=16.85 total=402180565 full avg10=0.00 avg60=0.00 avg300=0.00 total=0
$ uptime
10:03:20 up 1:37, 1 user, load average: 1.88, 0.73, 0.74

Fifteen seconds later the instant readings are back to normal: r is 0 and the CPUs are idle. The averages lag. PSI avg10 has dropped to 20 and keeps falling, avg60 still reads 46, and the 1-minute load average is 1.88, because each is an exponentially decaying average over a longer window. A load average that falls slowly after a fix is not a problem that is still running; check vmstat before you conclude anything.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /var/tmp/perf-method-load.log
Running as unit: perf-method-load.scope; invocation ID: bf8f22f3407349f2b1488b073c80275d … stress-ng: info: [372879] dispatching hogs: 4 cpu stress-ng: info: [372879] stopping 4 stressors … stress-ng: info: [372879] successful run completed in 55.14 secs

Stopping the scope sent SIGTERM to every process in it, and stress-ng shut down cleanly after 55.14 seconds. Nothing from the workload is left running, which is the point of starting test load in a unit you can stop by name.

Try this

A CPU quota is a software limit of the kind USE counts as a resource. CPUQuota=100% lets a unit's processes use one CPU's worth of time per 100 ms period, however many CPUs the machine has; once they have used it, the kernel throttles them (holds them off every CPU) until the next period starts. The scheduler lesson measures this and the resource-control lesson covers the settings.

Start the same workload with -p CPUQuota=100% added and the unit named perf-method-quota, wait ten seconds, and run mpstat 1 1 and pidstat 1 1. Predict first: four workers share one CPU's worth of time on two CPUs, so expect about half of the CPU time to be idle while each worker gets roughly 25% of a CPU and shows a %wait around 75%. Then read the scope's cgroup files, a directory per unit under /sys/fs/cgroup, starting with /sys/fs/cgroup/system.slice/perf-method-quota.scope/cpu.stat: nr_throttled counts the periods in which the scope used up its quota, and the scope's cpu.pressure file shows a non-zero full line. Name the saturated resource before you look at the answer: it is the quota, not the CPUs. Finish with sudo systemctl stop perf-method-quota.scope and confirm with pgrep stress-ng that nothing is left.

Takeaway

Before you change anything, name the saturated resource and the workload behind it, with numbers. Keep those numbers: the same commands after the fix are your proof that it worked.

Quick check
01A 4-CPU server runs an API service with CPUQuota=100%. The API is slow. mpstat shows every CPU about 70% idle, and pidstat shows the API's worker processes with a high %wait. Which resource is saturated?
Incorrect — iowait is a subset of idle, but pidstat's %wait is time spent ready to run and waiting for a CPU, not waiting for I/O.
Incorrect — The CPUs are mostly idle, so they cannot be what the workers are queueing for. Something is keeping the workers off idle CPUs.
Correct — Once the cgroup has used its quota for the period, its tasks wait even on idle CPUs. cpu.stat nr_throttled and the cgroup's cpu.pressure confirm it.
Incorrect — A task waiting on a socket is asleep, not runnable, so it adds nothing to %wait, which measures waiting for a CPU.
02You copy a PSI-based alert from your Ubuntu 26.04 servers to RHEL 10 servers. On RHEL the collector fails because /proc/pressure/cpu does not exist. What is the cause?
Correct — CONFIG_PSI_DEFAULT_DISABLED=y leaves tracking off until the kernel boots with psi=1, which grubby can add to every boot entry.
Incorrect — PSI is an upstream kernel feature, and the RHEL kernel config shows CONFIG_PSI=y. It is present but switched off.
Incorrect — The /proc/pressure files exist from boot whenever PSI is enabled, with zeros if nothing has stalled.
Incorrect — An unprivileged user can read /proc/pressure on Ubuntu. A missing file is a different error from a permission error.
03You stop a runaway batch job. Twenty seconds later uptime still shows a 1-minute load average of 2.1 on a 2-CPU host, while vmstat 1 shows r and b at 0 and the CPUs 100% idle. What do you conclude?
Incorrect — Load average also counts waiting and uninterruptible tasks, and vmstat shows none of either right now.
Incorrect — Surviving processes that used CPU would show up as runnable tasks and busy CPUs in vmstat, and both are at zero.
Incorrect — It is computed the same way in a VM. It is slow to react, which is not the same as being wrong.
Correct — The 1-minute figure is an exponentially decaying average, so it keeps falling for a while after the cause is removed.

Related