A method for performance problems
USE, workload characterisation and the first minute.
This course explains how Linux works underneath and how to get from a symptom ("the server is slow", "the job will not die") to its cause with measurements instead of guesses. It is not a hardening course. The only prerequisite is the Linux essentials course: processes, systemd, the journal and the basics of /proc. The hardening course is optional; it gives background where sysctl settings and security modules come up, and each lesson explains what it needs from them. The advanced security course is not required. This first lesson gives you the method every later lesson uses and the commands for the first minute on a slow server. By the end you will be able to start a bounded test workload safely, name the resource it saturates and the process behind it, and prove that the machine recovered.
The method: from symptom to next step
Every lesson in this course works through the same six steps. The symptom is what someone noticed: an alert, a timeout, a slow page. The layer is the part of the system that could produce that symptom: a CPU, memory, a disk, the network, a process, or a limit someone configured. The measurement is the number that would confirm or rule that layer out, and the tool is the command that reads it. Interpretation compares the value with what is normal for this machine. The next step is either a fix or a move one layer deeper.
Two checklists stop you from skipping a layer. The USE method, from Brendan Gregg, asks three questions of every resource. Utilisation is the average time the resource was busy. Saturation is extra work it cannot service yet, usually waiting in a queue. Errors are the count of failures. A resource can be fully utilised and healthy; saturation is what users feel. Resources include software ones too: a cgroup CPU quota, a task limit or a thread pool can saturate while the hardware is idle.
USE looks at the supply. Workload characterisation looks at the demand, with four questions: who is causing the load, why it is being done, what it is (operations, bytes, request types), and how it changes over time. A disk saturated by a backup that should have run at 2 a.m. needs a schedule, not tuning.
A workload you control
Learn the method on a lab machine with a load you started yourself, so you know the answer before you measure. stress-ng generates CPU, memory or I/O load on demand (sudo apt install stress-ng on Ubuntu; sudo dnf install stress-ng from AppStream on RHEL 10). Run it inside a transient systemd scope: systemd-run --scope puts the command in a cgroup of its own, here named perf-method-load.scope, so you can apply limits to it, find it by name and stop everything in it with one command.
--uid=$USER runs the workload as you, not as root, although sudo is needed to create a system scope. MemoryMax=256M and TasksMax=16 are guard rails in case you mistype a stressor, and --timeout 10m ends the run even if you forget it. Four CPU stressors on this two-vCPU machine ask for twice the CPU it has. The trailing & puts the scope in the background, and the redirect sends its messages to a file instead of your terminal. If sudo asks you for a password, run sudo -v first: a background job cannot answer the prompt.
sudo apt install stress-ng fio inotify-tools gdb ltrace gcc iperf3 iotop-c. On RHEL 10, sudo dnf install finds the same package names in BaseOS and AppStream, except inotify-tools, which is in neither.The first sixty seconds
Brendan Gregg and the Netflix performance team published a ten-command checklist for the first minute on a Linux server: uptime, dmesg | tail, vmstat 1, mpstat -P ALL 1, pidstat 1, iostat -xz 1, free -m, sar -n DEV 1, sar -n TCP,ETCP 1 and top. Together they cover utilisation, saturation and errors for CPU, memory, disks and network, and each takes a second to run. mpstat, pidstat, iostat and sar come from the sysstat package, which Ubuntu Server 26.04 installs by default; on RHEL, sudo dnf install sysstat. The lab let the load run for 45 seconds before starting.
The three load averages cover 1, 5 and 15 minutes. The 1-minute value is well above the 15-minute one, so something started recently and load is rising. Load average counts tasks that are running, waiting for a CPU, or in uninterruptible sleep, so it tells you that there is demand, not which resource it is on; the CPU saturation lesson takes it apart.
Both Ubuntu 26.04 and RHEL 10 set kernel.dmesg_restrict=1, so dmesg needs sudo. journalctl -k reads the same kernel messages from the journal, and members of adm (Ubuntu) or wheel (RHEL) may read it without sudo. You are looking for I/O errors, out-of-memory kills, hung-task warnings and network link resets. --since -1min limits the search to roughly the time the load has been running, and -p warning to warnings and worse: nothing, so the kernel reported no errors while the machine slowed down. Over a whole boot a virtual machine logs a few harmless lines about its virtual hardware. One to know: hrtimer: interrupt took ... ns in a VM usually means the host delayed the virtual CPU, while on a physical server repeated lines like it are worth correlating with latency spikes.
As the essentials course showed, the first line averages everything since boot, so read the rest. r (runnable tasks) holds at 4 on two CPUs: at any moment about half of them are waiting for a CPU, which is CPU saturation. b is zero, us at 100 with id at 0 is full utilisation, and si/so show no swapping. wa is a hint, not CPU use; the lesson "CPU saturation: load, run queue and pressure" explains why.
mpstat splits the same picture per CPU. Both are at 100% user time, so the load is spread across CPUs. One CPU at 100% with the others idle would point at a single-threaded bottleneck, which no amount of extra CPUs fixes.
pidstat 1 names the processes using CPU in that second. Four stress-ng-cpu workers owned by UID 1001 each got about half a CPU, and %wait, the share of time a task was ready to run but waiting for a CPU, is about 50% each. That is saturation measured per process. ps cannot give you this: its %CPU column is CPU time divided by the process's whole lifetime, not what it is using now.
The rest of the sweep rules layers out. iostat -xz prints a report averaged since boot (its device rows are left out above) and then one per interval, and -z hides idle devices: the second report is 100% user time and has no device rows at all. When a disk is involved, aqu-sz (average queue length) is its saturation and r_await/w_await are the latency; %util only means saturation for a device that serves one request at a time, not for SSDs and RAID. free -m shows 3383 MiB available, so memory is fine, and sar -n DEV shows no network traffic. sar -n TCP,ETCP would show retransmissions and resets.
All ten read host-wide counters. On a host that runs services or containers under limits, add an eleventh check for the limits themselves: grep throttled /sys/fs/cgroup/system.slice/*/cpu.stat shows which service has used up its CPU quota, and a unit's memory.events counts how often it reached its memory limit (the resource-control lesson covers both). Inside a container, uptime, vmstat, free and mpstat read the host's /proc, so they report the host's load, memory and CPUs, not the container's limits; read the container's own /sys/fs/cgroup/cpu.stat, memory.current and memory.events instead.
Pressure stall information
Pressure stall information (PSI) measures saturation directly. The essentials course introduced the format: some is the share of time at least one task was stalled on the resource, full the share in which all non-idle tasks were stalled at once (always zero for CPU at system level), with averages over 10, 60 and 300 seconds and a total in microseconds.
some avg10=97.47 says that in the last ten seconds some task was waiting for a CPU almost all the time. What this course adds is the per-cgroup view: every cgroup has its own cpu.pressure, memory.pressure and io.pressure files, and in those the CPU full line does matter, because it says when all of one service's tasks were waiting at once.
The essentials course also showed that RHEL 10 builds PSI into its kernel but turns it off.
CONFIG_PSI_DEFAULT_DISABLED means tracking starts only when the kernel is booted with psi=1. The kernel's own help text says the cost is too small to affect common workloads but shows up in artificial scheduler stress tests. psi=1 is a kernel command-line parameter: one of the options the boot loader passes to the kernel when it starts it, which the boot lesson traces from the boot loader's configuration to /proc/cmdline. On RHEL you add such parameters with grubby, which edits every boot entry; the change takes effect at the next reboot, and --remove-args rolls it back.
psi=1 and removed it again without rebooting. On a real host, schedule the reboot, then confirm with cat /proc/cmdline and cat /proc/pressure/cpu. Until then, RHEL hosts need vmstat and pidstat for saturation, and an alert rule written against /proc/pressure fails there.Who, what and since when
The sweep answered which resource. Workload characterisation answers who. Because the load runs in a named unit, systemd can show all of it at once; on a real host, systemctl status PID finds the unit that owns any process.
Who: a scope started by user 1001 (deploy) running stress-ng --cpu 4, five tasks against a limit of 16. What: 1 minute 49 seconds of CPU time in the 54 seconds since it started, which is two CPUs' worth, and only 17 MiB of memory. Why is the question you answer with the owner of the job; for the code path inside a process, the flame-graph lesson shows where the CPU time goes.
How it changes over time comes from history. sysstat's sysstat-collect.timer records a sample every 10 minutes into a daily file (/var/log/sysstat/saDD on Ubuntu, /var/log/sa/saDD on RHEL), and sar reads it back. Samples come only every 10 minutes, so start one yourself during the load and today's file contains it.
The first command prints the column header (it carries the time of the file's first sample), the second the last three samples and the day's average. runq-sz is the run queue, plist-sz the number of tasks, ldavg-* the load averages and blocked the tasks in I/O wait. The 10:03 row, the sample you started, has four runnable tasks and a 1-minute load of 2.42, where the earlier rows show loads below 1 (the 2 at 10:00 is a momentary count): the problem started between 10:00 and 10:03. sar -q -f /var/log/sysstat/sa26 reads yesterday's file. On RHEL, installing sysstat enables the timer but does not start it, so collection begins only after a reboot unless you start it yourself.
Stop the load and prove the recovery
The next step here is to remove the cause. On a real server it could equally be capping the job, moving it to another time or adding capacity; whichever you choose, measure the same things again afterwards.
Fifteen seconds later the instant readings are back to normal: r is 0 and the CPUs are idle. The averages lag. PSI avg10 has dropped to 20 and keeps falling, avg60 still reads 46, and the 1-minute load average is 1.88, because each is an exponentially decaying average over a longer window. A load average that falls slowly after a fix is not a problem that is still running; check vmstat before you conclude anything.
Stopping the scope sent SIGTERM to every process in it, and stress-ng shut down cleanly after 55.14 seconds. Nothing from the workload is left running, which is the point of starting test load in a unit you can stop by name.
Try this
A CPU quota is a software limit of the kind USE counts as a resource. CPUQuota=100% lets a unit's processes use one CPU's worth of time per 100 ms period, however many CPUs the machine has; once they have used it, the kernel throttles them (holds them off every CPU) until the next period starts. The scheduler lesson measures this and the resource-control lesson covers the settings.
Start the same workload with -p CPUQuota=100% added and the unit named perf-method-quota, wait ten seconds, and run mpstat 1 1 and pidstat 1 1. Predict first: four workers share one CPU's worth of time on two CPUs, so expect about half of the CPU time to be idle while each worker gets roughly 25% of a CPU and shows a %wait around 75%. Then read the scope's cgroup files, a directory per unit under /sys/fs/cgroup, starting with /sys/fs/cgroup/system.slice/perf-method-quota.scope/cpu.stat: nr_throttled counts the periods in which the scope used up its quota, and the scope's cpu.pressure file shows a non-zero full line. Name the saturated resource before you look at the answer: it is the quota, not the CPUs. Finish with sudo systemctl stop perf-method-quota.scope and confirm with pgrep stress-ng that nothing is left.
Takeaway
Before you change anything, name the saturated resource and the workload behind it, with numbers. Keep those numbers: the same commands after the fix are your proof that it worked.