CPU saturation: load, run queue and pressure
Load average, vmstat, mpstat and PSI.
"The load is high" is usually the first thing anyone says about a slow Linux server, and on its own it tells you very little. This lesson takes the load average apart: what the kernel counts, why a load above the CPU count can come with idle CPUs, and why the same load can mean a healthy machine or a saturated one. You will produce each case on purpose and separate them with vmstat, mpstat, pidstat and CPU pressure, so that the next time a load alert fires you know within a minute whether CPU is the problem. The method lesson introduced these tools; the scheduler lesson explained how tasks share a CPU.
What the load average counts
Every five seconds the kernel counts the tasks that are runnable (running or waiting for a CPU) plus the tasks in uninterruptible sleep, state D, across all CPUs. It folds that count into three exponentially decaying averages whose time constants are 1, 5 and 15 minutes; those are the three numbers uptime prints. Four consequences follow from the definition in the kernel source.
It counts tasks, so every thread of a multithreaded program counts separately. It is not divided by the number of CPUs, so a load of 4 is heavy demand for 2 CPUs and light work for 64; always read it next to nproc. It includes D tasks, which use no CPU at all: they wait inside the kernel, usually for block I/O, NFS or a kernel lock (idle kernel threads, state I, are excluded). And it lags: after a sudden change, the 1-minute figure has moved only about 63% of the way to the new level a minute later, so it is a trend, not a reading.
This machine has two CPUs and is idle now; the 5- and 15-minute values still carry earlier work, which is why the direction matters more than any single figure. For an instant view, vmstat splits the count, but not completely. r is the number of runnable tasks. b is procs_blocked from /proc/stat, which the kernel documents as tasks blocked waiting for I/O to complete: a subset of the D tasks. A task in D that waits for a kernel lock or a hung filesystem outside an I/O wait raises the load but not b. ps sees every D task, so it is the check that closes the gap.
Load with idle CPUs: blocked tasks
To produce load that uses no CPU, the lab reuses the trick from the process lifecycle lesson: a small disk made from a file, behind a device-mapper device that is then suspended, so every read sent to it waits. Four dd readers with iflag=direct (bypassing the page cache) are started in one transient service.
A minute later the load looks alarming.
A 1-minute load of 2.95 on two CPUs, and rising: the 5-minute value is 1.66. vmstat explains it: r is 0, nothing wants a CPU, and b is 4, four tasks in uninterruptible sleep waiting for I/O. The CPU columns show id 0 and wa 100. wa, iowait, is idle time during which some task that last ran on that CPU was waiting for I/O. It is idle time: these CPUs could run other work at once. The kernel documentation calls iowait unreliable on multi-CPU systems, so treat it as a hint that something waits for I/O, never as CPU use.
Per CPU the picture is the same: almost no user or system time and 98 to 100% iowait. Pressure stall information (PSI) then gives the verdict directly.
The first two lines are /proc/pressure/cpu: some at 0.18 says almost nothing waited for a CPU in the last ten seconds. The next two are /proc/pressure/io: some at 98.58% means at least one task was stalled on I/O nearly all the time, and full at 97.16% means all non-idle tasks were stalled on I/O at once, so the machine made almost no progress. The load is I/O, not CPU. Find the blocked tasks by their state:
Four dd processes in D; the l marks them multithreaded, because Ubuntu's dd is the uutils version with a helper thread. The wchan column names the kernel function each one waits in; the readers run as root, so reading it needs sudo, as in the lifecycle lesson. Here it is submit_bio_wait: block I/O, which is why vmstat counted them in b. A D task waiting in a lock or filesystem function would show up here and not in b. The next step belongs to the storage layer: which device they wait on (the disk I/O lesson), and where in the kernel they are stuck (/proc/PID/stack, as in the lifecycle lesson). Resume the device, and the reads complete and the readers exit.
The same run queue, two verdicts
CPU saturation means tasks wait for a CPU. Utilisation means CPUs are busy. They often come together, but not always, and system-wide numbers can hide which one you have. Start two CPU-bound workers on this two-CPU machine and look after twenty seconds.
r is 2 and both CPUs are at 100% user time: fully utilised. But almost nothing waits. Each worker has a CPU to itself, %wait is under 2%, and CPU pressure some is 1.85%, mostly the lab's own short commands queueing behind the workers now and then. The load average would settle at 2, exactly the CPU count, and the machine is healthy. Now stop them and start the same two workers pinned to CPU 1.
r is still 2 (the 3 in two samples includes the lab's own commands), the load average would settle at the same 2, and vmstat now reports the machine half idle. A dashboard showing 50% CPU would call this fine.
mpstat -P ALL shows the truth: CPU 0 is 100% idle and CPU 1 is at 100%. CPU pressure some is 87.96% and still rising, and each worker spends half its time waiting (%wait 50.00). This is saturation on a half-idle machine, and only the per-CPU view and the pressure figure reveal it.
The cause is affinity: the workers may only run on CPU 1. On real servers the same shape comes from taskset in a start script, CPUAffinity= or AllowedCPUs= in a unit, a container's CPU set, or a program with one busy thread doing the work of many. Every cgroup has its own cpu.pressure, which names the stalled workload: here the scope at 93.83%. Its full line, all of its tasks waiting at once, stays near zero because one of the two is almost always running; at system level, full for CPU is always reported as 0. The next step is to find out why the work is confined, as the scheduler lesson shows, and to remove the pinning or give the work more CPUs.
A slow service on an idle host: CPU quotas
One more case produces CPU waiting with the host looking fine: a CPU quota. A unit with CPUQuota= (a container's CPU limit is the same mechanism) may use only so much CPU time per period, 100 ms by default; once it has, the kernel throttles its tasks until the next period, even on idle CPUs. Host load, host PSI and mpstat all look healthy, and only the service's own cgroup shows it. The method lesson's exercise and the scheduler lesson both produced it. Setting a quota enables the cgroup cpu controller for the unit, so its cpu.stat file gains three counters: nr_periods, nr_throttled (periods in which the quota ran out) and throttled_usec (the time its tasks were held off). nr_throttled climbing in step with nr_periods means the service spends most periods at its limit. The cgroup's cpu.pressure says the same where PSI is on; the resource-control lesson covers the fixes.
When PSI is missing, and steal time
Where PSI is off, as on RHEL 10 by default, the same verdicts are available without it. Per task, pidstat -u reports %wait, the share of time a task was runnable but not running, and the second field of /proc/PID/schedstat is the same wait in nanoseconds. For alerting across a fleet, per-task numbers are awkward; two cheaper sources work without psi=1. /proc/schedstat has one line per CPU whose eighth number is the total time tasks waited to run on that CPU, in nanoseconds, and a service with a quota has the cpu.stat throttling counters above. On the host as a whole, compare vmstat's r with the CPU count and check mpstat -P ALL for uneven CPUs.
On a virtual machine, one more column matters: st in vmstat and %steal in mpstat, the time this machine's virtual CPU was ready to run but the hypervisor ran something else. It reads 0 throughout this lab, but that is not a measurement: the guest can only report steal when the hypervisor supplies it (KVM and Xen can), and this lab's hypervisor does not. On arm64 the kernel logs arm-pv: using stolen time PV at boot when steal is available, and neither lab machine has that line. Where steal is reported, a steady non-zero value means the host's physical CPUs are short, and the remedy is at the host or instance level, not inside the guest. Where it is not reported, the same contention shows up only as unexplained latency, so compare with the provider's host metrics.
Try this
Start a single worker without pinning: sudo systemd-run --scope --unit=perf-cpu-one --uid=$USER -p MemoryMax=256M -p TasksMax=16 stress-ng --cpu 1 --timeout 10m > /var/tmp/perf-cpu-one.log 2>&1 &. Before you measure, predict mpstat -P ALL 1 1 and /proc/pressure/cpu. Expect one CPU at 100% and the other almost idle (the scheduler picks which), and some close to 0 once the earlier tests have decayed: one CPU fully used, nothing waiting. That is a single-threaded program at full speed, and more CPUs would not make it faster. Stop it with sudo systemctl stop perf-cpu-one.scope and delete the log.
Takeaway
Read a load average only next to nproc, vmstat's r and b, the count of D tasks and the per-CPU view; then let CPU pressure (or %wait where PSI is off, and cpu.stat for a service with a quota) decide whether anything is actually waiting for a CPU.