CPU saturation: load, run queue and pressure

Load average, vmstat, mpstat and PSI.

Advanced14 min · lesson 5 of 21

"The load is high" is usually the first thing anyone says about a slow Linux server, and on its own it tells you very little. This lesson takes the load average apart: what the kernel counts, why a load above the CPU count can come with idle CPUs, and why the same load can mean a healthy machine or a saturated one. You will produce each case on purpose and separate them with vmstat, mpstat, pidstat and CPU pressure, so that the next time a load alert fires you know within a minute whether CPU is the problem. The method lesson introduced these tools; the scheduler lesson explained how tasks share a CPU.

What the load average counts

Every five seconds the kernel counts the tasks that are runnable (running or waiting for a CPU) plus the tasks in uninterruptible sleep, state D, across all CPUs. It folds that count into three exponentially decaying averages whose time constants are 1, 5 and 15 minutes; those are the three numbers uptime prints. Four consequences follow from the definition in the kernel source.

It counts tasks, so every thread of a multithreaded program counts separately. It is not divided by the number of CPUs, so a load of 4 is heavy demand for 2 CPUs and light work for 64; always read it next to nproc. It includes D tasks, which use no CPU at all: they wait inside the kernel, usually for block I/O, NFS or a kernel lock (idle kernel threads, state I, are excluded). And it lags: after a sudden change, the 1-minute figure has moved only about 63% of the way to the new level a minute later, so it is a trend, not a reading.

deploy@web01 · Ubuntu 26.04 LTS
$ uptime nproc
09:42:04 up 1:15, 1 user, load average: 0.22, 0.98, 1.20 2

This machine has two CPUs and is idle now; the 5- and 15-minute values still carry earlier work, which is why the direction matters more than any single figure. For an instant view, vmstat splits the count, but not completely. r is the number of runnable tasks. b is procs_blocked from /proc/stat, which the kernel documents as tasks blocked waiting for I/O to complete: a subset of the D tasks. A task in D that waits for a kernel lock or a hung filesystem outside an I/O wait raises the load but not b. ps sees every D task, so it is the check that closes the gap.

Load with idle CPUs: blocked tasks

To produce load that uses no CPU, the lab reuses the trick from the process lifecycle lesson: a small disk made from a file, behind a device-mapper device that is then suspended, so every read sent to it waits. Four dd readers with iflag=direct (bypassing the page cache) are started in one transient service.

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 64M /var/tmp/perf-cpu.img loop=$(sudo losetup --find --show /var/tmp/perf-cpu.img) sudo dmsetup create perf-cpu-slow --table "0 131072 linear $loop 0" sudo dmsetup suspend perf-cpu-slow
$ sudo systemd-run --unit=perf-cpu-readers sh -c 'for i in 1 2 3 4; do dd if=/dev/mapper/perf-cpu-slow of=/dev/null bs=4k count=1 skip=$i iflag=direct & done; wait'
Running as unit: perf-cpu-readers.service; invocation ID: be6446132b72404eaca8cd1e3f49b272

A minute later the load looks alarming.

deploy@web01 · Ubuntu 26.04 LTS
$ uptime
09:43:19 up 1:17, 1 user, load average: 2.95, 1.66, 1.42
$ vmstat 1 3
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 0 4 0 2697900 38844 900248 0 30 782 2750 2247 15 11 5 78 6 0 0 0 4 0 2697900 38844 900248 0 0 0 0 48 45 0 0 0 100 0 0 0 4 0 2697788 38844 900316 0 0 0 0 47 44 0 0 0 100 0 0

A 1-minute load of 2.95 on two CPUs, and rising: the 5-minute value is 1.66. vmstat explains it: r is 0, nothing wants a CPU, and b is 4, four tasks in uninterruptible sleep waiting for I/O. The CPU columns show id 0 and wa 100. wa, iowait, is idle time during which some task that last ran on that CPU was waiting for I/O. It is idle time: these CPUs could run other work at once. The kernel documentation calls iowait unreliable on multi-CPU systems, so treat it as a hint that something waits for I/O, never as CPU use.

deploy@web01 · Ubuntu 26.04 LTS
$ mpstat -P ALL 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:43:21 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle 09:43:22 all 0.50 0.00 0.50 99.01 0.00 0.00 0.00 0.00 0.00 0.00 09:43:22 0 0.98 0.00 0.98 98.04 0.00 0.00 0.00 0.00 0.00 0.00 09:43:22 1 0.00 0.00 0.00 100.00 0.00 0.00 0.00 0.00 0.00 0.00 …

Per CPU the picture is the same: almost no user or system time and 98 to 100% iowait. Pressure stall information (PSI) then gives the verdict directly.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/pressure/cpu /proc/pressure/io
some avg10=0.18 avg60=0.22 avg300=8.02 total=304986776 full avg10=0.00 avg60=0.00 avg300=0.00 total=0 some avg10=98.58 avg60=71.61 avg300=33.95 total=359135097 full avg10=97.16 avg60=71.17 avg300=33.32 total=343968081

The first two lines are /proc/pressure/cpu: some at 0.18 says almost nothing waited for a CPU in the last ten seconds. The next two are /proc/pressure/io: some at 98.58% means at least one task was stalled on I/O nearly all the time, and full at 97.16% means all non-idle tasks were stalled on I/O at once, so the machine made almost no progress. The load is I/O, not CPU. Find the blocked tasks by their state:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ps -eo pid,stat,wchan:20,comm | grep -E '^ *[0-9]+ D'
259198 Dl submit_bio_wait dd 259199 Dl submit_bio_wait dd 259200 Dl submit_bio_wait dd 259201 Dl submit_bio_wait dd

Four dd processes in D; the l marks them multithreaded, because Ubuntu's dd is the uutils version with a helper thread. The wchan column names the kernel function each one waits in; the readers run as root, so reading it needs sudo, as in the lifecycle lesson. Here it is submit_bio_wait: block I/O, which is why vmstat counted them in b. A D task waiting in a lock or filesystem function would show up here and not in b. The next step belongs to the storage layer: which device they wait on (the disk I/O lesson), and where in the kernel they are stuck (/proc/PID/stack, as in the lifecycle lesson). Resume the device, and the reads complete and the readers exit.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo dmsetup resume perf-cpu-slow sleep 1 ps -o pid,stat,comm -C dd
PID STAT COMMAND
$ sudo dmsetup remove perf-cpu-slow sudo losetup --detach "$(sudo losetup --noheadings --output NAME --associated /var/tmp/perf-cpu.img)" rm /var/tmp/perf-cpu.img

The same run queue, two verdicts

CPU saturation means tasks wait for a CPU. Utilisation means CPUs are busy. They often come together, but not always, and system-wide numbers can hide which one you have. Start two CPU-bound workers on this two-CPU machine and look after twenty seconds.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=perf-cpu-spread --uid=$USER -p MemoryMax=256M -p TasksMax=16 \ stress-ng --cpu 2 --timeout 10m > /var/tmp/perf-cpu-spread.log 2>&1 &
$ vmstat 1 3
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 3 0 0 2707844 38192 905460 0 30 780 2735 2245 15 11 5 78 6 0 0 2 0 0 2708248 38192 905460 0 0 0 0 2000 61 100 0 0 0 0 0 2 0 0 2708248 38192 905460 0 0 0 0 2057 152 100 1 0 0 0 0
$ mpstat -P ALL 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:43:46 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle 09:43:47 all 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 09:43:47 0 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 09:43:47 1 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 …
$ cat /proc/pressure/cpu
some avg10=1.85 avg60=0.70 avg300=7.46 total=305479081 full avg10=0.00 avg60=0.00 avg300=0.00 total=0
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:43:47 UID PID %usr %system %guest %wait %CPU CPU Command 09:43:52 1001 259451 99.60 0.00 0.00 0.60 99.60 1 stress-ng-cpu 09:43:52 1001 259452 100.00 0.00 0.00 0.00 100.00 0 stress-ng-cpu …

r is 2 and both CPUs are at 100% user time: fully utilised. But almost nothing waits. Each worker has a CPU to itself, %wait is under 2%, and CPU pressure some is 1.85%, mostly the lab's own short commands queueing behind the workers now and then. The load average would settle at 2, exactly the CPU count, and the machine is healthy. Now stop them and start the same two workers pinned to CPU 1.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-cpu-spread.scope
$ sudo systemd-run --scope --unit=perf-cpu-pinned --uid=$USER -p MemoryMax=256M -p TasksMax=16 \ stress-ng --cpu 2 --taskset 1 --timeout 10m > /var/tmp/perf-cpu-pinned.log 2>&1 &
$ vmstat 1 3
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 3 0 0 2717312 38200 905540 0 30 773 2713 2234 15 11 5 78 6 0 0 3 0 0 2717260 38200 905540 0 0 0 0 1064 552 50 0 50 0 0 0 2 0 0 2717260 38200 905540 0 0 0 0 1040 532 50 0 50 0 0 0

r is still 2 (the 3 in two samples includes the lab's own commands), the load average would settle at the same 2, and vmstat now reports the machine half idle. A dashboard showing 50% CPU would call this fine.

deploy@web01 · Ubuntu 26.04 LTS
$ mpstat -P ALL 1 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:44:24 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %gnice %idle 09:44:25 all 50.25 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 49.75 09:44:25 0 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 100.00 09:44:25 1 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 …
$ cat /proc/pressure/cpu
some avg10=87.96 avg60=30.96 avg300=13.82 total=328705171 full avg10=0.00 avg60=0.00 avg300=0.00 total=0
$ pidstat -u -C stress-ng-cpu 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:44:25 UID PID %usr %system %guest %wait %CPU CPU Command 09:44:30 1001 259635 50.00 0.00 0.00 50.00 50.00 1 stress-ng-cpu 09:44:30 1001 259636 50.00 0.00 0.00 50.00 50.00 1 stress-ng-cpu …

mpstat -P ALL shows the truth: CPU 0 is 100% idle and CPU 1 is at 100%. CPU pressure some is 87.96% and still rising, and each worker spends half its time waiting (%wait 50.00). This is saturation on a half-idle machine, and only the per-CPU view and the pressure figure reveal it.

deploy@web01 · Ubuntu 26.04 LTS
$ taskset -cp $(pgrep -ox stress-ng-cpu)
pid 259635's current affinity list: 1
$ cat /sys/fs/cgroup/system.slice/perf-cpu-pinned.scope/cpu.pressure
some avg10=93.83 avg60=37.16 avg300=9.14 total=28398645 full avg10=0.00 avg60=0.00 avg300=0.00 total=12773

The cause is affinity: the workers may only run on CPU 1. On real servers the same shape comes from taskset in a start script, CPUAffinity= or AllowedCPUs= in a unit, a container's CPU set, or a program with one busy thread doing the work of many. Every cgroup has its own cpu.pressure, which names the stalled workload: here the scope at 93.83%. Its full line, all of its tasks waiting at once, stays near zero because one of the two is almost always running; at system level, full for CPU is always reported as 0. The next step is to find out why the work is confined, as the scheduler lesson shows, and to remove the pinning or give the work more CPUs.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-cpu-pinned.scope

A slow service on an idle host: CPU quotas

One more case produces CPU waiting with the host looking fine: a CPU quota. A unit with CPUQuota= (a container's CPU limit is the same mechanism) may use only so much CPU time per period, 100 ms by default; once it has, the kernel throttles its tasks until the next period, even on idle CPUs. Host load, host PSI and mpstat all look healthy, and only the service's own cgroup shows it. The method lesson's exercise and the scheduler lesson both produced it. Setting a quota enables the cgroup cpu controller for the unit, so its cpu.stat file gains three counters: nr_periods, nr_throttled (periods in which the quota ran out) and throttled_usec (the time its tasks were held off). nr_throttled climbing in step with nr_periods means the service spends most periods at its limit. The cgroup's cpu.pressure says the same where PSI is on; the resource-control lesson covers the fixes.

Load or slowness: finding the CPU demand
Load above nproc, or a slow service
vmstat r and b, ps D count, mpstat -P ALL, PSI
b or D count high, CPUs idle
Blocked in D state
wchan and io pressure: storage, NFS or a lock
r above nproc, all CPUs busy
CPU saturation everywhere
pidstat: who; then cap, move or add CPUs
one CPU busy, cpu some high
Saturation on some CPUs
affinity, a cpuset or one busy thread
CPUs idle, service slow, quota set
Throttled by its cgroup
cpu.stat nr_throttled, cgroup cpu.pressure
r, b and D count near 0 now
The average is catching up
the cause has gone; watch the trend
Each branch is one of the measurements in this lesson or the method lesson.

When PSI is missing, and steal time

Where PSI is off, as on RHEL 10 by default, the same verdicts are available without it. Per task, pidstat -u reports %wait, the share of time a task was runnable but not running, and the second field of /proc/PID/schedstat is the same wait in nanoseconds. For alerting across a fleet, per-task numbers are awkward; two cheaper sources work without psi=1. /proc/schedstat has one line per CPU whose eighth number is the total time tasks waited to run on that CPU, in nanoseconds, and a service with a quota has the cpu.stat throttling counters above. On the host as a whole, compare vmstat's r with the CPU count and check mpstat -P ALL for uneven CPUs.

On a virtual machine, one more column matters: st in vmstat and %steal in mpstat, the time this machine's virtual CPU was ready to run but the hypervisor ran something else. It reads 0 throughout this lab, but that is not a measurement: the guest can only report steal when the hypervisor supplies it (KVM and Xen can), and this lab's hypervisor does not. On arm64 the kernel logs arm-pv: using stolen time PV at boot when steal is available, and neither lab machine has that line. Where steal is reported, a steady non-zero value means the host's physical CPUs are short, and the remedy is at the host or instance level, not inside the guest. Where it is not reported, the same contention shows up only as unexplained latency, so compare with the provider's host metrics.

Load is not a CPU metric
Alert on load average only together with the CPU count, and never treat it as CPU use. A load of twice the CPU count can be an idle machine waiting for a dead NFS server, and a load equal to the CPU count can hide half the workload queueing on one CPU.

Try this

Start a single worker without pinning: sudo systemd-run --scope --unit=perf-cpu-one --uid=$USER -p MemoryMax=256M -p TasksMax=16 stress-ng --cpu 1 --timeout 10m > /var/tmp/perf-cpu-one.log 2>&1 &. Before you measure, predict mpstat -P ALL 1 1 and /proc/pressure/cpu. Expect one CPU at 100% and the other almost idle (the scheduler picks which), and some close to 0 once the earlier tests have decayed: one CPU fully used, nothing waiting. That is a single-threaded program at full speed, and more CPUs would not make it faster. Stop it with sudo systemctl stop perf-cpu-one.scope and delete the log.

Takeaway

Read a load average only next to nproc, vmstat's r and b, the count of D tasks and the per-CPU view; then let CPU pressure (or %wait where PSI is off, and cpu.stat for a service with a quota) decide whether anything is actually waiting for a CPU.

Quick check
01A 2-CPU VM shows a 1-minute load of 3.2. vmstat shows r 0, b 4, id 0 and wa 100, and users say file uploads hang. What is the right first conclusion?
Correct — The load comes from the b column, iowait is idle time, and the next measurement is io pressure and the device the tasks wait on.
Incorrect — r is 0, so nothing is waiting for a CPU. The load average also counts blocked tasks, which is what raised it here.
Incorrect — iowait is time the CPU sat idle while I/O was outstanding. The CPUs could run other work immediately; they are not doing the I/O.
Incorrect — A lagging average would come with r, b and the D-state count near zero now. b is 4 at this moment, so the demand is current.
02A monitoring dashboard shows a 2-CPU server at 50% CPU with a load of 2.0, yet a service on it is slow. mpstat -P ALL shows CPU 0 idle and CPU 1 at 100%, and CPU pressure some is 86. What explains it?
Incorrect — The free capacity is on a CPU the service cannot or does not use. CPU pressure says its tasks spend most of the time waiting.
Incorrect — PSI measures the share of time tasks were stalled; it is not scaled by idle CPUs, and 86 is what the tasks experienced.
Correct — Affinity, a cpuset or a single busy thread keeps the work on one CPU; taskset -cp or the unit's CPU settings show which.
Incorrect — One CPU is idle, so the host is not short of CPUs; the work cannot reach them. More CPUs would sit idle as well.
03Your team wants a CPU saturation alert on RHEL 10 hosts, where PSI is off by default. Someone proposes alerting when the 1-minute load average exceeds the CPU count. Why is that a poor signal?
Incorrect — The Linux load average includes uninterruptible tasks on every distribution; RHEL does not change the definition.
Correct — %wait and the schedstat wait field are time spent runnable but not running, which is what saturation means. For fleet alerts, the per-CPU wait in /proc/schedstat and cpu.stat nr_throttled for limited services are cheaper to collect.
Incorrect — The longer window lags even more; the problem is what the number counts, not how quickly it reacts.
Incorrect — The 5-second interval is how often the average is updated; the value is already a count of tasks, not a sum of samples.

Related