Disk I/O: latency, queues and saturation
iostat, pidstat -d and biolatency.
When a service is slow and the disk is suspected, three questions decide what to do: how long each request takes, how many requests are waiting, and whether the device is at its limit. This lesson answers them with iostat -x, pressure stall information and a latency histogram, on a test disk whose real capacity is known, so you can check every reading against the truth. You will see why %util cannot tell you whether a disk is saturated, find the processes doing the I/O and how long they wait, and then handle the other storage emergency: a full filesystem where df and du disagree. The previous lesson followed a request through the block layer; this one times it there.
Where disk latency is measured
iostat reads /proc/diskstats, a set of counters per block device that the kernel has kept since boot, and turns the difference between two readings into rates. The time counters start when the block layer allocates a request for a bio and stop when the driver completes it. So r_await and w_await, the average milliseconds per read and write request, include the time a request waits in the I/O scheduler plus the time the device takes to serve it. They do not include the page cache or the filesystem above, which is why a read served from memory never appears in them.
Three more columns complete the picture. r/s and w/s count completed requests per second, and rareq-sz is their average size in KiB. aqu-sz is the average number of requests in the block layer at once, queued or being served. %util is the share of the interval in which at least one request was in flight. These are tied together by Little's law: requests in flight equal completions per second times the time each takes, so aqu-sz is about r/s × r_await / 1000.
A disk whose answers you know
To learn to read these columns you need a disk whose truth you know, as the method lesson used a CPU load you started yourself. null_blk is the kernel's test block driver: it creates a disk that stores nothing and completes each request after a fixed time. Loaded as below, every request takes 10 ms (irqmode=2 completes each request from a timer after completion_nsec, here 10,000,000 ns) and the device has four request slots (hw_queue_depth=4), so its capacity is exactly 400 requests per second. It starts with the scheduler none; the lab switches it to mq-deadline, the scheduler the kernel gives a single-queue disk such as vda. The load generator, fio, is not part of a default Ubuntu Server install: sudo apt install fio (on RHEL 10, sudo dnf install fio).
RQ-SIZE 8 means the scheduler can hold eight requests for this disk, twice its four hardware slots. The workload is fio doing random 4 KiB reads with --direct=1 (bypassing the page cache) and --ioengine=psync: each job issues one synchronous read and waits for it, like a worker thread in a database. It runs as root because it reads the raw device, in a transient scope with a five-minute limit. Start one reader and look after it has settled.
Stop that reader and start four, as many as the device has slots, then read iostat again.
One reader: 85.6 reads per second, each taking 11.60 ms, one request in flight (aqu-sz 0.99) and %util 98.48. Four readers: 360.6 reads per second, four times as many, with about the same latency (10.99 ms) and four requests in flight, and %util 99.94. Both check out with Little's law. The device had spare capacity in the first case and used all of it in the second, and %util could not tell them apart: it only says that the device was never idle.
%util measures the time a device had at least one request in flight. For a device that serves one request at a time, such as a single spinning disk, that is a saturation measure. SSDs, NVMe drives, RAID arrays and most virtual disks serve many requests in parallel and can show 100% while they have room for several times the load. Judge them by latency and queue length.Saturation: when requests queue
Now eight readers, twice what the device can serve at once.
Throughput did not grow, 348.1 reads per second, but r_await doubled to 22.85 ms and aqu-sz rose to 7.95. Four requests are always being served and four wait in the scheduler for a slot, each for about one service time. That is saturation: more demand than the device can serve, visible as a queue longer than its parallelism and a latency above its service time, while throughput stays flat. On a real disk you seldom know the parallelism, so compare await with what the device shows when it is quiet and watch whether throughput still grows when the queue does.
On cloud block storage and in VMs, read that plateau with care. A cloud volume usually has a provisioned limit on IOPS and throughput, sometimes with burst credits that run out, and hitting it looks exactly like this: await rises and throughput stops growing. iostat cannot tell a provider's cap from the device's own parallelism, and in a VM await also includes queueing on the host, shared with other guests. Before tuning queue depths or schedulers in the guest, compare the plateau with the volume's provisioned numbers and its burst balance in the provider's metrics; no guest setting raises a provider cap.
I/O pressure agrees from the tasks' side. some at 96.67 means at least one task waited for I/O almost all of the last ten seconds; full at 94.50 means all non-idle tasks were waiting at once, so little else could run. (PSI is off by default on RHEL, as the method lesson showed.) Next, who is doing the I/O.
pidstat -d shows each process's read and write rate from /proc/PID/io. It needs sudo to see other users' processes, here the eight fio jobs owned by root, each reading about 171 KiB/s, which is 43 reads of 4 KiB per second. iodelay is 0, a problem the next section solves.
An average of 23 ms can hide very different distributions, for example most requests at 2 ms and a few at 500. A latency histogram shows the shape. The bcc tool biolatency-bpfcc does not compile against Ubuntu 26.04's kernel headers with bcc 0.35; the bpftrace version, biolatency.bt, works. On RHEL the bcc tool, /usr/share/bcc/tools/biolatency, works.
biolatency.bt runs until interrupted, so timeout -s INT 10 stops it after ten seconds, and --preserve-status keeps its exit status. It prints warnings about its own script first (left out here); they are harmless. The histogram counts requests by latency in microseconds, in power-of-two buckets, and it covers every block device on the machine. The test disk is the hump of 3273 requests between 16 and 32 ms, queue plus service. The few dozen requests below 8 ms cannot be the test disk, where nothing finishes in under 10 ms. The tool times each request from the moment its bio is queued, so it also includes any time spent waiting for a free request slot. It keys its start times on the sector number alone, though, so it cannot say which device a request went to. The same few lines of bpftrace keyed on the device number as well split the histogram per disk (the device number packs the major number above bit 20):
[253, 0] is vda, the VM's system disk: twenty fast requests from the rest of the system. [251, 0] is nullb0, and every one of its requests took between 8 and 64 ms. That is what a histogram shows and an average hides: two populations, one fast and one slow, which here belong to two devices.
On RHEL, run the bcc tool by its path, sudo /usr/share/bcc/tools/biolatency 5 1; it prints the same kind of histogram in bcc's format.
Who is waiting: delay accounting
The kernel can record, per task, how long it waited for block I/O, swap-in and other resources: delay accounting. pidstat's iodelay column and iotop's IO column come from it. Both Ubuntu 26.04 and RHEL 10 leave it off by default; the delayacct boot option or the kernel.task_delayacct sysctl turns it on.
That is RHEL's value; Ubuntu's reads the same below. The C rewrite of iotop, iotop-c, is packaged on both; on Ubuntu it also provides the iotop command. Switch accounting on and install it:
The eight readers are waiting all the time, and IO still reads 0.00%. The kernel documentation explains it: only tasks started after delay accounting was switched on have delay information. Restart the workload.
Now each new fio job shows 100% in the IO column, the share of time it waited for block I/O, and pidstat's iodelay is no longer 0. pidstat(1) gives that column in clock ticks, 100 per second here, yet it reports more than the 500 ticks a 5-second interval can hold for one process, so read it only as a relative number and use iotop's percentage for the share of time. On a real server this is how you find the process that suffers from a slow disk, which is not always the one that causes the load. Switch accounting off again when you are done; it only costs a little, but it is not the platform default.
Stop the load and confirm recovery
Fifteen seconds after the stop the device is idle. PSI's ten-second average has fallen to 22.63 while the 60-second one still reads 64.01: averages lag, as in the CPU lessons. fio's own report is the application's view and agrees with the block layer's: 350 reads per second at an average of 22.8 ms (22841 µs), with the median at 22.7 ms and the 99th percentile at 30.3 ms. Remove the test device; the exercise at the end loads it again.
When df and du disagree
The other storage emergency is a full filesystem where df and du disagree. The files lesson showed the mechanism: a deleted file that a process still holds open keeps its blocks until the last descriptor closes. Here is how it looks from the disk's side, on a small loop filesystem with a writer that behaves like a logging daemon: it appends to app.log through standard output, writes 40 MiB, then adds a line every second.
df asks the filesystem how many blocks are in use: 41M. du adds up the files it can reach by name: 20K. The difference is a file with no name left. lsof +L1 lists open files with a link count below one.
(lsof first warns that it cannot inspect the tracing filesystem as a normal user; that line is left out.) NLINK 0 and (deleted) mark the file, 41943098 bytes, held on descriptor 1 by the shell and by the sleep it started, because a child inherits its parent's descriptors. The clean fix is to make the writer close and reopen its file: restart it, or send the signal it uses for that (for rsyslog, systemctl kill -s HUP rsyslog, which is what logrotate does). When you cannot do that right now, truncate the file through the descriptor.
The space came back at once and the writer carried on: three seconds later the file holds 87 bytes, three new lines. That works here because the shell opened the log with >> (O_APPEND); a writer without O_APPEND would carry on at its old offset and leave a sparse file of the old size, as the files lesson showed with copytruncate. Truncation also destroys the log's contents, so copy anything you need from /proc/PID/fd/N first. Finally remove the test filesystem:
Try this
Load the test disk again with the same modprobe command, but leave out the switch to mq-deadline, and check lsblk -t /dev/nullb0: this time it keeps its default scheduler, none. Start the eight readers again, wait 15 seconds, and predict iostat -dxy nullb0 5 1 and /proc/pressure/io before you run them. Expect RQ-SIZE 4, and iostat to show about 358 reads per second at 11.1 ms with aqu-sz near 4, exactly like four readers, while some pressure stays above 80 and fio reports about 22.4 ms per read. With no scheduler the queue holds only as many requests as the device has slots; the other four readers wait for a free request before the block layer starts timing them, so iostat cannot see their wait. Stop the scope and remove the device with sudo modprobe -r null_blk.
Takeaway
Judge a disk by its latency compared with its quiet-time latency and by whether throughput still grows when the queue grows; %util only tells you it was never idle, and PSI or the application's own timings tell you when iostat is missing part of the wait.