Disk I/O: latency, queues and saturation

iostat, pidstat -d and biolatency.

Advanced14 min · lesson 14 of 21

When a service is slow and the disk is suspected, three questions decide what to do: how long each request takes, how many requests are waiting, and whether the device is at its limit. This lesson answers them with iostat -x, pressure stall information and a latency histogram, on a test disk whose real capacity is known, so you can check every reading against the truth. You will see why %util cannot tell you whether a disk is saturated, find the processes doing the I/O and how long they wait, and then handle the other storage emergency: a full filesystem where df and du disagree. The previous lesson followed a request through the block layer; this one times it there.

Where disk latency is measured

iostat reads /proc/diskstats, a set of counters per block device that the kernel has kept since boot, and turns the difference between two readings into rates. The time counters start when the block layer allocates a request for a bio and stop when the driver completes it. So r_await and w_await, the average milliseconds per read and write request, include the time a request waits in the I/O scheduler plus the time the device takes to serve it. They do not include the page cache or the filesystem above, which is why a read served from memory never appears in them.

Three more columns complete the picture. r/s and w/s count completed requests per second, and rareq-sz is their average size in KiB. aqu-sz is the average number of requests in the block layer at once, queued or being served. %util is the share of the interval in which at least one request was in flight. These are tied together by Little's law: requests in flight equal completions per second times the time each takes, so aqu-sz is about r/s × r_await / 1000.

A disk whose answers you know

To learn to read these columns you need a disk whose truth you know, as the method lesson used a CPU load you started yourself. null_blk is the kernel's test block driver: it creates a disk that stores nothing and completes each request after a fixed time. Loaded as below, every request takes 10 ms (irqmode=2 completes each request from a timer after completion_nsec, here 10,000,000 ns) and the device has four request slots (hw_queue_depth=4), so its capacity is exactly 400 requests per second. It starts with the scheduler none; the lab switches it to mq-deadline, the scheduler the kernel gives a single-queue disk such as vda. The load generator, fio, is not part of a default Ubuntu Server install: sudo apt install fio (on RHEL 10, sudo dnf install fio).

deploy@web01 · Ubuntu 26.04 LTS
$ sudo modprobe null_blk nr_devices=1 gb=1 irqmode=2 completion_nsec=10000000 hw_queue_depth=4 echo mq-deadline | sudo tee /sys/block/nullb0/queue/scheduler
mq-deadline
$ lsblk -t /dev/nullb0
NAME ALIGNMENT MIN-IO OPT-IO PHY-SEC LOG-SEC ROTA SCHED RQ-SIZE RA WSAME nullb0 0 512 0 512 512 0 mq-deadline 8 128 0B

RQ-SIZE 8 means the scheduler can hold eight requests for this disk, twice its four hardware slots. The workload is fio doing random 4 KiB reads with --direct=1 (bypassing the page cache) and --ioengine=psync: each job issues one synchronous read and waits for it, like a worker thread in a database. It runs as root because it reads the raw device, in a transient scope with a five-minute limit. Start one reader and look after it has settled.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --scope --unit=perf-io-load -p MemoryMax=256M -p TasksMax=32 \ fio --name=reader --filename=/dev/nullb0 --rw=randread --bs=4k --direct=1 --ioengine=psync \ --numjobs=1 --time_based --runtime=5m --group_reporting > /var/tmp/perf-io-fio.log 2>&1 &
$ iostat -dxy nullb0 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util nullb0 85.60 342.40 0.00 0.00 11.60 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.99 98.48

Stop that reader and start four, as many as the device has slots, then read iostat again.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-io-load.scope
$ sudo systemd-run --scope --unit=perf-io-load -p MemoryMax=256M -p TasksMax=32 \ fio --name=reader --filename=/dev/nullb0 --rw=randread --bs=4k --direct=1 --ioengine=psync \ --numjobs=4 --time_based --runtime=5m --group_reporting > /var/tmp/perf-io-fio.log 2>&1 &
$ iostat -dxy nullb0 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util nullb0 360.60 1442.40 0.00 0.00 10.99 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 3.96 99.94

One reader: 85.6 reads per second, each taking 11.60 ms, one request in flight (aqu-sz 0.99) and %util 98.48. Four readers: 360.6 reads per second, four times as many, with about the same latency (10.99 ms) and four requests in flight, and %util 99.94. Both check out with Little's law. The device had spare capacity in the first case and used all of it in the second, and %util could not tell them apart: it only says that the device was never idle.

%util is not saturation
%util measures the time a device had at least one request in flight. For a device that serves one request at a time, such as a single spinning disk, that is a saturation measure. SSDs, NVMe drives, RAID arrays and most virtual disks serve many requests in parallel and can show 100% while they have room for several times the load. Judge them by latency and queue length.

Saturation: when requests queue

Now eight readers, twice what the device can serve at once.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-io-load.scope
$ sudo systemd-run --scope --unit=perf-io-load -p MemoryMax=256M -p TasksMax=32 \ fio --name=reader --filename=/dev/nullb0 --rw=randread --bs=4k --direct=1 --ioengine=psync \ --numjobs=8 --time_based --runtime=5m --group_reporting > /var/tmp/perf-io-fio.log 2>&1 &
$ iostat -dxy nullb0 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util nullb0 348.10 1392.42 0.00 0.00 22.85 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 7.95 99.84

Throughput did not grow, 348.1 reads per second, but r_await doubled to 22.85 ms and aqu-sz rose to 7.95. Four requests are always being served and four wait in the scheduler for a slot, each for about one service time. That is saturation: more demand than the device can serve, visible as a queue longer than its parallelism and a latency above its service time, while throughput stays flat. On a real disk you seldom know the parallelism, so compare await with what the device shows when it is quiet and watch whether throughput still grows when the queue does.

On cloud block storage and in VMs, read that plateau with care. A cloud volume usually has a provisioned limit on IOPS and throughput, sometimes with burst credits that run out, and hitting it looks exactly like this: await rises and throughput stops growing. iostat cannot tell a provider's cap from the device's own parallelism, and in a VM await also includes queueing on the host, shared with other guests. Before tuning queue depths or schedulers in the guest, compare the plateau with the volume's provisioned numbers and its burst balance in the provider's metrics; no guest setting raises a provider cap.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/pressure/io
some avg10=96.67 avg60=53.64 avg300=15.94 total=60898161 full avg10=94.50 avg60=52.01 avg300=15.37 total=57855852

I/O pressure agrees from the tasks' side. some at 96.67 means at least one task waited for I/O almost all of the last ten seconds; full at 94.50 means all non-idle tasks were waiting at once, so little else could run. (PSI is off by default on RHEL, as the method lesson showed.) Next, who is doing the I/O.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo pidstat -d 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:26:49 UID PID kB_rd/s kB_wr/s kB_ccwr/s iodelay Command 09:26:54 0 223303 171.54 0.00 0.00 0 fio 09:26:54 0 223304 170.75 0.00 0.00 0 fio 09:26:54 0 223305 170.75 0.00 0.00 0 fio 09:26:54 0 223306 171.54 0.00 0.00 0 fio 09:26:54 0 223307 169.96 0.00 0.00 0 fio 09:26:54 0 223308 169.96 0.00 0.00 0 fio 09:26:54 0 223309 170.75 0.00 0.00 0 fio 09:26:54 0 223310 170.75 0.00 0.00 0 fio …

pidstat -d shows each process's read and write rate from /proc/PID/io. It needs sudo to see other users' processes, here the eight fio jobs owned by root, each reading about 171 KiB/s, which is 43 reads of 4 KiB per second. iodelay is 0, a problem the next section solves.

An average of 23 ms can hide very different distributions, for example most requests at 2 ms and a few at 500. A latency histogram shows the shape. The bcc tool biolatency-bpfcc does not compile against Ubuntu 26.04's kernel headers with bcc 0.35; the bpftrace version, biolatency.bt, works. On RHEL the bcc tool, /usr/share/bcc/tools/biolatency, works.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo biolatency-bpfcc 5 1
… Exception: Failed to compile BPF module <text>
$ sudo timeout --preserve-status -s INT 10 biolatency.bt
… Tracing block device I/O... Hit Ctrl-C to end. … @usecs: [64, 128) 1 | | [128, 256) 2 | | [256, 512) 26 | | [512, 1K) 9 | | [1K, 2K) 25 | | [2K, 4K) 1 | | [4K, 8K) 0 | | [8K, 16K) 26 | | [16K, 32K) 3273 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [32K, 64K) 13 | |

biolatency.bt runs until interrupted, so timeout -s INT 10 stops it after ten seconds, and --preserve-status keeps its exit status. It prints warnings about its own script first (left out here); they are harmless. The histogram counts requests by latency in microseconds, in power-of-two buckets, and it covers every block device on the machine. The test disk is the hump of 3273 requests between 16 and 32 ms, queue plus service. The few dozen requests below 8 ms cannot be the test disk, where nothing finishes in under 10 ms. The tool times each request from the moment its bio is queued, so it also includes any time spent waiting for a free request slot. It keys its start times on the sector number alone, though, so it cannot say which device a request went to. The same few lines of bpftrace keyed on the device number as well split the histogram per disk (the device number packs the major number above bit 20):

deploy@web01 · Ubuntu 26.04 LTS
$ lsblk -dno NAME,MAJ:MIN /dev/vda /dev/nullb0 sudo timeout --preserve-status -s INT 10 bpftrace -e ' tracepoint:block:block_bio_queue { @start[args.dev, args.sector] = nsecs; } tracepoint:block:block_rq_complete, tracepoint:block:block_bio_complete /@start[args.dev, args.sector]/ { @usecs[args.dev >> 20, args.dev & 0xfffff] = hist((nsecs - @start[args.dev, args.sector]) / 1000); delete(@start[args.dev, args.sector]); } END { clear(@start); }'
nullb0 251:0 vda 253:0 … @usecs[253, 0]: [32, 64) 1 |@@@@ | [64, 128) 0 | | [128, 256) 2 |@@@@@@@@@ | [256, 512) 2 |@@@@@@@@@ | [512, 1K) 0 | | [1K, 2K) 11 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [2K, 4K) 4 |@@@@@@@@@@@@@@@@@@ | @usecs[251, 0]: [8K, 16K) 21 | | [16K, 32K) 3322 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| [32K, 64K) 13 | |

[253, 0] is vda, the VM's system disk: twenty fast requests from the rest of the system. [251, 0] is nullb0, and every one of its requests took between 8 and 64 ms. That is what a histogram shows and an average hides: two populations, one fast and one slow, which here belong to two devices.

On RHEL, run the bcc tool by its path, sudo /usr/share/bcc/tools/biolatency 5 1; it prints the same kind of histogram in bcc's format.

Who is waiting: delay accounting

The kernel can record, per task, how long it waited for block I/O, swap-in and other resources: delay accounting. pidstat's iodelay column and iotop's IO column come from it. Both Ubuntu 26.04 and RHEL 10 leave it off by default; the delayacct boot option or the kernel.task_delayacct sysctl turns it on.

deploy@rocky10 · Rocky Linux 10.2
$ sysctl kernel.task_delayacct
kernel.task_delayacct = 0

That is RHEL's value; Ubuntu's reads the same below. The C rewrite of iotop, iotop-c, is packaged on both; on Ubuntu it also provides the iotop command. Switch accounting on and install it:

deploy@web01 · Ubuntu 26.04 LTS
$ sysctl kernel.task_delayacct sudo sysctl kernel.task_delayacct=1
kernel.task_delayacct = 0 kernel.task_delayacct = 1
$ sudo apt install -y iotop-c
… Setting up iotop-c (1.31-1) ... update-alternatives: using /usr/sbin/iotop-c to provide /usr/sbin/iotop (iotop) in auto mode …
$ sudo iotop-c -b -o -P -n 2 -d 5
… Total DISK READ: 1400.00 K/s | Total DISK WRITE: 48.41 K/s Current DISK READ: 1400.00 K/s | Current DISK WRITE: 274.60 K/s PID PRIO USER DISK READ DISK WRITE SWAPIN IO COMMAND 311 be/3 root 0.00 B/s 812.70 B/s 0.00 % 0.00 % <unknown> … 223303 be/4 root 175.40 K/s 0.00 B/s 0.00 % 0.00 % fio 223304 be/4 root 175.40 K/s 0.00 B/s 0.00 % 0.00 % fio 223305 be/4 root 175.40 K/s 0.00 B/s 0.00 % 0.00 % fio 223306 be/4 root 175.40 K/s 0.00 B/s 0.00 % 0.00 % fio 223307 be/4 root 175.40 K/s 0.00 B/s 0.00 % 0.00 % fio 223308 be/4 root 174.60 K/s 0.00 B/s 0.00 % 0.00 % fio 223309 be/4 root 174.60 K/s 0.00 B/s 0.00 % 0.00 % fio 223310 be/4 root 173.81 K/s 0.00 B/s 0.00 % 0.00 % fio

The eight readers are waiting all the time, and IO still reads 0.00%. The kernel documentation explains it: only tasks started after delay accounting was switched on have delay information. Restart the workload.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-io-load.scope sudo systemd-run --scope --unit=perf-io-load -p MemoryMax=256M -p TasksMax=32 \ fio --name=reader --filename=/dev/nullb0 --rw=randread --bs=4k --direct=1 --ioengine=psync \ --numjobs=8 --time_based --runtime=5m --group_reporting > /var/tmp/perf-io-fio.log 2>&1 &
$ sudo iotop-c -b -o -P -n 2 -d 5
… Total DISK READ: 1372.33 K/s | Total DISK WRITE: 53.75 K/s Current DISK READ: 1372.33 K/s | Current DISK WRITE: 236.36 K/s PID PRIO USER DISK READ DISK WRITE SWAPIN IO COMMAND 232494 be/4 root 172.33 K/s 0.00 B/s 0.00 % 100.00 % fio 232495 be/4 root 171.54 K/s 0.00 B/s 0.00 % 100.00 % fio 232496 be/4 root 170.75 K/s 0.00 B/s 0.00 % 100.00 % fio 232497 be/4 root 171.54 K/s 0.00 B/s 0.00 % 100.00 % fio 232498 be/4 root 173.12 K/s 0.00 B/s 0.00 % 100.00 % fio 232499 be/4 root 172.33 K/s 0.00 B/s 0.00 % 100.00 % fio 232500 be/4 root 171.54 K/s 0.00 B/s 0.00 % 100.00 % fio 232501 be/4 root 169.17 K/s 0.00 B/s 0.00 % 100.00 % fio 311 be/3 root 0.00 B/s 6.32 K/s 0.00 % 0.00 % <unknown> …
$ sudo pidstat -d 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) 09:27:50 UID PID kB_rd/s kB_wr/s kB_ccwr/s iodelay Command 09:27:55 0 232494 177.25 0.00 0.00 779 fio 09:27:55 0 232495 177.25 0.00 0.00 730 fio 09:27:55 0 232496 177.25 0.00 0.00 808 fio 09:27:55 0 232497 178.04 0.00 0.00 752 fio 09:27:55 0 232498 177.25 0.00 0.00 755 fio 09:27:55 0 232499 178.84 0.00 0.00 767 fio 09:27:55 0 232500 178.04 0.00 0.00 750 fio 09:27:55 0 232501 177.25 0.00 0.00 835 fio …

Now each new fio job shows 100% in the IO column, the share of time it waited for block I/O, and pidstat's iodelay is no longer 0. pidstat(1) gives that column in clock ticks, 100 per second here, yet it reports more than the 500 ticks a 5-second interval can hold for one process, so read it only as a relative number and use iotop's percentage for the share of time. On a real server this is how you find the process that suffers from a slow disk, which is not always the one that causes the load. Switch accounting off again when you are done; it only costs a little, but it is not the platform default.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo sysctl kernel.task_delayacct=0
kernel.task_delayacct = 0

Stop the load and confirm recovery

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop perf-io-load.scope
$ iostat -dxy nullb0 5 1
Linux 7.0.0-34-generic (web01) 09/27/26 _aarch64_ (2 CPU) Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util nullb0 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
$ cat /proc/pressure/io
some avg10=22.63 avg60=64.01 avg300=30.63 total=123975731 full avg10=21.22 avg60=58.00 avg300=28.08 total=113528926
$ cat /var/tmp/perf-io-fio.log
… reader: (groupid=0, jobs=8): err= 0: pid=232494: Sun Sep 27 09:27:55 2026 read: IOPS=350, BW=1400KiB/s (1434kB/s)(27.9MiB/20374msec) clat (usec): min=10089, max=74579, avg=22840.92, stdev=2838.33 lat (usec): min=10089, max=74579, avg=22841.21, stdev=2838.41 clat percentiles (usec): | 1.00th=[19006], 5.00th=[20317], 10.00th=[20579], 20.00th=[21365], | 30.00th=[21890], 40.00th=[22152], 50.00th=[22676], 60.00th=[22938], | 70.00th=[23462], 80.00th=[23987], 90.00th=[25035], 95.00th=[26346], | 99.00th=[30278], 99.50th=[34866], 99.90th=[65799], 99.95th=[73925], | 99.99th=[74974] …

Fifteen seconds after the stop the device is idle. PSI's ten-second average has fallen to 22.63 while the 60-second one still reads 64.01: averages lag, as in the CPU lessons. fio's own report is the application's view and agrees with the block layer's: 350 reads per second at an average of 22.8 ms (22841 µs), with the median at 22.7 ms and the 99th percentile at 30.3 ms. Remove the test device; the exercise at the end loads it again.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo modprobe -r null_blk ls /dev/nullb0
ls: cannot access '/dev/nullb0': No such file or directory

When df and du disagree

The other storage emergency is a full filesystem where df and du disagree. The files lesson showed the mechanism: a deleted file that a process still holds open keeps its blocks until the last descriptor closes. Here is how it looks from the disk's side, on a small loop filesystem with a writer that behaves like a logging daemon: it appends to app.log through standard output, writes 40 MiB, then adds a line every second.

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 64M /var/tmp/perf-io-fs.img sudo mkfs.ext4 -q /var/tmp/perf-io-fs.img sudo mkdir /mnt/perf-io sudo mount -o loop /var/tmp/perf-io-fs.img /mnt/perf-io sudo chown $USER: /mnt/perf-io
$ sudo systemd-run --unit=perf-io-writer --uid=$USER sh -c 'exec >>/mnt/perf-io/app.log; head -c 40M /dev/urandom; while sleep 1; do date; done'
Running as unit: perf-io-writer.service; invocation ID: 614b0930d0c245d4aedb4e6842d880c8
$ rm /mnt/perf-io/app.log df -h /mnt/perf-io sudo du -sh /mnt/perf-io
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 41M 12M 78% /mnt/perf-io 20K /mnt/perf-io

df asks the filesystem how many blocks are in use: 41M. du adds up the files it can reach by name: 20K. The difference is a file with no name left. lsof +L1 lists open files with a link count below one.

deploy@web01 · Ubuntu 26.04 LTS
$ lsof +L1 /mnt/perf-io
… COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME sh 232986 deploy 1w REG 7,0 41943098 0 13 /mnt/perf-io/app.log (deleted) sleep 233019 deploy 1w REG 7,0 41943098 0 13 /mnt/perf-io/app.log (deleted)
$ ls -l /proc/$(systemctl show -P MainPID perf-io-writer)/fd
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:28 0 -> /dev/null l-wx------ 1 deploy deploy 64 Sep 27 09:28 1 -> /mnt/perf-io/app.log (deleted) lrwx------+ 1 deploy deploy 64 Sep 27 09:28 2 -> socket:[807548]

(lsof first warns that it cannot inspect the tracing filesystem as a normal user; that line is left out.) NLINK 0 and (deleted) mark the file, 41943098 bytes, held on descriptor 1 by the shell and by the sleep it started, because a child inherits its parent's descriptors. The clean fix is to make the writer close and reopen its file: restart it, or send the signal it uses for that (for rsyslog, systemctl kill -s HUP rsyslog, which is what logrotate does). When you cannot do that right now, truncate the file through the descriptor.

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 0 /proc/$(systemctl show -P MainPID perf-io-writer)/fd/1 df -h /mnt/perf-io
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 152K 52M 1% /mnt/perf-io
$ stat -L -c "%s bytes" /proc/$(systemctl show -P MainPID perf-io-writer)/fd/1
87 bytes
$ sudo systemctl stop perf-io-writer df -h /mnt/perf-io
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 152K 52M 1% /mnt/perf-io

The space came back at once and the writer carried on: three seconds later the file holds 87 bytes, three new lines. That works here because the shell opened the log with >> (O_APPEND); a writer without O_APPEND would carry on at its old offset and leave a sparse file of the old size, as the files lesson showed with copytruncate. Truncation also destroys the log's contents, so copy anything you need from /proc/PID/fd/N first. Finally remove the test filesystem:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo umount /mnt/perf-io sudo rmdir /mnt/perf-io rm /var/tmp/perf-io-fs.img /var/tmp/perf-io-fio.log

Try this

Load the test disk again with the same modprobe command, but leave out the switch to mq-deadline, and check lsblk -t /dev/nullb0: this time it keeps its default scheduler, none. Start the eight readers again, wait 15 seconds, and predict iostat -dxy nullb0 5 1 and /proc/pressure/io before you run them. Expect RQ-SIZE 4, and iostat to show about 358 reads per second at 11.1 ms with aqu-sz near 4, exactly like four readers, while some pressure stays above 80 and fio reports about 22.4 ms per read. With no scheduler the queue holds only as many requests as the device has slots; the other four readers wait for a free request before the block layer starts timing them, so iostat cannot see their wait. Stop the scope and remove the device with sudo modprobe -r null_blk.

Takeaway

Judge a disk by its latency compared with its quiet-time latency and by whether throughput still grows when the queue grows; %util only tells you it was never idle, and PSI or the application's own timings tell you when iostat is missing part of the wait.

Quick check
01A database on an NVMe drive gets more traffic every hour. iostat -x shows %util at 100, r_await steady at 0.2 ms, aqu-sz around 3, and r/s rising with the traffic. An alert says the disk is saturated. What do you conclude?
Incorrect — %util only records that some request was always in flight. An NVMe drive serves many at once, so it can do more at 100%.
Correct — Saturation shows as await rising above the service time while throughput stops growing; neither is happening.
Incorrect — aqu-sz counts requests being served as well as waiting ones. Three in flight on a parallel device can all be in service at once.
Incorrect — The scheduler does not create capacity, and nothing here shows a shortage of it; NVMe drives normally run with none.
02During an incident you enable kernel.task_delayacct=1 and run iotop-c. PSI io shows heavy stalls, yet the IO column is 0.00% for the application's worker processes, which were started last week. Why?
Incorrect — -a changes rates into totals; it does not create delay data that the kernel never recorded for these tasks.
Incorrect — Both see the same block I/O waits. PSI counts them system-wide; delay accounting would count them per task if it had been recording.
Correct — The kernel sets up delay tracking for a task when it is created; restart the workers, or boot with delayacct, to see their wait.
Incorrect — The IO column is the share of time a task waited on block I/O of any direction, plus swap-in in its own column.
03A database on a cloud VM slows down every day at peak. On its block volume, iostat -x shows w/s flat at about 3000 for an hour while w_await climbs from 1 ms to 40 ms and aqu-sz from 3 to 120. What is the most useful next step?
Incorrect — A scheduler reorders requests; it does not add capacity, and if the limit is outside the guest nothing in the guest's queue changes it.
Incorrect — A deeper queue only lets aqu-sz grow further. By Little's law, at a fixed 3000 writes per second more queued requests mean a longer await.
Incorrect — Steady throughput with a growing queue and latency is the signature of saturation, not of a healthy device.
Correct — A flat ceiling with rising await is what a provider's IOPS cap or exhausted burst credits look like, and iostat cannot tell that apart from a device limit.

Related