The block layer, device-mapper and LVM

The I/O path down to the disk, with LVM.

Advanced16 min · lesson 13 of 21

Between a filesystem and a disk sit two more layers: the block layer, which queues, merges and schedules requests, and on many servers device-mapper, which builds virtual disks such as LVM volumes, encrypted volumes and snapshots. This lesson follows one I/O request through them, reads each layer's settings, and then builds an LVM stack on two loop devices so you can see what LVM asks the kernel to do. By the end you can tell which layer a setting or a problem belongs to, grow a volume and its filesystem without unmounting it, size and watch a snapshot so that it does not break, and deal with the LVM devices file on RHEL. The filesystem lesson covered what happens above this; the next lesson measures the I/O that passes through it.

The path of a block I/O request

When a filesystem needs data from a disk, or has dirty pages to write back, it builds a bio: the kernel's description of one I/O, naming the device, the starting sector (a 512-byte unit), the direction and the memory pages involved. The bio is handed to the block layer, which on current kernels is blk-mq, the multi-queue block layer. blk-mq puts it on a software staging queue, one per CPU, where bios become requests and adjacent requests are merged into larger ones. An I/O scheduler may then reorder them. Finally the requests move to a hardware dispatch queue, where the device driver takes them and sends them to the device; the device's completion interrupt ends each one.

One read or write, from the filesystem to the disk
1Filesystem or page cache
builds a bio: device, sector, pages, direction
2device-mapper (optional)
LVM, dm-crypt: remaps the bio to another device
3blk-mq software queue
one per CPU: bios become requests, merged
4I/O scheduler
mq-deadline or none; RHEL also kyber and bfq
5Hardware dispatch queue
the driver takes requests: virtio, NVMe, SCSI
6Device
completion interrupt ends the request
device-mapper works on bios before they reach the lower device's queue.

lsblk -t (topology) shows the block layer's view of a disk.

deploy@web01 · Ubuntu 26.04 LTS
$ lsblk -t /dev/vda
NAME ALIGNMENT MIN-IO OPT-IO PHY-SEC LOG-SEC ROTA SCHED RQ-SIZE RA WSAME vda 0 512 0 512 512 1 mq-deadline 256 8192 0B ├─vda1 0 512 0 512 512 1 mq-deadline 256 8192 0B ├─vda13 0 512 0 512 512 1 mq-deadline 256 8192 0B └─vda15 0 512 0 512 512 1 mq-deadline 256 8192 0B

ROTA is 1, a rotational disk, because that is what the virtual disk reports, not a measurement (the /proc and /sys lesson showed the same). SCHED is the active I/O scheduler. RQ-SIZE is nr_requests, how many requests the queue may hold for this disk. LOG-SEC and PHY-SEC are the logical and physical sector sizes. Partitions share their disk's queue, so they repeat its values.

deploy@web01 · Ubuntu 26.04 LTS
$ ls /sys/block/vda/mq cat /sys/block/vda/mq/0/cpu_list
0 0, 1
$ cat /sys/block/vda/queue/scheduler
none [mq-deadline]

/sys/block/vda/mq has one directory per hardware queue. This virtio disk has one, and both CPUs feed it. The number depends on the device and driver and is never more than the number of CPUs; NVMe drives usually offer enough for one per CPU, so CPUs do not contend for a single queue. The scheduler file lists the choices and brackets the active one. The kernel gives a disk with a single hardware queue mq-deadline, which gives every request a deadline, favouring reads, so that no request waits indefinitely behind others. Devices with several hardware queues get none, which passes requests through in order and leaves reordering to the device, the usual choice for NVMe.

deploy@rocky10 · Rocky Linux 10.2
$ cat /sys/block/vda/queue/scheduler
none [mq-deadline] kyber bfq

The RHEL kernel builds two more schedulers in, so they are always offered: kyber, which adjusts queue depths to meet target read and write latencies on fast devices, and bfq, a proportional-share scheduler that can divide a disk between cgroups by weight (the cgroup lesson explains why IOWeight= does nothing under the default schedulers). Ubuntu builds both as modules, which are not loaded, so its list shows only mq-deadline and none until one is. Changing the scheduler is a tuning decision that belongs in the last lesson, after you can measure its effect.

An LVM stack on two loop devices

LVM (the Logical Volume Manager) pools disks and hands out volumes that can grow, span disks and be snapshotted. A physical volume (PV) is a disk or partition given to LVM: pvcreate writes a label and a metadata area at its start. A volume group (VG) pools one or more PVs and divides them into extents, 4 MiB each by default. A logical volume (LV) is a set of extents that LVM presents as a block device. The Ubuntu Server installer uses an LVM layout by default and so does RHEL's automatic partitioning; cloud images, including the lab VMs, use plain partitions. The lab uses two 256 MiB files as disks, attached as loop devices.

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 256M /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img sudo losetup --find --show /var/tmp/st-lvm-a.img sudo losetup --find --show /var/tmp/st-lvm-b.img
/dev/loop0 /dev/loop1
$ sudo pvcreate /dev/loop0 /dev/loop1
Physical volume "/dev/loop0" successfully created. Physical volume "/dev/loop1" successfully created.
$ sudo vgcreate stlvm /dev/loop0 /dev/loop1
Volume group "stlvm" successfully created
$ sudo lvcreate --name data --size 200M stlvm
Logical volume "data" created.
$ sudo pvs sudo vgs sudo lvs
PV VG Fmt Attr PSize PFree /dev/loop0 stlvm lvm2 a-- 252.00m 52.00m /dev/loop1 stlvm lvm2 a-- 252.00m 252.00m VG #PV #LV #SN Attr VSize VFree stlvm 2 1 0 wz--n- 504.00m 304.00m LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert data stlvm -wi-a----- 200.00m

losetup --find --show attaches a file to the first free loop device and prints its name; use the names it prints in the commands that follow. Each PV offers 252 MiB: LVM keeps the first MiB of the 256 MiB for its label and metadata, and the remaining 255 MiB hold 63 whole 4 MiB extents. The VG pools both into 504 MiB, and the 200 MiB LV took its extents from /dev/loop0, which is why that PV has 52 MiB free. In the Attr column, w is writable and a is active; a sixth letter o appears once the volume is open, for example mounted.

What LVM asks device-mapper to build

LVM itself is a set of user-space tools that keep metadata on the PVs. The mapping happens in the kernel: LVM loads a table into device-mapper, and device-mapper creates a block device that translates every bio according to that table.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo dmsetup table stlvm-data
0 409600 linear 7:0 2048
$ lsblk -o NAME,MAJ:MIN,SIZE,TYPE,SCHED /dev/loop0 /dev/loop1
NAME MAJ:MIN SIZE TYPE SCHED loop0 7:0 256M loop none └─stlvm-data 252:0 200M lvm loop1 7:1 256M loop none
$ ls -l /dev/stlvm/data cat /sys/block/dm-0/queue/scheduler ls /sys/block/dm-0/mq
lrwxrwxrwx 1 root root 7 Sep 27 09:25 /dev/stlvm/data -> ../dm-0 cat: /sys/block/dm-0/queue/scheduler: No such file or directory ls: cannot access '/sys/block/dm-0/mq': No such file or directory

Read the table line as: sectors 0 to 409600 of this device (200 MiB in 512-byte sectors) use the linear target, which sends them to device 7:0 (loop0) starting at sector 2048, just past LVM's first MiB. The device appears as /dev/dm-0 with the friendly names /dev/stlvm/data and /dev/mapper/stlvm-data; in mapper names a dash inside a VG or LV name is doubled, so a VG called my-vg becomes my--vg.

The last step fails on purpose. A device-mapper device like this one has no scheduler and no hardware queues. It remaps each bio and submits it to the device below, and scheduling happens in the queue of loop0, whose scheduler is none (the SCHED column above). Two practical consequences follow: queue settings such as the scheduler belong on the physical disk under an LV, not on the LV, and iostat reports the same I/O twice, once for the dm- device and once for the disk below it; comparing the two rows shows which layer adds the latency.

Growing a volume and its filesystem online

Put an ext4 filesystem on the LV, mount it, and fill it close to capacity.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo mkfs.ext4 -q /dev/stlvm/data sudo mkdir /mnt/st-lvm sudo mount /dev/stlvm/data /mnt/st-lvm sudo chown $USER: /mnt/st-lvm
$ dd if=/dev/urandom of=/mnt/st-lvm/data.bin bs=1M count=150 status=none df -h /mnt/st-lvm
Filesystem Size Used Avail Use% Mounted on /dev/mapper/stlvm-data 172M 151M 7.2M 96% /mnt/st-lvm
$ sudo lvextend --size +200M stlvm/data
Size of logical volume stlvm/data changed from 200.00 MiB (50 extents) to 400.00 MiB (100 extents). Logical volume stlvm/data successfully resized.
$ df -h /mnt/st-lvm
Filesystem Size Used Avail Use% Mounted on /dev/mapper/stlvm-data 172M 151M 7.2M 96% /mnt/st-lvm

lvextend added 200 MiB, 50 extents, and df still reports 172M. Nothing is wrong with the volume: the filesystem inside it was created for 200 MiB and does not know the device grew. This is the most common mistake when adding space. The filesystem has to be grown as well, and ext4 can do that while it is mounted and in use.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo resize2fs /dev/stlvm/data
resize2fs 1.47.2 (1-Jan-2025) Filesystem at /dev/stlvm/data is mounted on /mnt/st-lvm; on-line resizing required old_desc_blocks = 1, new_desc_blocks = 1 The filesystem on /dev/stlvm/data is now 102400 (4k) blocks long.
$ df -h /mnt/st-lvm
Filesystem Size Used Avail Use% Mounted on /dev/mapper/stlvm-data 359M 151M 191M 45% /mnt/st-lvm
$ sudo dmsetup table stlvm-data
0 516096 linear 7:0 2048 516096 303104 linear 7:1 2048
$ sudo lvs --segments -o lv_name,seg_start,seg_size,devices stlvm
LV Start SSize Devices data 0 252.00m /dev/loop0(0) data 252.00m 148.00m /dev/loop1(0)

The table now has two lines: the first 516096 sectors (252 MiB) still come from loop0, and the rest from loop1, because loop0 had only 52 MiB left. One LV now spans two disks, and losing either disk loses the filesystem; spanning adds capacity, not redundancy. lvs --segments shows the same layout from LVM's side. For XFS the grow step is xfs_growfs on the mount point; XFS can shrink only within its last allocation group, and ext4 shrinks only when unmounted. lvextend --resizefs runs the right filesystem tool for you.

Snapshots and their copy-on-write space

An LVM snapshot is a second device that shows the origin volume as it was when the snapshot was taken. It does not copy the volume. Instead, the first time a chunk of the origin is about to be overwritten, device-mapper copies the old chunk into the snapshot's copy-on-write (COW) space, then lets the write go ahead. Reading the snapshot returns the saved chunk if it changed and the origin's chunk if it did not. The COW space only has to hold the chunks that change during the snapshot's life.

deploy@web01 · Ubuntu 26.04 LTS
$ echo "version 1" > /mnt/st-lvm/config.txt sync
$ sudo lvcreate --snapshot --name snap --size 32M stlvm/data
Logical volume "snap" created.
$ sudo lvs stlvm
LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert data stlvm owi-aos--- 400.00m snap stlvm swi-a-s--- 32.00m data 19.63
$ sudo dmsetup table | grep ^stlvm
stlvm-data: 0 819200 snapshot-origin 252:1 stlvm-data-real: 0 516096 linear 7:0 2048 stlvm-data-real: 516096 303104 linear 7:1 2048 stlvm-snap: 0 819200 snapshot 252:1 252:2 P 8 stlvm-snap-cow: 0 65536 linear 7:1 305152

The origin's attributes now begin with o (origin) and the snapshot's with s; Data% is the share of the 32 MiB COW space in use, and it already reads 19.63% although the test has not written anything yet: every first write to a chunk of the origin since the snapshot counts, including the filesystem's own background work (ext4 initialises the inode tables of a new or just-grown filesystem lazily, after mounting). Creating one snapshot made device-mapper build four devices. stlvm-data-real holds the original linear table. stlvm-data, the device you mounted, is now a snapshot-origin on top of it, which copies chunks before they are overwritten. stlvm-snap-cow is the 32 MiB of COW space on loop1, and stlvm-snap combines the real origin (252:1) with the COW device (252:2): P means the exceptions are stored persistently, and 8 is the chunk size in sectors, 4 KiB. To take the snapshot, LVM suspended the origin for a moment, and a device-mapper suspend first syncs the filesystem on the device (dmsetup(8)), so the snapshot holds a consistent ext4. A database still needs its own flush or backup mode for a consistent copy of its files.

deploy@web01 · Ubuntu 26.04 LTS
$ echo "version 2" > /mnt/st-lvm/config.txt dd if=/dev/urandom of=/mnt/st-lvm/new.bin bs=1M count=12 conv=fsync status=none sudo lvs stlvm
LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert data stlvm owi-aos--- 400.00m snap stlvm swi-a-s--- 32.00m data 57.41
$ sudo mkdir /mnt/st-lvm-snap sudo mount -o ro,noload /dev/stlvm/snap /mnt/st-lvm-snap cat /mnt/st-lvm/config.txt /mnt/st-lvm-snap/config.txt ls /mnt/st-lvm-snap
version 2 version 1 config.txt data.bin lost+found
$ sudo umount /mnt/st-lvm-snap

Writing 12 MiB to the origin added about 37.5 points (12 of 32 MiB), to 57.41%. The snapshot still shows version 1 and does not contain the new file. It is mounted with noload, which tells ext4 not to replay its journal, so that reading the snapshot changes nothing on it; for XFS the equivalent is ro,norecovery,nouuid, because a snapshot carries the same filesystem UUID as its origin. Now write more than the COW space can hold.

deploy@web01 · Ubuntu 26.04 LTS
$ dd if=/dev/urandom of=/mnt/st-lvm/big.bin bs=1M count=40 conv=fsync status=none sudo lvs stlvm
LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert data stlvm owi-aos--- 400.00m snap stlvm swi-I-s--- 32.00m data 100.00
$ journalctl -k --since '-2 min' --no-hostname --grep snapshot
Sep 27 09:25:56 kernel: device-mapper: snapshots: Invalidating snapshot: Unable to allocate exception.
$ ls -l /mnt/st-lvm df -h /mnt/st-lvm
total 206868 -rw-rw-r-- 1 deploy deploy 41943040 Sep 27 09:25 big.bin -rw-rw-r-- 1 deploy deploy 10 Sep 27 09:25 config.txt -rw-rw-r-- 1 deploy deploy 157286400 Sep 27 09:25 data.bin drwx------ 2 root root 16384 Sep 27 09:25 lost+found -rw-rw-r-- 1 deploy deploy 12582912 Sep 27 09:25 new.bin …
$ sudo lvremove --yes stlvm/snap
Logical volume "snap" successfully removed.

The fifth attribute letter I means invalid snapshot: the COW space filled, the kernel could not store the next chunk and dropped the snapshot. The origin carried on without an error, as ls shows, and nothing told the application. The snapshot is useless now; remove it and take a larger one.

Size snapshots for their lifetime, and watch Data%
A classic snapshot needs COW space for every chunk the origin changes while the snapshot exists, and every first write to a chunk costs an extra read and write while it exists. Size it for the writes you expect during the backup, alert on Data% from lvs, and remove it as soon as the backup is done. Thin-provisioned snapshots share one pool instead; when that pool is full, writes to every thin volume in it are queued for 60 seconds by default and then fail (lvmthin(7)), so a thin pool needs the same monitoring.

On RHEL: the LVM devices file

Ubuntu's LVM looks at every block device it finds, because its configuration leaves the devices file off:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo lvmconfig --typeconfig full devices/use_devicesfile
use_devicesfile=0

RHEL 10 turns the devices file on: LVM uses only the devices listed in /etc/lvm/devices/system.devices.

deploy@rocky10 · Rocky Linux 10.2
$ sudo lvmconfig --typeconfig full devices/use_devicesfile
use_devicesfile=1
$ sudo ls -A /etc/lvm/devices
$ truncate -s 256M /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img sudo losetup --find --show /var/tmp/st-lvm-a.img sudo losetup --find --show /var/tmp/st-lvm-b.img
/dev/loop0 /dev/loop1
$ sudo vgcreate stlvm /dev/loop0
Physical volume "/dev/loop0" successfully created. Creating devices file /etc/lvm/devices/system.devices Volume group "stlvm" successfully created
$ sudo lvmdevices
Device /dev/loop0 IDTYPE=loop_file IDNAME=/var/tmp/st-lvm-a.img DEVNAME=/dev/loop0 PVID=LrasaKfBHGeRHahDqhxFPrf8HPfTFZT6

The cloud image has no devices file, and without one LVM uses every device. The first vgcreate on a system with no VGs created it. Each entry identifies a device by a stable ID: a WWID or serial number for real disks, and for a loop device its backing file (IDTYPE=loop_file). pvcreate, vgcreate and vgextend add the devices they use. A disk that already carries a PV, for example one moved from another server, is not added by anything. The lab simulates one by creating a VG with --devicesfile "", which tells LVM to ignore the file for that command.

deploy@rocky10 · Rocky Linux 10.2
$ sudo vgcreate --devicesfile "" stlvm-moved /dev/loop1
Physical volume "/dev/loop1" successfully created. Volume group "stlvm-moved" successfully created
$ sudo vgs
VG #PV #LV #SN Attr VSize VFree stlvm 1 0 0 wz--n- 252.00m 252.00m
$ sudo vgs --devicesfile ""
VG #PV #LV #SN Attr VSize VFree stlvm 1 0 0 wz--n- 252.00m 252.00m stlvm-moved 1 0 0 wz--n- 252.00m 252.00m
$ sudo vgimportdevices stlvm-moved
Added 1 devices to devices file.
$ sudo vgs
VG #PV #LV #SN Attr VSize VFree stlvm 1 0 0 wz--n- 252.00m 252.00m stlvm-moved 1 0 0 wz--n- 252.00m 252.00m
$ sudo lvmdevices
Device /dev/loop0 IDTYPE=loop_file IDNAME=/var/tmp/st-lvm-a.img DEVNAME=/dev/loop0 PVID=LrasaKfBHGeRHahDqhxFPrf8HPfTFZT6 Device /dev/loop1 IDTYPE=loop_file IDNAME=/var/tmp/st-lvm-b.img DEVNAME=/dev/loop1 PVID=9VMdZuPzYFd85lTYKB0AeuQq5pOvEMqk

vgs does not list stlvm-moved although the disk is present, and nothing warns you. vgs --devicesfile "" shows it, and vgimportdevices adds all of that VG's PVs to the file; lvmdevices --adddev does the same for a single device. The file also changes how you remove a disk.

deploy@rocky10 · Rocky Linux 10.2
$ sudo vgremove --yes stlvm stlvm-moved for img in /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img; do dev=$(losetup --noheadings --output NAME --associated $img) sudo pvremove $dev sudo lvmdevices --deldev $dev sudo losetup --detach $dev done rm /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img
Volume group "stlvm" successfully removed Volume group "stlvm-moved" successfully removed Labels on physical volume "/dev/loop0" successfully wiped. Labels on physical volume "/dev/loop1" successfully wiped.
$ sudo grep -v "^#" /etc/lvm/devices/system.devices
PRODUCT_UUID=1198e3e1-37d4-fd4d-b5e0-dcaeb68efbb9 VERSION=1.1.6

pvremove wipes the label but leaves the entry, so remove it with lvmdevices --deldev. The file itself stays, now empty, and an existing empty devices file means LVM sees no devices at all until something adds one.

Try this

With the Ubuntu stack still in place and the snapshot removed, grow the volume once more in one command: sudo lvextend --resizefs --size +40M stlvm/data. Before you run it, predict what df -h /mnt/st-lvm will report and whether sudo dmsetup table stlvm-data will gain a third line. Expect lvextend to report 100 extents growing to 110 and to run resize2fs itself, df to show about 399M, and still two table lines: the new extents follow the existing segment on loop1, so the second line just grows from 303104 to 385024 sectors. Then take everything apart in order, top layer first. The loop devices are found again from their image files with losetup --associated, so the commands work whatever numbers your machine gave them:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo umount /mnt/st-lvm sudo rmdir /mnt/st-lvm /mnt/st-lvm-snap sudo vgremove --yes stlvm for img in /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img; do dev=$(losetup --noheadings --output NAME --associated $img) sudo pvremove $dev sudo losetup --detach $dev done rm /var/tmp/st-lvm-a.img /var/tmp/st-lvm-b.img
Logical volume "data" successfully removed. Volume group "stlvm" successfully removed Labels on physical volume "/dev/loop0" successfully wiped. Labels on physical volume "/dev/loop1" successfully wiped.

Takeaway

Before you change storage, read the stack with lsblk and dmsetup table: every operation acts on one layer, and the usual failures come from forgetting the layer above (a filesystem that was not grown) or below (a queue setting on a device-mapper device, a full COW space, a disk missing from the devices file).

Quick check
01To try a different I/O scheduler for a database on the LV /dev/vgdata/pg (dm-3), you look for /sys/block/dm-3/queue/scheduler and it does not exist. The LV sits on the disk /dev/sdb. Where does the setting belong, and why?
Incorrect — LVM has no scheduler option. The dm device has no scheduler because it only remaps bios; there is nothing for lvchange to set.
Incorrect — The requests are scheduled, just not on the dm device. After device-mapper remaps them they enter sdb's queue, where its scheduler applies.
Correct — The LV's I/O is queued, merged and scheduled on the physical disk, so the scheduler and nr_requests are set on sdb.
Incorrect — The LV was already active; a bio-based device-mapper device has no scheduler file at any point, mounted or not.
02A nightly job copies files from an LVM snapshot of /srv. This morning the copy failed with I/O errors, lvs shows the snapshot as swi-I-s--- with Data% 100.00, and the application writing to /srv reports no errors at all. What happened?
Correct — The I flag is an invalid snapshot. The origin is unaffected; remove the snapshot and retake it with room for the night's writes.
Incorrect — The origin keeps its own extents and kept accepting writes; only the snapshot's COW space was exhausted.
Incorrect — LVM has no time limit on snapshots. Only the amount of changed data counts, and reading the snapshot uses no COW space.
Incorrect — noload avoids changing the snapshot during forensic reads, but a journal replay cannot fill the COW space to 100% and mark it invalid.
03You move a data disk holding a VG from one RHEL 10 server to another. lsblk on the new server shows the disk and its size, but vgs and pvs do not list anything from it. What is the most likely cause and the safe fix?
Incorrect — pvcreate writes a new label over the old one, which risks the data. lsblk and the unchanged disk give no reason to think the label is gone.
Incorrect — The cache is not what hides the disk. On RHEL 10 LVM ignores any device missing from the devices file, however often it rescans.
Incorrect — When a devices file is in use, LVM ignores the regex filter settings, so editing the filter changes nothing.
Correct — With use_devicesfile=1, LVM uses only listed devices. vgimportdevices, or lvmdevices --adddev for one disk, adds it without touching the data.

Related