The block layer, device-mapper and LVM
The I/O path down to the disk, with LVM.
Between a filesystem and a disk sit two more layers: the block layer, which queues, merges and schedules requests, and on many servers device-mapper, which builds virtual disks such as LVM volumes, encrypted volumes and snapshots. This lesson follows one I/O request through them, reads each layer's settings, and then builds an LVM stack on two loop devices so you can see what LVM asks the kernel to do. By the end you can tell which layer a setting or a problem belongs to, grow a volume and its filesystem without unmounting it, size and watch a snapshot so that it does not break, and deal with the LVM devices file on RHEL. The filesystem lesson covered what happens above this; the next lesson measures the I/O that passes through it.
The path of a block I/O request
When a filesystem needs data from a disk, or has dirty pages to write back, it builds a bio: the kernel's description of one I/O, naming the device, the starting sector (a 512-byte unit), the direction and the memory pages involved. The bio is handed to the block layer, which on current kernels is blk-mq, the multi-queue block layer. blk-mq puts it on a software staging queue, one per CPU, where bios become requests and adjacent requests are merged into larger ones. An I/O scheduler may then reorder them. Finally the requests move to a hardware dispatch queue, where the device driver takes them and sends them to the device; the device's completion interrupt ends each one.
lsblk -t (topology) shows the block layer's view of a disk.
ROTA is 1, a rotational disk, because that is what the virtual disk reports, not a measurement (the /proc and /sys lesson showed the same). SCHED is the active I/O scheduler. RQ-SIZE is nr_requests, how many requests the queue may hold for this disk. LOG-SEC and PHY-SEC are the logical and physical sector sizes. Partitions share their disk's queue, so they repeat its values.
/sys/block/vda/mq has one directory per hardware queue. This virtio disk has one, and both CPUs feed it. The number depends on the device and driver and is never more than the number of CPUs; NVMe drives usually offer enough for one per CPU, so CPUs do not contend for a single queue. The scheduler file lists the choices and brackets the active one. The kernel gives a disk with a single hardware queue mq-deadline, which gives every request a deadline, favouring reads, so that no request waits indefinitely behind others. Devices with several hardware queues get none, which passes requests through in order and leaves reordering to the device, the usual choice for NVMe.
The RHEL kernel builds two more schedulers in, so they are always offered: kyber, which adjusts queue depths to meet target read and write latencies on fast devices, and bfq, a proportional-share scheduler that can divide a disk between cgroups by weight (the cgroup lesson explains why IOWeight= does nothing under the default schedulers). Ubuntu builds both as modules, which are not loaded, so its list shows only mq-deadline and none until one is. Changing the scheduler is a tuning decision that belongs in the last lesson, after you can measure its effect.
An LVM stack on two loop devices
LVM (the Logical Volume Manager) pools disks and hands out volumes that can grow, span disks and be snapshotted. A physical volume (PV) is a disk or partition given to LVM: pvcreate writes a label and a metadata area at its start. A volume group (VG) pools one or more PVs and divides them into extents, 4 MiB each by default. A logical volume (LV) is a set of extents that LVM presents as a block device. The Ubuntu Server installer uses an LVM layout by default and so does RHEL's automatic partitioning; cloud images, including the lab VMs, use plain partitions. The lab uses two 256 MiB files as disks, attached as loop devices.
losetup --find --show attaches a file to the first free loop device and prints its name; use the names it prints in the commands that follow. Each PV offers 252 MiB: LVM keeps the first MiB of the 256 MiB for its label and metadata, and the remaining 255 MiB hold 63 whole 4 MiB extents. The VG pools both into 504 MiB, and the 200 MiB LV took its extents from /dev/loop0, which is why that PV has 52 MiB free. In the Attr column, w is writable and a is active; a sixth letter o appears once the volume is open, for example mounted.
What LVM asks device-mapper to build
LVM itself is a set of user-space tools that keep metadata on the PVs. The mapping happens in the kernel: LVM loads a table into device-mapper, and device-mapper creates a block device that translates every bio according to that table.
Read the table line as: sectors 0 to 409600 of this device (200 MiB in 512-byte sectors) use the linear target, which sends them to device 7:0 (loop0) starting at sector 2048, just past LVM's first MiB. The device appears as /dev/dm-0 with the friendly names /dev/stlvm/data and /dev/mapper/stlvm-data; in mapper names a dash inside a VG or LV name is doubled, so a VG called my-vg becomes my--vg.
The last step fails on purpose. A device-mapper device like this one has no scheduler and no hardware queues. It remaps each bio and submits it to the device below, and scheduling happens in the queue of loop0, whose scheduler is none (the SCHED column above). Two practical consequences follow: queue settings such as the scheduler belong on the physical disk under an LV, not on the LV, and iostat reports the same I/O twice, once for the dm- device and once for the disk below it; comparing the two rows shows which layer adds the latency.
Growing a volume and its filesystem online
Put an ext4 filesystem on the LV, mount it, and fill it close to capacity.
lvextend added 200 MiB, 50 extents, and df still reports 172M. Nothing is wrong with the volume: the filesystem inside it was created for 200 MiB and does not know the device grew. This is the most common mistake when adding space. The filesystem has to be grown as well, and ext4 can do that while it is mounted and in use.
The table now has two lines: the first 516096 sectors (252 MiB) still come from loop0, and the rest from loop1, because loop0 had only 52 MiB left. One LV now spans two disks, and losing either disk loses the filesystem; spanning adds capacity, not redundancy. lvs --segments shows the same layout from LVM's side. For XFS the grow step is xfs_growfs on the mount point; XFS can shrink only within its last allocation group, and ext4 shrinks only when unmounted. lvextend --resizefs runs the right filesystem tool for you.
Snapshots and their copy-on-write space
An LVM snapshot is a second device that shows the origin volume as it was when the snapshot was taken. It does not copy the volume. Instead, the first time a chunk of the origin is about to be overwritten, device-mapper copies the old chunk into the snapshot's copy-on-write (COW) space, then lets the write go ahead. Reading the snapshot returns the saved chunk if it changed and the origin's chunk if it did not. The COW space only has to hold the chunks that change during the snapshot's life.
The origin's attributes now begin with o (origin) and the snapshot's with s; Data% is the share of the 32 MiB COW space in use, and it already reads 19.63% although the test has not written anything yet: every first write to a chunk of the origin since the snapshot counts, including the filesystem's own background work (ext4 initialises the inode tables of a new or just-grown filesystem lazily, after mounting). Creating one snapshot made device-mapper build four devices. stlvm-data-real holds the original linear table. stlvm-data, the device you mounted, is now a snapshot-origin on top of it, which copies chunks before they are overwritten. stlvm-snap-cow is the 32 MiB of COW space on loop1, and stlvm-snap combines the real origin (252:1) with the COW device (252:2): P means the exceptions are stored persistently, and 8 is the chunk size in sectors, 4 KiB. To take the snapshot, LVM suspended the origin for a moment, and a device-mapper suspend first syncs the filesystem on the device (dmsetup(8)), so the snapshot holds a consistent ext4. A database still needs its own flush or backup mode for a consistent copy of its files.
Writing 12 MiB to the origin added about 37.5 points (12 of 32 MiB), to 57.41%. The snapshot still shows version 1 and does not contain the new file. It is mounted with noload, which tells ext4 not to replay its journal, so that reading the snapshot changes nothing on it; for XFS the equivalent is ro,norecovery,nouuid, because a snapshot carries the same filesystem UUID as its origin. Now write more than the COW space can hold.
The fifth attribute letter I means invalid snapshot: the COW space filled, the kernel could not store the next chunk and dropped the snapshot. The origin carried on without an error, as ls shows, and nothing told the application. The snapshot is useless now; remove it and take a larger one.
Data% from lvs, and remove it as soon as the backup is done. Thin-provisioned snapshots share one pool instead; when that pool is full, writes to every thin volume in it are queued for 60 seconds by default and then fail (lvmthin(7)), so a thin pool needs the same monitoring.On RHEL: the LVM devices file
Ubuntu's LVM looks at every block device it finds, because its configuration leaves the devices file off:
RHEL 10 turns the devices file on: LVM uses only the devices listed in /etc/lvm/devices/system.devices.
The cloud image has no devices file, and without one LVM uses every device. The first vgcreate on a system with no VGs created it. Each entry identifies a device by a stable ID: a WWID or serial number for real disks, and for a loop device its backing file (IDTYPE=loop_file). pvcreate, vgcreate and vgextend add the devices they use. A disk that already carries a PV, for example one moved from another server, is not added by anything. The lab simulates one by creating a VG with --devicesfile "", which tells LVM to ignore the file for that command.
vgs does not list stlvm-moved although the disk is present, and nothing warns you. vgs --devicesfile "" shows it, and vgimportdevices adds all of that VG's PVs to the file; lvmdevices --adddev does the same for a single device. The file also changes how you remove a disk.
pvremove wipes the label but leaves the entry, so remove it with lvmdevices --deldev. The file itself stays, now empty, and an existing empty devices file means LVM sees no devices at all until something adds one.
Try this
With the Ubuntu stack still in place and the snapshot removed, grow the volume once more in one command: sudo lvextend --resizefs --size +40M stlvm/data. Before you run it, predict what df -h /mnt/st-lvm will report and whether sudo dmsetup table stlvm-data will gain a third line. Expect lvextend to report 100 extents growing to 110 and to run resize2fs itself, df to show about 399M, and still two table lines: the new extents follow the existing segment on loop1, so the second line just grows from 303104 to 385024 sectors. Then take everything apart in order, top layer first. The loop devices are found again from their image files with losetup --associated, so the commands work whatever numbers your machine gave them:
Takeaway
Before you change storage, read the stack with lsblk and dmsetup table: every operation acts on one layer, and the usual failures come from forgetting the layer above (a filesystem that was not grown) or below (a queue setting on a device-mapper device, a full COW space, a disk missing from the devices file).