Filesystem internals: ext4 and XFS

Inodes, extents, journaling and repair.

Advanced16 min · lesson 12 of 21

After a power cut, a settings file exists but is empty; after a disk error, a directory's files turn up in lost+found under numbers; fsck reports an XFS volume clean although it is damaged. Each of these follows from how the filesystem stores files and when it writes them, and this lesson reproduces all three. Ubuntu installs its root filesystem as ext4 and RHEL installs XFS, so those are the two you will build, on loop devices where a file stands in for a disk. You will see how each stores a file (inodes and extents), when written data actually reaches the disk (page cache, writeback, delayed allocation and fsync), what the journal protects and what it costs, and how to check and repair a filesystem that you damage on purpose. The previous lesson covered the VFS objects above the filesystem. Mount options as a security control are covered in the Linux hardening course (optional), in "Mount options for /tmp and friends".

Two filesystems in files

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 256M /var/tmp/st-fs-ext4.img truncate -s 512M /var/tmp/st-fs-xfs.img ls -ls /var/tmp/st-fs-*.img
0 -rw-rw-r-- 1 deploy deploy 268435456 Sep 27 09:25 /var/tmp/st-fs-ext4.img 0 -rw-rw-r-- 1 deploy deploy 536870912 Sep 27 09:25 /var/tmp/st-fs-xfs.img
$ mkfs.ext4 -L st-fs-ext4 /var/tmp/st-fs-ext4.img
mke2fs 1.47.2 (1-Jan-2025) Discarding device blocks: 0/65536 done Creating filesystem with 65536 4k blocks and 65536 inodes Filesystem UUID: 03b6ef89-658d-4d37-8421-0696e764dbfa Superblock backups stored on blocks: 32768 Allocating group tables: 0/2 done Writing inode tables: 0/2 done Creating journal (4096 blocks): done Writing superblocks and filesystem accounting information: 0/2 done

truncate creates sparse files: 256 MiB and 512 MiB long, but ls -s shows 0 blocks allocated until something is written. deploy can run mkfs on a file it owns; on a real disk the same command needs sudo. mke2fs made 65,536 blocks of 4 KiB and 65,536 inodes. An inode is the on-disk record of one file, and ext4 fixes how many there are when the filesystem is made, here one per 4 KiB of space (the small profile in /etc/mke2fs.conf, which mke2fs uses for filesystems from 3 MiB up to 512 MiB). Blocks are grouped into block groups of 32,768, a backup copy of the superblock sits in group 1, and the journal takes 4,096 blocks.

deploy@web01 · Ubuntu 26.04 LTS
$ mkfs.xfs -L st-fs-xfs /var/tmp/st-fs-xfs.img
meta-data=/var/tmp/st-fs-xfs.img isize=512 agcount=4, agsize=32768 blks = sectsz=512 attr=2, projid32bit=1 = crc=1 finobt=1, sparse=1, rmapbt=1 = reflink=1 bigtime=1 inobtcount=1 nrext64=1 = exchange=0 metadir=0 data = bsize=4096 blocks=131072, imaxpct=25 = sunit=0 swidth=0 blks naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=0 log =internal log bsize=4096 blocks=16384, version=2 = sectsz=512 sunit=0 blks, lazy-count=1 realtime =none extsz=4096 blocks=0, rtextents=0 = rgcount=0 rgsize=0 extents = zoned=0 start=0 reserved=0

XFS divides the space into allocation groups (agcount=4), each managing its own free space and inodes, so parallel writers do not contend for one allocator. Inodes are 512 bytes (isize) and are created in chunks as files appear; imaxpct=25 only caps how much space they may take. crc=1 means checksummed metadata, reflink=1 allows files to share extents (cp --reflink), and the internal log has 16,384 blocks.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo mkdir /mnt/st-fs-ext4 /mnt/st-fs-xfs sudo mount -o loop /var/tmp/st-fs-ext4.img /mnt/st-fs-ext4 sudo mount -o loop /var/tmp/st-fs-xfs.img /mnt/st-fs-xfs sudo chown $USER: /mnt/st-fs-ext4 /mnt/st-fs-xfs findmnt -o TARGET,SOURCE,FSTYPE,OPTIONS /mnt/st-fs-ext4 findmnt -o TARGET,SOURCE,FSTYPE,OPTIONS /mnt/st-fs-xfs
TARGET SOURCE FSTYPE OPTIONS /mnt/st-fs-ext4 /dev/loop0 ext4 rw,relatime TARGET SOURCE FSTYPE OPTIONS /mnt/st-fs-xfs /dev/loop1 xfs rw,relatime,inode64,logbufs=8,logbsize=32k,noquota

mount -o loop attached each file to the first free loop device (here /dev/loop0 and /dev/loop1; on your machine the numbers depend on what else is attached) and mounted it with the default options. Such a loop device is detached again automatically when the filesystem is unmounted. From here on the kernel treats them like any other disk.

The path of a write
1write() on a descriptor
copies the data into the page cache and returns
2Dirty pages
wait in memory until writeback or fsync
3Filesystem (ext4, XFS)
allocates extents, journals the metadata
4Block layer
queues the I/O for the device (next lesson)
5Device
data is durable once the device confirms a flush
fsync() pushes one file down this path and waits for the device to confirm.

Inodes and extents

deploy@web01 · Ubuntu 26.04 LTS
$ echo hello > /mnt/st-fs-ext4/small stat /mnt/st-fs-ext4/small
File: /mnt/st-fs-ext4/small Size: 6 Blocks: 8 IO Block: 4096 regular file Device: 7,0 Inode: 13 Links: 1 Access: (0664/-rw-rw-r--) Uid: ( 1001/ deploy) Gid: ( 1001/ deploy) Access: 2026-09-27 09:25:48.586188670 +0000 Modify: 2026-09-27 09:25:48.586188670 +0000 Change: 2026-09-27 09:25:48.586188670 +0000 Birth: 2026-09-27 09:25:48.586188670 +0000
$ filefrag -v /mnt/st-fs-ext4/small
Filesystem type is: ef53 File size of /mnt/st-fs-ext4/small is 6 (1 block of 4096 bytes) ext: logical_offset: physical_offset: length: expected: flags: 0: 0.. 0: 0.. 0: 0: last,unknown_loc,delalloc,eof /mnt/st-fs-ext4/small: 1 extent found

stat reads the inode: number 13, one link (one name), 8 blocks of 512 bytes (one 4 KiB block reserved) and four timestamps; Birth is the creation time that ext4 stores. filefrag -v asks the filesystem where the data lives, and the answer is nowhere yet: the flags unknown_loc,delalloc mean ext4 has reserved space but not chosen a location. That is delayed allocation. ext4 and XFS pick the physical blocks when the data is written back, when they know how much there is, so a file written in many small pieces can still get one contiguous extent.

deploy@web01 · Ubuntu 26.04 LTS
$ sync grep -E '^(Dirty|Writeback):' /proc/meminfo dd if=/dev/zero of=/mnt/st-fs-ext4/big bs=1M count=16 status=none grep -E '^(Dirty|Writeback):' /proc/meminfo filefrag -v /mnt/st-fs-ext4/big
Dirty: 24 kB Writeback: 4 kB Dirty: 15896 kB Writeback: 8192 kB Filesystem type is: ef53 File size of /mnt/st-fs-ext4/big is 16777216 (4096 blocks of 4096 bytes) ext: logical_offset: physical_offset: length: expected: flags: 0: 0.. 4095: 38912.. 43007: 4096: last,eof /mnt/st-fs-ext4/big: 1 extent found
$ sync grep -E '^(Dirty|Writeback):' /proc/meminfo filefrag -v /mnt/st-fs-ext4/small /mnt/st-fs-ext4/big
Dirty: 56 kB Writeback: 0 kB Filesystem type is: ef53 File size of /mnt/st-fs-ext4/small is 6 (1 block of 4096 bytes) ext: logical_offset: physical_offset: length: expected: flags: 0: 0.. 0: 4171.. 4171: 1: last,eof /mnt/st-fs-ext4/small: 1 extent found File size of /mnt/st-fs-ext4/big is 16777216 (4096 blocks of 4096 bytes) ext: logical_offset: physical_offset: length: expected: flags: 0: 0.. 4095: 38912.. 43007: 4096: last,eof /mnt/st-fs-ext4/big: 1 extent found

The first sync flushed everything else, so Dirty in /proc/meminfo started near zero. After dd had written 16 MiB and exited, Dirty read about 15.5 MB and Writeback (pages being written out right now) 8 MB. Together that is more than the 16 MiB written, because both counters cover the whole machine, and a loop device writes its image file through the page cache of /var/tmp in turn: the same data is counted once for big on the test filesystem and again for the image file under it. On a real disk you would see only the first. Writeback of big had already started, so ext4 had allocated its blocks, and filefrag shows one extent. After sync, Dirty is close to zero and small has a real block (4171).

deploy@web01 · Ubuntu 26.04 LTS
$ debugfs -R 'stat /big' /var/tmp/st-fs-ext4.img
debugfs 1.47.2 (1-Jan-2025) Inode: 14 Type: regular Mode: 0664 Flags: 0x80000 Generation: 292513313 Version: 0x00000000:00000008 User: 1001 Group: 1001 Project: 0 Size: 16777216 File ACL: 0 Links: 1 Blockcount: 32768 Fragment: Address: 0 Number: 0 Size: 0 ctime: 0x6ab8e11c:9f11f434 -- Sun Sep 27 09:25:48 2026 atime: 0x6ab8e11c:9c729194 -- Sun Sep 27 09:25:48 2026 mtime: 0x6ab8e11c:9f11f434 -- Sun Sep 27 09:25:48 2026 crtime: 0x6ab8e11c:9c729194 -- Sun Sep 27 09:25:48 2026 Size of extra inode fields: 32 Inode checksum: 0xbdf92818 EXTENTS: (0-4095):38912-43007

debugfs reads the image directly. Inode 14 has flag 0x80000 (it uses extents), Blockcount in 512-byte sectors, timestamps to the nanosecond and a checksum. EXTENTS maps logical blocks 0 to 4095 of the file to physical blocks 38912 to 43007: one extent, a start and a length, where ext2 and ext3 kept a list of every single block.

deploy@web01 · Ubuntu 26.04 LTS
$ dd if=/dev/zero of=/mnt/st-fs-xfs/big bs=1M count=16 status=none sync xfs_bmap -v /mnt/st-fs-xfs/big stat -c "%n: inode %i" /mnt/st-fs-xfs/big /mnt/st-fs-xfs /mnt/st-fs-ext4
/mnt/st-fs-xfs/big: EXT: FILE-OFFSET BLOCK-RANGE AG AG-OFFSET TOTAL 0: [0..32767]: 192..32959 0 (192..32959) 32768 /mnt/st-fs-xfs/big: inode 131 /mnt/st-fs-xfs: inode 128 /mnt/st-fs-ext4: inode 2

xfs_bmap counts in 512-byte sectors, so [0..32767] is the whole 16 MiB file in one extent, with the allocation group it sits in and its offset inside that group. The root directory is inode 128 on this XFS filesystem and inode 2 on ext4; XFS inode numbers encode the allocation group and position of the inode.

Writeback and fsync

write() returns once the data is in the page cache. The kernel's flusher threads write dirty pages back when they are older than vm.dirty_expire_centisecs (3000, 30 seconds) or when too much memory is dirty (vm.dirty_background_ratio 10 and vm.dirty_ratio 20 percent). sync() writes back everything. fsync(fd) writes one file's data and metadata and waits until the device confirms them; fdatasync() skips metadata that is not needed to read the data back. Before relying on the sync command, check which call it makes.

deploy@web01 · Ubuntu 26.04 LTS
$ strace -e trace=sync,syncfs,fsync,fdatasync sync /mnt/st-fs-ext4/small strace -e trace=sync,syncfs,fsync,fdatasync gnusync /mnt/st-fs-ext4/small
sync() = 0 +++ exited with 0 +++ fsync(3) = 0 +++ exited with 0 +++

Ubuntu 26.04's Rust sync given a file name still calls sync(), flushing every filesystem; GNU sync (installed as gnusync) calls fsync() on that file. The Rust version's sync --data FILE calls fdatasync(). The next test uses dd conv=fsync, which calls fsync() on its output with either implementation. xfs_io -x -c shutdown then stops the filesystem as a power cut would: nothing more is written, not even the journal. It works on ext4 as well as XFS.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /mnt/st-fs-ext4 echo 'written, never synced' > unsynced echo 'written and fsynced' | dd of=synced conv=fsync status=none ls -l unsynced synced
-rw-rw-r-- 1 deploy deploy 20 Sep 27 09:25 synced -rw-rw-r-- 1 deploy deploy 22 Sep 27 09:25 unsynced
$ sudo xfs_io -x -c shutdown /mnt/st-fs-ext4 echo more >> /mnt/st-fs-ext4/synced
-bash: line 2: /mnt/st-fs-ext4/synced: Input/output error
$ journalctl -k -o cat -n 2
EXT4-fs (loop0): shut down requested (2) Aborting journal on device loop0-8.

Before the shutdown the two files looked the same. After it, any write fails with Input/output error, and the kernel log shows the requested shutdown and the aborted journal (loop0-8 names the journal by its device and the journal's inode, 8; the thread that writes it is jbd2/loop0-8). Unmount, mount again and look.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo umount /mnt/st-fs-ext4 sudo mount -o loop /var/tmp/st-fs-ext4.img /mnt/st-fs-ext4 ls -l /mnt/st-fs-ext4/unsynced /mnt/st-fs-ext4/synced cat /mnt/st-fs-ext4/synced
-rw-rw-r-- 1 deploy deploy 20 Sep 27 09:25 /mnt/st-fs-ext4/synced -rw-rw-r-- 1 deploy deploy 0 Sep 27 09:25 /mnt/st-fs-ext4/unsynced written and fsynced
$ journalctl -k -o cat -n 2
EXT4-fs (loop0): recovery complete EXT4-fs (loop0): mounted filesystem 03b6ef89-658d-4d37-8421-0696e764dbfa r/w with ordered data mode. Quota mode: none.

Mounting replayed the journal (recovery complete). synced has its contents. unsynced exists but is empty: its creation was committed to the journal when the fsync of synced flushed the running transaction, but its data was still waiting for delayed allocation in memory. That is the classic empty-file-after-a-crash pattern, and the fix is in the program: write a temporary file, fsync it, rename it over the old name, and fsync the directory.

Journaling modes

The journal records metadata changes before they are written in place, so after a crash the filesystem replays it at mount instead of scanning everything. ext4 has three modes, set with data= at mount. In ordered, the default, only metadata goes through the journal, but a file's data blocks are written before the metadata that points to them is committed, so a crash never exposes old contents from freed blocks. writeback drops that ordering: faster, and after a crash a file can contain stale data. journal sends data through the journal as well.

deploy@web01 · Ubuntu 26.04 LTS
$ dev=$(basename "$(findmnt -no SOURCE /mnt/st-fs-ext4)") grep data= /proc/fs/ext4/$dev/options before=$(cat /sys/fs/ext4/$dev/lifetime_write_kbytes) dd if=/dev/zero of=/mnt/st-fs-ext4/ordered bs=1M count=32 status=none && sync echo "$(( $(cat /sys/fs/ext4/$dev/lifetime_write_kbytes) - before )) KiB written for 32768 KiB of data"
data=ordered 32844 KiB written for 32768 KiB of data
$ sudo umount /mnt/st-fs-ext4 sudo mount -o loop,data=journal /var/tmp/st-fs-ext4.img /mnt/st-fs-ext4 dev=$(basename "$(findmnt -no SOURCE /mnt/st-fs-ext4)") grep data= /proc/fs/ext4/$dev/options before=$(cat /sys/fs/ext4/$dev/lifetime_write_kbytes) dd if=/dev/zero of=/mnt/st-fs-ext4/journalled bs=1M count=32 status=none && sync echo "$(( $(cat /sys/fs/ext4/$dev/lifetime_write_kbytes) - before )) KiB written for 32768 KiB of data"
data=journal 65884 KiB written for 32768 KiB of data

lifetime_write_kbytes counts what the filesystem sent to its device. In ordered mode 32 MiB of data cost about 32 MiB of writes plus a little metadata; in journal mode it cost twice that, because every block went to the journal first and to its place later. The ext4 documentation adds that data=journal disables delayed allocation and direct I/O. XFS journals metadata only and has no data= option; its log is the internal log in the mkfs.xfs output.

Checking and repairing

The journal covers crashes. Corruption from a failing disk, a bug or a stray write needs a checker, and the checker needs the filesystem unmounted. To have something to repair, fill a directory, unmount, find the directory's data block and overwrite it with zeros.

deploy@web01 · Ubuntu 26.04 LTS
$ mkdir /mnt/st-fs-ext4/docs for i in $(seq 20); do echo "report $i" > /mnt/st-fs-ext4/docs/report-$i.txt; done sudo umount /mnt/st-fs-ext4 debugfs -R "blocks /docs" /var/tmp/st-fs-ext4.img
debugfs 1.47.2 (1-Jan-2025) 21053
$ dd if=/dev/zero of=/var/tmp/st-fs-ext4.img bs=4096 seek=21053 count=1 conv=notrunc status=none
$ e2fsck -fn /var/tmp/st-fs-ext4.img
e2fsck 1.47.2 (1-Jan-2025) Pass 1: Checking inodes, blocks, and sizes Pass 2: Checking directory structure Directory inode 32769, block #0, offset 0: directory corrupted Salvage? no e2fsck: aborted st-fs-ext4: ********** WARNING: Filesystem still has errors **********

Always look before you repair. -n opens the filesystem read-only and answers no to every question; here it stopped at the corrupted directory, because it cannot go further without changing something, and exited 12: 4 (errors left uncorrected) plus 8 (operational error). -f forces a full check even when the superblock says the filesystem is clean.

A modifying repair cannot be undone, and the next run shows it throwing names away. On real data, take a copy of the device first, a block-level image or a snapshot, which is your only way back; e2image (ext4) and xfs_metadump (XFS) save just the metadata when a full copy is too large. Here the image file is the lab copy.

deploy@web01 · Ubuntu 26.04 LTS
$ e2fsck -fy /var/tmp/st-fs-ext4.img
e2fsck 1.47.2 (1-Jan-2025) Pass 1: Checking inodes, blocks, and sizes Pass 2: Checking directory structure Directory inode 32769, block #0, offset 0: directory corrupted Salvage? yes Missing '.' in directory inode 32769. Fix? yes Missing '..' in directory inode 32769. Fix? yes Pass 3: Checking directory connectivity '..' in /docs (32769) is <The NULL inode> (0), should be / (2). Fix? yes Pass 4: Checking reference counts Unattached inode 19 Connect to /lost+found? yes Inode 19 ref count is 2, should be 1. Fix? yes … Pass 5: Checking group summary information st-fs-ext4: ***** FILE SYSTEM WAS MODIFIED ***** st-fs-ext4: 39/65536 files (2.6% non-contiguous), 28803/65536 blocks

With -y e2fsck salvaged the directory, recreated its . and .. entries, and reconnected the twenty files it found with no name into lost+found; exit status 1 means errors were corrected. Run the check again until it is clean, then mount and look.

deploy@web01 · Ubuntu 26.04 LTS
$ e2fsck -fn /var/tmp/st-fs-ext4.img
e2fsck 1.47.2 (1-Jan-2025) Pass 1: Checking inodes, blocks, and sizes Pass 2: Checking directory structure Pass 3: Checking directory connectivity Pass 4: Checking reference counts Pass 5: Checking group summary information st-fs-ext4: 39/65536 files (2.6% non-contiguous), 28803/65536 blocks
$ sudo mount -o loop /var/tmp/st-fs-ext4.img /mnt/st-fs-ext4 ls /mnt/st-fs-ext4/docs sudo sh -c 'grep -H . /mnt/st-fs-ext4/lost+found/*' | head -4
/mnt/st-fs-ext4/lost+found/#19:report 1 /mnt/st-fs-ext4/lost+found/#20:report 2 /mnt/st-fs-ext4/lost+found/#21:report 3 /mnt/st-fs-ext4/lost+found/#22:report 4

docs is empty, and lost+found holds the files under names made from their inode numbers, contents intact. The names lived in the block you zeroed; the inodes and data did not.

XFS does not check itself at mount and has no working fsck: its repair tool is xfs_repair. It also refuses to run while the log holds changes. Simulate a crash again, then attach the image to a loop device yourself so the tools see a block device. losetup --find --show picks the first free loop device and prints its name; keep the name in a shell variable, loop, and use "$loop" in the commands that follow rather than typing a number, because the number depends on your machine. The later commands expect the same shell; in a new one, loop=$(losetup --noheadings --output NAME --associated /var/tmp/st-fs-xfs.img) finds the device again.

deploy@web01 · Ubuntu 26.04 LTS
$ mkdir /mnt/st-fs-xfs/docs for i in $(seq 20); do echo "report $i" > /mnt/st-fs-xfs/docs/report-$i.txt; done sync xfs_bmap -v /mnt/st-fs-xfs/docs
/mnt/st-fs-xfs/docs: EXT: FILE-OFFSET BLOCK-RANGE AG AG-OFFSET TOTAL 0: [0..7]: 120..127 0 (120..127) 8
$ sudo xfs_io -x -c shutdown /mnt/st-fs-xfs sudo umount /mnt/st-fs-xfs loop=$(sudo losetup --find --show /var/tmp/st-fs-xfs.img) echo "$loop"
/dev/loop1
$ sudo xfs_repair "$loop"
Phase 1 - find and verify superblock... Phase 2 - using internal log - zero log... ERROR: The filesystem has valuable metadata changes in a log which needs to be replayed. Mount the filesystem to replay the log, and unmount it before re-running xfs_repair. If the filesystem is a snapshot of a mounted filesystem, you may need to give mount the nouuid option. If you are unable to mount the filesystem, then use the -L option to destroy the log and attempt a repair. Note that destroying the log may cause corruption -- please attempt a mount of the filesystem before doing this.

xfs_repair stopped at the dirty log. Its advice is the right order: mount to replay the log, unmount, then repair. -L zeroes the log and throws away the metadata changes in it, so it is the last resort for a filesystem that will not mount.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo mount "$loop" /mnt/st-fs-xfs ls /mnt/st-fs-xfs/docs | wc -l sudo umount /mnt/st-fs-xfs
20
$ journalctl -k -o cat -n 3
XFS (loop1): Starting recovery (logdev: internal) XFS (loop1): Ending recovery (logdev: internal) XFS (loop1): Unmounting Filesystem 47e01c97-4e66-4bda-8c35-919da91305b0

The mount replayed the log and all twenty files are there. Now damage the directory block, whose sector range xfs_bmap showed above, and try the obvious tool first.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo dd if=/dev/zero of="$loop" bs=512 seek=120 count=8 status=none
$ sudo fsck "$loop"
fsck from util-linux 2.41.3 If you wish to check the consistency of an XFS filesystem or repair a damaged filesystem, see xfs_repair(8).

fsck succeeded without looking at anything. fsck.xfs is a script that runs xfs_repair only when called with -f from a non-interactive boot (fsck.mode=force on the kernel command line); otherwise it prints this pointer and exits 0. A clean fsck exit on XFS proves nothing.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo xfs_repair -n "$loop"
Phase 1 - find and verify superblock... Phase 2 - using internal log - zero log... - scan filesystem freespace and inode maps... - found root inode chunk Phase 3 - for each AG... - scan (but don't clear) agi unlinked lists... - process known inodes and perform inode discovery... - agno = 0 Metadata CRC error detected at 0xb6e49eee8c54, xfs_dir3_block block 0x78/0x1000 bad directory block magic # 0 in block 0 for directory inode 132 corrupt block 0 in directory inode 132 would junk block corrupt directory block 0 for inode 132 no . entry for directory 132 no .. entry for directory 132 problem with directory contents in inode 132 would have cleared inode 132 … entry "docs" in shortform directory 128 references free inode 132 would have junked entry "docs" in directory inode 128 … No modify flag set, skipping phase 5 Phase 6 - check inode connectivity... - traversing filesystem ... entry "docs" in shortform directory inode 128 points to free inode 132, would junk entry - traversal finished ... - moving disconnected inodes to lost+found ... disconnected inode 133, would move to lost+found … Phase 7 - verify link counts... would have reset inode 128 nlinks from 3 to 2 No modify flag set, skipping filesystem flush and exiting.
$ sudo xfs_repair "$loop"
… - agno = 0 Metadata CRC error detected at 0xc2417d338c54, xfs_dir3_block block 0x78/0x1000 bad directory block magic # 0 in block 0 for directory inode 132 corrupt block 0 in directory inode 132 will junk block corrupt directory block 0 for inode 132 no . entry for directory 132 no .. entry for directory 132 problem with directory contents in inode 132 cleared inode 132 … entry "docs" in shortform directory 128 references free inode 132 junking entry "docs" in directory inode 128 … - moving disconnected inodes to lost+found ... disconnected inode 133, moving to lost+found … Phase 7 - verify and correct link counts... resetting inode 128 nlinks from 4 to 3 done
$ sudo mount "$loop" /mnt/st-fs-xfs ls /mnt/st-fs-xfs sudo ls /mnt/st-fs-xfs/lost+found | head -4 sudo cat /mnt/st-fs-xfs/lost+found/$(sudo ls /mnt/st-fs-xfs/lost+found | head -1)
big lost+found 133 134 135 136 report 1

xfs_repair -n found the bad checksum and the zeroed directory block and listed what it would do, exiting 1. The real run did it, and chose differently from e2fsck: it cleared the damaged directory inode and removed the docs entry from the root, then moved the disconnected files into lost+found under their inode numbers. Either way the data survives and the names do not.

The root filesystem cannot be unmounted on a running system. To check it, add fsck.mode=force to the kernel command line for one boot, from the boot menu (the boot lesson covers the kernel command line). systemd-fsck then checks the root filesystem during boot. On ext4 the default fsck.repair=preen fixes only what is safe to fix without asking, and fsck.repair=yes answers yes to every question unattended, so read the result of a preen boot in the journal before you consider that. On XFS the forced fsck.xfs runs a full xfs_repair straight away, with no read-only pass, so the copy of the disk comes first there. The boot takes as long as the check, which on a multi-terabyte filesystem can be hours with only the console to watch. For anything bigger than a quick check, boot rescue media and repair the unmounted device by hand.

Take the lab apart: unmount both filesystems, detach the loop device you attached yourself (the ones mount -o loop made went away at unmount), and delete the images.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo umount /mnt/st-fs-ext4 /mnt/st-fs-xfs sudo losetup --detach "$loop" sudo rmdir /mnt/st-fs-ext4 /mnt/st-fs-xfs rm /var/tmp/st-fs-ext4.img /var/tmp/st-fs-xfs.img

Try this

Make a 64 MiB ext4 image with only 64 inodes (mkfs.ext4 -q -N 64), mount it and create empty files in a loop until one fails. Predict first: touch fails at file 53 with No space left on device, df -h still shows the filesystem 1% used, and df -i shows all 64 inodes used: your 52 files and 12 the filesystem uses itself (reserved inodes and lost+found among them). An XFS filesystem would keep creating inodes until its space or imaxpct ran out. Clean up with umount, rmdir and rm.

Takeaway

Treat data as durable only after fsync has returned, and treat a filesystem check as a sequence: look with -n, copy the device, repair it unmounted (replaying an XFS log by mounting first), and check again until it is clean.

Quick check
01After a power cut, an application's settings file on ext4 exists but is 0 bytes. The application writes the new settings straight into the file and closes it. Which change prevents this?
Incorrect — data=writeback removes the ordering between data and metadata; it makes stale or missing data after a crash more likely, not less.
Incorrect — An hourly flush still leaves up to an hour of changes in memory, and the file can be caught between writes whenever the power fails.
Incorrect — A shorter expiry narrows the window but guarantees nothing; only fsync() waits for the device to confirm the data.
Correct — The fsync makes the new data durable before the rename switches names, so after a crash the name points to either the old or the new complete file.
02xfs_repair on an unmounted XFS device stops with "The filesystem has valuable metadata changes in a log which needs to be replayed". What do you do first?
Incorrect — -L zeroes the log and discards the changes in it; it is the last resort when the filesystem cannot be mounted.
Correct — Mounting replays the committed changes from the log; after a clean unmount xfs_repair can check the filesystem in its current state.
Incorrect — fsck.xfs does not replay anything; without -f at boot it just prints a pointer to xfs_repair and exits 0.
Incorrect — e2fsck checks ext2/3/4 filesystems only and cannot read an XFS log.
03A team plans to mount a write-heavy database volume with data=journal "for extra safety". What measurable cost should they expect?
Incorrect — Reads come from the page cache or the data blocks; the journal is written, not consulted, during normal operation.
Incorrect — A full journal forces a checkpoint that writes committed blocks to their places; it slows writes but does not switch the filesystem read-only.
Correct — In data=journal mode each block is written to the journal and later to its place, as lifetime_write_kbytes showed; delayed allocation and direct I/O are also disabled.
Incorrect — The default ordered mode journals metadata only and orders data writes before it; data=journal is what adds the data blocks to the journal.

Related