Files, descriptors and the VFS

fd tables, open files, inodes and limits.

Advanced14 min · lesson 11 of 21

"Too many open files" on a machine with a hard limit of half a million descriptors, a disk that stays full after the big log was deleted, a log file that fills with zero bytes after rotation, a port whose owner ss does not show: each of these is explained by how the kernel's virtual filesystem layer (the VFS) represents an open file. This lesson follows a file descriptor down through the objects the VFS keeps, shows what dup and fork share and what exec passes on, recovers a deleted file that a process still holds, maps a socket to its process through its inode, and produces the errors behind each file limit in bounded processes. Linux essentials introduced /proc/PID/fd and lsof +L1; this is the layer underneath them.

From a descriptor to the filesystem

A file descriptor is a small integer, an index into the process's file descriptor table. Each entry points to an open file description, the object that open() creates: it holds the current offset, the access mode and status flags such as O_APPEND. The description points to a dentry, a cached directory entry that ties one name in one directory to an inode. The inode is the file itself: owner, mode, size and where its data lives. Every inode belongs to a superblock, the kernel's instance of one mounted filesystem, and a mount places that filesystem in the directory tree.

What an open file descriptor points to
1fd 3 in the process
an index into its fd table
2Open file description
offset and flags; one per open() call
3dentry
a name in a directory, cached
4inode
the file: owner, mode, size, data blocks
5superblock and mount
the filesystem and its place in the tree
dup() and fork() copy the pointer from the fd to the description; open() makes a new description.

The examples open files on numbered descriptors in the current shell, which the essentials course did not need. exec 3< FILE opens FILE for reading on descriptor 3 of the shell itself and keeps it open for the following commands; exec 4<&3 makes descriptor 4 a copy of descriptor 3; <&3 gives a single command descriptor 3 as its standard input. Open /etc/passwd on descriptor 3 in a shell and read what the kernel reports about it.

deploy@web01 · Ubuntu 26.04 LTS
$ exec 3< /etc/passwd ls -l /proc/$$/fd cat /proc/$$/fdinfo/3
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:25 0 -> pipe:[765972] l-wx------ 1 deploy deploy 64 Sep 27 09:25 1 -> pipe:[765091] l-wx------ 1 deploy deploy 64 Sep 27 09:25 2 -> pipe:[765091] lr-x------ 1 deploy deploy 64 Sep 27 09:25 3 -> /etc/passwd pos: 0 flags: 0400000 mnt_id: 49 ino: 4113

/proc/$$/fd is the shell's table. Descriptors 0 to 2 are pipes because the lab runs these commands non-interactively; in an SSH session they point to /dev/pts/N. fdinfo/3 shows the description behind descriptor 3: offset 0, since nothing has been read, and the flags in octal. 0400000 is O_RDONLY (0) plus O_LARGEFILE, which the kernel adds on 64-bit systems; that bit is 0400000 on aarch64 and 0100000 on x86_64. mnt_id and ino identify the mount and the inode.

deploy@web01 · Ubuntu 26.04 LTS
$ findmnt -o ID,TARGET,SOURCE,FSTYPE --target /etc/passwd stat -c 'inode %i, %h link(s), %s bytes' /etc/passwd
ID TARGET SOURCE FSTYPE 49 / /dev/vda1 ext4 inode 4113, 1 link(s), 1890 bytes
$ stat -f /etc/passwd
File: "/etc/passwd" ID: 7e9f9a745a795e20 Namelen: 255 Type: ext2/ext3 Block Size: 4096 Fundamental block size: 4096 Blocks: Total: 5821056 Free: 4723578 Available: 4719482 Inodes: Total: 3014656 Free: 2853862

findmnt shows mount 49 is the root filesystem on /dev/vda1, and stat shows the same inode number with one link (one name). stat -f calls statfs(), which reports counters the superblock keeps for the whole filesystem: block size, total and free blocks, total and free inodes. ext2/ext3 is the name stat gives the magic number that ext4 also uses.

Every path lookup walks dentries, and the kernel caches them, including negative dentries that record that a name does not exist, so a repeated failed lookup (a PATH search, a program probing for optional configuration files) costs almost nothing.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/sys/fs/dentry-state python3 -c 'import os; [os.path.exists(f"/var/tmp/missing-{os.getpid()}-{i}") for i in range(100000)]' cat /proc/sys/fs/dentry-state
257873 249668 45 0 17010 0 357873 349668 45 0 117010 0

The fields are all dentries, the unused ones kept for reuse, the age after which unused entries may be reclaimed, a flag for a pending shrink, and the negative dentries (unused entries for names that do not exist). Looking up 100,000 names that do not exist added 100,000 negative dentries. They count as reclaimable slab memory (SReclaimable in /proc/meminfo) and the kernel frees them under memory pressure, but a program that probes millions of unique missing names can grow the cache to gigabytes, which looks like a leak in free.

Offsets live in the open file description

Every open() creates a new description. dup() (and the shell's 4<&3) creates a second descriptor for an existing description, and so does fork() for every descriptor the child inherits.

deploy@web01 · Ubuntu 26.04 LTS
$ exec 3< /etc/passwd exec 4<&3 exec 5< /etc/passwd read -r line <&3 grep pos /proc/$$/fdinfo/3 /proc/$$/fdinfo/4 /proc/$$/fdinfo/5
/proc/219181/fdinfo/3:pos: 32 /proc/219181/fdinfo/4:pos: 32 /proc/219181/fdinfo/5:pos: 0
$ exec 3< /etc/passwd sh -c 'read -r line; echo "child: $line"' <&3 read -r line <&3; echo "parent: $line"
child: root:x:0:0:root:/root:/bin/bash parent: daemon:x:1:1:daemon:/usr/sbin:/usr/sbin/nologin

Reading one line through descriptor 3 moved descriptor 4 as well, to offset 32, because they share one description; descriptor 5 came from its own open() and is still at 0. In the second example the child shell inherited descriptor 3 as its standard input and read the first line, and the parent carried on from the second line: the offset belongs to the description, not to either process. That is why (cmd1; cmd2) > file puts both outputs one after the other.

The same sharing explains a classic log problem. A service writes to a log opened without O_APPEND (a plain > redirection), and a rotation tool copies the log and truncates it in place, as logrotate's copytruncate does.

deploy@web01 · Ubuntu 26.04 LTS
$ (echo 'first line'; echo 'second line'; sleep 2; echo 'after truncate') > /var/tmp/k-vfs.log & sleep 1; truncate -s 0 /var/tmp/k-vfs.log; wait ls -l /var/tmp/k-vfs.log od -c /var/tmp/k-vfs.log
-rw-rw-r-- 1 deploy deploy 38 Sep 27 09:25 /var/tmp/k-vfs.log 0000000 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 \0 0000020 \0 \0 \0 \0 \0 \0 \0 a f t e r t r u 0000040 n c a t e \n 0000046

truncate set the file's size to 0, but the writer's description still had its offset at 23, just after the two lines. Its next write landed at offset 23, and the kernel reads the gap before it as zero bytes: a 38-byte file of 23 NULs and one line. On a busy service the gap is gigabytes, the file is sparse, and ls -l shows its old size. With O_APPEND every write goes to the current end of the file instead, which is what >> and most logging libraries use.

What a child inherits: close-on-exec

execve() keeps every open descriptor unless it is marked close-on-exec (FD_CLOEXEC, set at open time with O_CLOEXEC). The mark belongs to the descriptor, not to the description.

deploy@web01 · Ubuntu 26.04 LTS
$ exec 3< /etc/passwd ls -l /proc/self/fd
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:25 0 -> pipe:[765972] l-wx------ 1 deploy deploy 64 Sep 27 09:25 1 -> pipe:[766294] l-wx------ 1 deploy deploy 64 Sep 27 09:25 2 -> pipe:[766294] lr-x------ 1 deploy deploy 64 Sep 27 09:25 3 -> /etc/passwd lr-x------ 1 deploy deploy 64 Sep 27 09:25 4 -> /proc/219256/fd
$ exec 3< /etc/passwd grep flags /proc/$$/fdinfo/3 python3 -c 'f = open("/etc/passwd"); print(open(f"/proc/self/fdinfo/{f.fileno()}").read(), end="")'
flags: 0400000 pos: 0 flags: 02400000 mnt_id: 49 ino: 4113

ls inherited descriptor 3 from the shell; descriptor 4 is the directory ls opened itself. Bash redirections do not set close-on-exec, while Python's open() does by default, which fdinfo shows as 02000000 added to the flags (02400000). A descriptor leaked across exec has real effects: a helper program that inherited a log keeps its space allocated after the log is deleted, and a helper that inherited a listening socket keeps the port open after its parent daemon restarts, so the new daemon fails with "Address already in use".

Deleted but still open

rm removes a name. The inode and its data stay allocated while any open file description refers to them, and they are freed when the last one is closed. To see it, fill a 64 MiB log and follow it with tail -f in a transient service, then delete the log.

deploy@web01 · Ubuntu 26.04 LTS
$ yes 'k-vfs log line' | head -c 64M > /var/tmp/k-vfs-held.log sudo systemd-run --unit=k-vfs-tail --uid=$USER tail -f /var/tmp/k-vfs-held.log
Running as unit: k-vfs-tail.service; invocation ID: 3dc56344bf914f1c83fcfe3fc5ce79cd
$ rm /var/tmp/k-vfs-held.log ls -l /proc/$(systemctl show -P MainPID k-vfs-tail)/fd
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:25 0 -> /dev/null lrwx------+ 1 deploy deploy 64 Sep 27 09:25 1 -> socket:[766458] lrwx------+ 1 deploy deploy 64 Sep 27 09:25 2 -> socket:[766458] lr-x------ 1 deploy deploy 64 Sep 27 09:25 3 -> anon_inode:inotify lrwx------ 1 deploy deploy 64 Sep 27 09:25 4 -> anon_inode:[eventpoll] lrwx------ 1 deploy deploy 64 Sep 27 09:25 5 -> anon_inode:[eventfd] lr-x------ 1 deploy deploy 64 Sep 27 09:25 6 -> /var/tmp/k-vfs-held.log (deleted) lr-x------ 1 deploy deploy 64 Sep 27 09:25 7 -> pipe:[766469] l-wx------ 1 deploy deploy 64 Sep 27 09:25 8 -> pipe:[766469]

Descriptor 6 still reaches the deleted log. The rest of tail's table shows that a descriptor can point to things that are not files on a disk: sockets (standard output and error go to the journal), an inotify instance, an epoll set, an eventfd and a pipe, all backed by inodes on the kernel's internal filesystems. The + that Ubuntu's Rust ls prints on socket entries marks an extended attribute (sockets expose system.sockprotoname); GNU ls shows + only for ACLs.

deploy@web01 · Ubuntu 26.04 LTS
$ cd /proc/$(systemctl show -P MainPID k-vfs-tail) grep -E '^(pos|flags|ino|inotify)' fdinfo/6 fdinfo/3 printf 'inode of /var/tmp in hex: %x\n' $(stat -c %i /var/tmp)
fdinfo/6:pos: 67108864 fdinfo/6:flags: 02400000 fdinfo/6:ino: 3508 fdinfo/3:pos: 0 fdinfo/3:flags: 02004000 fdinfo/3:ino: 64 fdinfo/3:inotify wd:1 ino:1a21e sdev:fd00001 mask:fee ignored_mask:0 fhandle-bytes:8 fhandle-type:1 f_handle:1ea201009c9339fb inode of /var/tmp in hex: 1a21e

tail has read to the end, offset 67108864, and its descriptors are close-on-exec. The inotify descriptor lists each watch: wd:1 watches inode 1a21e (hexadecimal) on device fd00001 (major 253, minor 1: vda1), which is /var/tmp itself, so this tail learns about changes to the file's name from its directory.

deploy@web01 · Ubuntu 26.04 LTS
$ stat -L -c '%h link(s), %s bytes' /proc/$(systemctl show -P MainPID k-vfs-tail)/fd/6
0 link(s), 67108864 bytes
$ cp /proc/$(systemctl show -P MainPID k-vfs-tail)/fd/6 /var/tmp/k-vfs-copy.log ls -l /var/tmp/k-vfs-copy.log
-rw-rw-r-- 1 deploy deploy 67108864 Sep 27 09:25 /var/tmp/k-vfs-copy.log
$ : > /proc/$(systemctl show -P MainPID k-vfs-tail)/fd/6 stat -L -c '%h link(s), %s bytes' /proc/$(systemctl show -P MainPID k-vfs-tail)/fd/6
0 link(s), 0 bytes

stat -L follows the /proc link to the inode: no names left, 64 MiB still allocated. Opening /proc/PID/fd/6 opens that same inode again, so cp recovers the whole file, and : > truncates it in place, which frees the space without restarting the process. Both need the kernel's ptrace read-access check on the process (the same owner, or root) plus permission on the file itself, which is why they worked here as the file's owner. That check is weaker than the one for attaching a debugger, which Ubuntu's Yama module restricts further. lsof +L1 finds such files across the system, and the disk I/O lesson shows how they make df and du disagree.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop k-vfs-tail rm /var/tmp/k-vfs-copy.log

Sockets are inodes too

A socket is an open file description with an inode on the kernel's socket filesystem, and that inode number is what ties the network tables to processes. Start a small listener and follow it.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-vfs-http --uid=$USER python3 -m http.server --bind 127.0.0.1 8471
Running as unit: k-vfs-http.service; invocation ID: 5eaf0bc665dc42f5a70129f5f09a54db
$ ss -tlnpe 'sport = :8471'
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess LISTEN 0 5 127.0.0.1:8471 0.0.0.0:* users:(("python3",pid=219513,fd=3)) uid:1001 ino:766758 sk:10c8 cgroup:/system.slice/k-vfs-http.service <->
$ ls -l /proc/$(systemctl show -P MainPID k-vfs-http)/fd
total 0 lr-x------ 1 deploy deploy 64 Sep 27 09:25 0 -> /dev/null lrwx------+ 1 deploy deploy 64 Sep 27 09:25 1 -> socket:[766755] lrwx------+ 1 deploy deploy 64 Sep 27 09:25 2 -> socket:[766755] lrwx------+ 1 deploy deploy 64 Sep 27 09:25 3 -> socket:[766758]
$ grep -i ':2117 ' /proc/net/tcp
1: 0100007F:2117 00000000:0000 0A 00000000:00000000 00:00000000 00000000 1001 0 766758 1 0000000000000000 100 0 0 10 0

ss -e prints the socket's inode (ino:766758) and owner UID, and -p names the process, which ss finds by scanning /proc/*/fd for socket:[766758]: descriptor 3 of python3. As an ordinary user ss -p can read only your own processes' descriptors, so a socket that shows no process may simply belong to another user; run it with sudo before concluding that nothing owns it. /proc/net/tcp is the raw table ss summarises: 0100007F:2117 is 127.0.0.1 port 8471 in hexadecimal (the address in the machine's byte order), state 0A is LISTEN, then the UID and the inode.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop k-vfs-http

Running out: EMFILE, ENFILE and inotify

Three limits decide how many files can be open. RLIMIT_NOFILE is per process, with a soft limit the process may raise up to its hard limit; exceeding it fails with EMFILE, "Too many open files". fs.file-max caps open file descriptions system-wide (fs.file-nr shows how many exist); exceeding it fails with ENFILE, "Too many open files in system". fs.nr_open is the highest value RLIMIT_NOFILE can be set to.

deploy@web01 · Ubuntu 26.04 LTS
$ ulimit -Sn; ulimit -Hn cat /proc/sys/fs/file-nr /proc/sys/fs/file-max /proc/sys/fs/nr_open
1024 524288 1125 0 9223372036854775807 9223372036854775807 2147483584

The soft limit is 1024 and the hard limit 524288, systemd's defaults for services and login sessions. The soft limit stays at 1024 because select(), the oldest system call for waiting on several descriptors at once, uses a fixed-size bitmap and cannot handle descriptors above 1023; the systemd.exec manual advises programs that do not use select() to raise their own soft limit to the hard one at startup, and LimitNOFILE= sets it for a unit (the tuning lesson covers when). Since version 240, systemd raises fs.file-max and fs.nr_open to their maximums at boot, because cgroup memory accounting already charges open files to whoever opened them. So on these systems ENFILE is practically out of reach (and processes with CAP_SYS_ADMIN are exempt from it anyway); the error you will meet is EMFILE.

deploy@web01 · Ubuntu 26.04 LTS
$ ( ulimit -Sn 32 python3 -c ' import os fds = [] try: while True: fds.append(os.open("/etc/hostname", os.O_RDONLY)) except OSError as err: print(len(fds), "opened, then:", err)' )
29 opened, then: [Errno 24] Too many open files: '/etc/hostname'

The subshell lowered only its own soft limit to 32, so nothing else was affected. Python opened 29 files, because descriptors 0 to 2 were already in use, and the 30th open() failed with errno 24. On a real service, count its descriptors over time (ls /proc/PID/fd | wc -l) before raising anything: a count that climbs with connections is load, a count that climbs forever is a leak, and a higher limit only delays a leak.

inotify, the interface programs use to watch files for changes, has limits of its own, per user rather than per process, and one of them fails with the same message. The test uses inotifywait from the inotify-tools package, which a default Ubuntu Server does not include: sudo apt install inotify-tools (RHEL 10's own repositories do not carry it).

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/sys/fs/inotify/max_user_instances /proc/sys/fs/inotify/max_user_watches
128 30890
$ for i in $(seq 130); do inotifywait -qq -m /var/tmp & done; sleep 2 pgrep -c -u $USER -x inotifywait pkill -u $USER -x inotifywait
Couldn't initialize inotify: Too many open files Try increasing the value of /proc/sys/fs/inotify/max_user_instances Couldn't initialize inotify: Too many open files Try increasing the value of /proc/sys/fs/inotify/max_user_instances 128

Each inotifywait creates one inotify instance. 128 of them started; the other two failed in inotify_init() with EMFILE, "Too many open files", although each process had only a handful of descriptors open. The watch limit, sized by the kernel to about 1% of memory (30,890 watches on this 4 GiB machine), fails differently: inotify_add_watch() returns ENOSPC, "No space left on device", on a disk with plenty of space. When either message does not match the descriptor count or df, check inotify. sudo find /proc/*/fd -lname anon_inode:inotify lists every instance and the process that holds it.

Try this

Predict, then run: repeat the truncation example with >> instead of > and a new file name, /var/tmp/k-vfs-append.log. Because the writer's description now has O_APPEND, its write after the truncation goes to the new end of the file: ls -l shows 15 bytes and od -c shows only after truncate and a newline, no NUL bytes. Remove both files afterwards.

Takeaway

When a file-related error does not make sense, go down the chain: the descriptor table and fdinfo for offsets, flags and inherited descriptors, the inode for space that a deleted name still holds, and the specific limit (per process, system-wide or inotify) that the error number points to.

Quick check
01After logrotate runs with copytruncate, an application's log shows 12 GB in ls -l but du reports a few megabytes, and the start of the file is all zero bytes. What happened?
Incorrect — copytruncate copies the log and truncates the original; it never rewrites the original's contents with compressed data.
Correct — The offset lives in the application's open file description. After truncation its next write lands at the old offset, leaving a sparse gap that reads as zeros.
Incorrect — ls -l shows the file's size, not its allocated blocks; du shows the blocks are freed, which is why they disagree.
Incorrect — A deleted-but-open file is a separate inode with no name; it cannot change the size of the new file's inode.
02A file-sync agent logs "Too many open files" and stops watching directories. /proc/PID/fd shows 40 descriptors and the soft limit is 1024. What should you check next?
Incorrect — A system-wide shortage fails with ENFILE, "Too many open files in system", and systemd raises file-max to its maximum.
Incorrect — The soft limit is the one enforced; 40 descriptors are far below 1024 either way.
Incorrect — A full inode table makes file creation fail with ENOSPC, "No space left on device", not EMFILE.
Correct — inotify_init() fails with EMFILE when the user has reached fs.inotify.max_user_instances (128 by default), regardless of how few descriptors the process has.
03Someone deleted /etc/app/app.conf by mistake. The running service still holds it: ls -l /proc/PID/fd shows "4 -> /etc/app/app.conf (deleted)". What is the quickest safe way to get the file back?
Correct — Opening /proc/PID/fd/4 reaches the same inode, and cp opens its own description, so it reads the whole file from offset 0 whatever the service's offset is.
Incorrect — Services read their configuration, they do not write it back; a restart closes the last reference and the inode and its data are freed.
Incorrect — The kernel refuses to link an inode whose link count has reached zero (O_TMPFILE files excepted), so the only way back is to copy the data.
Incorrect — e2fsck cannot check a mounted root filesystem, and it frees orphaned inodes rather than reconnecting them.

Related