Files, descriptors and the VFS
fd tables, open files, inodes and limits.
"Too many open files" on a machine with a hard limit of half a million descriptors, a disk that stays full after the big log was deleted, a log file that fills with zero bytes after rotation, a port whose owner ss does not show: each of these is explained by how the kernel's virtual filesystem layer (the VFS) represents an open file. This lesson follows a file descriptor down through the objects the VFS keeps, shows what dup and fork share and what exec passes on, recovers a deleted file that a process still holds, maps a socket to its process through its inode, and produces the errors behind each file limit in bounded processes. Linux essentials introduced /proc/PID/fd and lsof +L1; this is the layer underneath them.
From a descriptor to the filesystem
A file descriptor is a small integer, an index into the process's file descriptor table. Each entry points to an open file description, the object that open() creates: it holds the current offset, the access mode and status flags such as O_APPEND. The description points to a dentry, a cached directory entry that ties one name in one directory to an inode. The inode is the file itself: owner, mode, size and where its data lives. Every inode belongs to a superblock, the kernel's instance of one mounted filesystem, and a mount places that filesystem in the directory tree.
The examples open files on numbered descriptors in the current shell, which the essentials course did not need. exec 3< FILE opens FILE for reading on descriptor 3 of the shell itself and keeps it open for the following commands; exec 4<&3 makes descriptor 4 a copy of descriptor 3; <&3 gives a single command descriptor 3 as its standard input. Open /etc/passwd on descriptor 3 in a shell and read what the kernel reports about it.
/proc/$$/fd is the shell's table. Descriptors 0 to 2 are pipes because the lab runs these commands non-interactively; in an SSH session they point to /dev/pts/N. fdinfo/3 shows the description behind descriptor 3: offset 0, since nothing has been read, and the flags in octal. 0400000 is O_RDONLY (0) plus O_LARGEFILE, which the kernel adds on 64-bit systems; that bit is 0400000 on aarch64 and 0100000 on x86_64. mnt_id and ino identify the mount and the inode.
findmnt shows mount 49 is the root filesystem on /dev/vda1, and stat shows the same inode number with one link (one name). stat -f calls statfs(), which reports counters the superblock keeps for the whole filesystem: block size, total and free blocks, total and free inodes. ext2/ext3 is the name stat gives the magic number that ext4 also uses.
Every path lookup walks dentries, and the kernel caches them, including negative dentries that record that a name does not exist, so a repeated failed lookup (a PATH search, a program probing for optional configuration files) costs almost nothing.
The fields are all dentries, the unused ones kept for reuse, the age after which unused entries may be reclaimed, a flag for a pending shrink, and the negative dentries (unused entries for names that do not exist). Looking up 100,000 names that do not exist added 100,000 negative dentries. They count as reclaimable slab memory (SReclaimable in /proc/meminfo) and the kernel frees them under memory pressure, but a program that probes millions of unique missing names can grow the cache to gigabytes, which looks like a leak in free.
Offsets live in the open file description
Every open() creates a new description. dup() (and the shell's 4<&3) creates a second descriptor for an existing description, and so does fork() for every descriptor the child inherits.
Reading one line through descriptor 3 moved descriptor 4 as well, to offset 32, because they share one description; descriptor 5 came from its own open() and is still at 0. In the second example the child shell inherited descriptor 3 as its standard input and read the first line, and the parent carried on from the second line: the offset belongs to the description, not to either process. That is why (cmd1; cmd2) > file puts both outputs one after the other.
The same sharing explains a classic log problem. A service writes to a log opened without O_APPEND (a plain > redirection), and a rotation tool copies the log and truncates it in place, as logrotate's copytruncate does.
truncate set the file's size to 0, but the writer's description still had its offset at 23, just after the two lines. Its next write landed at offset 23, and the kernel reads the gap before it as zero bytes: a 38-byte file of 23 NULs and one line. On a busy service the gap is gigabytes, the file is sparse, and ls -l shows its old size. With O_APPEND every write goes to the current end of the file instead, which is what >> and most logging libraries use.
What a child inherits: close-on-exec
execve() keeps every open descriptor unless it is marked close-on-exec (FD_CLOEXEC, set at open time with O_CLOEXEC). The mark belongs to the descriptor, not to the description.
ls inherited descriptor 3 from the shell; descriptor 4 is the directory ls opened itself. Bash redirections do not set close-on-exec, while Python's open() does by default, which fdinfo shows as 02000000 added to the flags (02400000). A descriptor leaked across exec has real effects: a helper program that inherited a log keeps its space allocated after the log is deleted, and a helper that inherited a listening socket keeps the port open after its parent daemon restarts, so the new daemon fails with "Address already in use".
Deleted but still open
rm removes a name. The inode and its data stay allocated while any open file description refers to them, and they are freed when the last one is closed. To see it, fill a 64 MiB log and follow it with tail -f in a transient service, then delete the log.
Descriptor 6 still reaches the deleted log. The rest of tail's table shows that a descriptor can point to things that are not files on a disk: sockets (standard output and error go to the journal), an inotify instance, an epoll set, an eventfd and a pipe, all backed by inodes on the kernel's internal filesystems. The + that Ubuntu's Rust ls prints on socket entries marks an extended attribute (sockets expose system.sockprotoname); GNU ls shows + only for ACLs.
tail has read to the end, offset 67108864, and its descriptors are close-on-exec. The inotify descriptor lists each watch: wd:1 watches inode 1a21e (hexadecimal) on device fd00001 (major 253, minor 1: vda1), which is /var/tmp itself, so this tail learns about changes to the file's name from its directory.
stat -L follows the /proc link to the inode: no names left, 64 MiB still allocated. Opening /proc/PID/fd/6 opens that same inode again, so cp recovers the whole file, and : > truncates it in place, which frees the space without restarting the process. Both need the kernel's ptrace read-access check on the process (the same owner, or root) plus permission on the file itself, which is why they worked here as the file's owner. That check is weaker than the one for attaching a debugger, which Ubuntu's Yama module restricts further. lsof +L1 finds such files across the system, and the disk I/O lesson shows how they make df and du disagree.
Sockets are inodes too
A socket is an open file description with an inode on the kernel's socket filesystem, and that inode number is what ties the network tables to processes. Start a small listener and follow it.
ss -e prints the socket's inode (ino:766758) and owner UID, and -p names the process, which ss finds by scanning /proc/*/fd for socket:[766758]: descriptor 3 of python3. As an ordinary user ss -p can read only your own processes' descriptors, so a socket that shows no process may simply belong to another user; run it with sudo before concluding that nothing owns it. /proc/net/tcp is the raw table ss summarises: 0100007F:2117 is 127.0.0.1 port 8471 in hexadecimal (the address in the machine's byte order), state 0A is LISTEN, then the UID and the inode.
Running out: EMFILE, ENFILE and inotify
Three limits decide how many files can be open. RLIMIT_NOFILE is per process, with a soft limit the process may raise up to its hard limit; exceeding it fails with EMFILE, "Too many open files". fs.file-max caps open file descriptions system-wide (fs.file-nr shows how many exist); exceeding it fails with ENFILE, "Too many open files in system". fs.nr_open is the highest value RLIMIT_NOFILE can be set to.
The soft limit is 1024 and the hard limit 524288, systemd's defaults for services and login sessions. The soft limit stays at 1024 because select(), the oldest system call for waiting on several descriptors at once, uses a fixed-size bitmap and cannot handle descriptors above 1023; the systemd.exec manual advises programs that do not use select() to raise their own soft limit to the hard one at startup, and LimitNOFILE= sets it for a unit (the tuning lesson covers when). Since version 240, systemd raises fs.file-max and fs.nr_open to their maximums at boot, because cgroup memory accounting already charges open files to whoever opened them. So on these systems ENFILE is practically out of reach (and processes with CAP_SYS_ADMIN are exempt from it anyway); the error you will meet is EMFILE.
The subshell lowered only its own soft limit to 32, so nothing else was affected. Python opened 29 files, because descriptors 0 to 2 were already in use, and the 30th open() failed with errno 24. On a real service, count its descriptors over time (ls /proc/PID/fd | wc -l) before raising anything: a count that climbs with connections is load, a count that climbs forever is a leak, and a higher limit only delays a leak.
inotify, the interface programs use to watch files for changes, has limits of its own, per user rather than per process, and one of them fails with the same message. The test uses inotifywait from the inotify-tools package, which a default Ubuntu Server does not include: sudo apt install inotify-tools (RHEL 10's own repositories do not carry it).
Each inotifywait creates one inotify instance. 128 of them started; the other two failed in inotify_init() with EMFILE, "Too many open files", although each process had only a handful of descriptors open. The watch limit, sized by the kernel to about 1% of memory (30,890 watches on this 4 GiB machine), fails differently: inotify_add_watch() returns ENOSPC, "No space left on device", on a disk with plenty of space. When either message does not match the descriptor count or df, check inotify. sudo find /proc/*/fd -lname anon_inode:inotify lists every instance and the process that holds it.
Try this
Predict, then run: repeat the truncation example with >> instead of > and a new file name, /var/tmp/k-vfs-append.log. Because the writer's description now has O_APPEND, its write after the truncation goes to the new end of the file: ls -l shows 15 bytes and od -c shows only after truncate and a newline, no NUL bytes. Remove both files afterwards.
Takeaway
When a file-related error does not make sense, go down the chain: the descriptor table and fdinfo for offsets, flags and inherited descriptors, the inode for space that a deleted name still holds, and the specific limit (per process, system-wide or inotify) that the error number points to.