Processes: fork, exec, wait and signals
How processes are born, run and die.
Three symptoms bring you to the process lifecycle: a process that will not die even with kill -9, zombies piling up under a service, and fork: Resource temporarily unavailable on a machine with free memory. Each has its cause in how the kernel creates, runs and ends processes. In this lesson you follow a process from creation to reaping under strace, read its state and signal masks in /proc, make a task that SIGKILL cannot end and see why, and run a service into its task limit. Linux essentials covered signals, kill and job control from the shell; this is what the kernel does underneath.
Tasks, threads and PIDs
The kernel schedules tasks. A single-threaded process is one task; a multithreaded process is a group of tasks that share one address space, one file descriptor table and one set of signal handlers. Each task has its own ID. The ID of the first task is the thread group ID (TGID), and that is what ps, kill and /proc call the process ID.
ps -L prints one line per thread. The rsyslog daemon is one process with four tasks: the LWP (light-weight process, the thread ID) of the first equals the PID, and the others have IDs of their own and names that rsyslog gave them. Every thread also appears under /proc/PID/task/TID. Threads count against the same task limits as processes, which matters at the end of this lesson.
fork, exec and wait
A new process starts as a copy of its parent. fork() creates a child that runs the same program with a new PID. Its memory is shared copy-on-write: parent and child see the same pages, and a page is copied only when one side writes to it. glibc implements fork() with the clone system call, the same call that creates threads with different flags.
The child then usually calls execve(), which replaces the program in its memory with a new one. The PID, the parent and the open files stay, except files marked close-on-exec, which the kernel closes at this point.
When a program ends it calls exit_group(), which ends all its threads, and the parent collects the exit status with wait4() or waitid(). strace, which the system-call lesson covers in depth, prints the system calls a program makes; -f follows its children, and -e trace=process shows only these calls.
Read it from the top. execve starts bash. bash calls clone twice, once for each side of the pipe, and each call returns the new child's PID to the parent. Lines tagged [pid N] come from the children: each calls execve for ls and wc, then exit_group(0), and strace reports +++ exited with 0 +++. The parent's wait4(-1, ...) returns each child's PID with its status. The kernel also sends the parent SIGCHLD when a child exits, but two children exited and only one SIGCHLD arrived: standard signals do not queue, which is why a parent must keep calling wait until it reports ECHILD (no children left), as bash does last. The = ? after exit_group means the call never returns.
The shell printed its PID and then replaced itself with ps, which reports the same PID with a different command. This is why a wrapper script that starts a service should end with exec /path/to/program: the program takes over the PID systemd is watching, so stop and reload signals reach the program rather than the shell.
Exit, zombies and orphans
When a process exits, the kernel frees its memory and closes its files at once, but keeps a small record, its PID and exit status, until the parent asks for it with wait(). A task in that state is a zombie, Z in ps. A parent that never waits leaves zombies behind. To see one, start a shell that runs sleep 1 in the background and then replaces itself with sleep 600, a program that never calls wait(). pgrep -u $USER -x sleep finds your own sleep processes and prints their PIDs, which ps then shows; limiting it to your user keeps other people's sleep processes on a shared machine out of the picture.
The parent sleeps in S state; its child exited after a second and is now Z, with no wait channel because it is not waiting for anything (ps columns that show the command line mark it <defunct>). A zombie uses no CPU and no memory beyond that record, but it keeps its PID. Signals cannot remove it, because it is already dead.
pgrep -r selects by state, a quick way to count zombies. kill -KILL changed nothing. Ending the parent did: its children were handed to PID 1, and systemd, like every init system, reaps whatever it inherits. The fix for piling zombies is therefore always in the parent: fix its code so it waits, or restart it. The last pgrep printed nothing because no sleep of yours is left, and exited with status 1.
A child whose parent exits while the child is still running is an orphan, and the kernel gives it a new parent straight away.
The shell exited as soon as it had started sleep, and the orphan's parent is now PID 1. A process can ask to adopt its orphaned descendants instead with prctl(PR_SET_CHILD_SUBREAPER). The systemd user manager does this, and so do container runtime shims such as containerd-shim and conmon, so orphans are reaped by the manager that started them. Inside a PID namespace, the separate PID numbering a container gets (the namespaces lesson builds one), orphans go to that namespace's PID 1, which is why a container's first process must reap children or be a small init such as tini.
Signals: caught, ignored, blocked and pending
Every signal has a default action: terminate, terminate with a core dump, ignore, stop or continue. A process can change that per signal by catching it with a handler or ignoring it, and each thread can block signals, which holds them pending until they are unblocked. SIGKILL and SIGSTOP cannot be caught, ignored or blocked. The kernel publishes all of this as bitmasks in status: SigCgt (caught), SigIgn (ignored), SigBlk (blocked), and two pending sets, SigPnd for signals sent to one thread and ShdPnd for signals sent to the whole process. This small program catches SIGTERM, blocks SIGUSR1 and sleeps; save it as /var/tmp/sigdemo, make it executable and start it as a service.
#!/usr/bin/python3# sigdemo: catches SIGTERM, blocks SIGUSR1, then sleeps.import signal, timesignal.signal(signal.SIGTERM, lambda signum, frame: print('SIGTERM caught, carrying on', flush=True))signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGUSR1})while True:time.sleep(3600)
Each mask is 64 bits in hexadecimal, and bit n-1 stands for signal n. SigCgt 0x4002 has bits 1 and 14 set: signals 2 (SIGINT, which Python handles by default to raise KeyboardInterrupt) and 15 (SIGTERM, our handler). SigIgn 0x1001000 is signals 13 and 25, SIGPIPE and SIGXFSZ, which the Python runtime ignores at startup. SigBlk 0x200 is signal 10, SIGUSR1. kill -l translates numbers to names.
SIGUSR1 did nothing visible: it is blocked, so it sits in ShdPnd until the program unblocks it. SIGTERM ran the handler, which logged a line to the journal, and the process carried on. SIGSTOP put it in state T, waiting in do_signal_stop, and cannot be caught; SIGCONT resumes a stopped process even if it has a handler. SIGKILL ended it, and systemd reports the unit as failed because its main process was killed. When a process ignores your SIGTERM, these masks tell you whether it caught, ignored or blocked the signal before you reach for SIGKILL.
Process states and the unkillable task
The essentials lesson "Processes and /proc" introduced the state letters (R, S, D, T, Z, I) and the zombie count. Two details matter for debugging. A lowercase t is a task in a ptrace stop, held by a debugger or tracer (gdb at a breakpoint, for example) rather than by a stop signal; it runs again when the tracer resumes it or detaches. And D, uninterruptible sleep, means the task is waiting inside the kernel for something that must finish before it can safely return, usually I/O. It counts in the load average although it uses no CPU, and whether a signal can end it depends on which kind of kernel sleep it is in.
To make one on purpose, the lab builds a 64 MiB disk from a file, puts a device-mapper device in front of it and suspends that device. A suspended device-mapper device holds every I/O sent to it, just like a storage path that stopped answering: a dead NFS server, a SAN path that is down, a failing disk. truncate creates the 64 MiB file, and losetup --find --show attaches it to a free loop device (a block device backed by a file) and prints its name. The dmsetup table 0 131072 linear $loop 0 says: the new device's sectors 0 to 131071 (512 bytes each, so 64 MiB) map one to one onto the loop device, starting at its sector 0. The block-layer lesson explains loop devices and device-mapper properly; here the device only has to stop answering. Two dd readers are then started as transient services: one reads with iflag=direct (O_DIRECT, bypassing the page cache), the other reads normally through the page cache.
The readers run as root, and another user's wait channel is behind the kernel's ptrace read-access check (same user or root), so this ps runs with sudo. Both are in D. The direct reader waits in submit_bio_wait for its block I/O to complete; the buffered one waits in folio_wait_bit_common for a page of the page cache to be filled. In Dsl, s marks a session leader (each is its service's main process) and l a multithreaded process: Ubuntu's dd is the Rust uutils version, which starts a helper thread. Now send both SIGKILL.
The buffered reader is gone and the direct one is still there. D in ps covers two kernel sleep types. A task in TASK_KILLABLE sleep ignores ordinary signals but wakes for a fatal one; the page-cache read waits that way. A task in plain TASK_UNINTERRUPTIBLE sleep, like this direct I/O wait, wakes only when the I/O completes. The signal is not lost: ShdPnd has bit 8 set, which is signal 9, SIGKILL, pending until the task returns from the kernel. The stack shows where it waits, which on a real server names the subsystem to investigate (NFS, a device driver, a filesystem).
Resuming the device let the I/O complete, and the pending SIGKILL ended the process the moment it left the kernel. On a real server the equivalent is repairing the storage path or waiting for the I/O to time out; when it never completes, only a reboot clears the task. Clean up the lab device when you are done; losetup --associated finds the loop device that backs the file.
When fork fails: task limits
Every task needs a PID, and several limits cap how many can exist. kernel.pid_max is the largest PID, kernel.threads-max the most tasks the kernel will create (sized from RAM), each cgroup can have a pids.max, which systemd sets from a unit's TasksMax=, and ulimit -u (RLIMIT_NPROC) caps the tasks of one user. When a limit is reached, fork() and clone() fail with EAGAIN, "Resource temporarily unavailable", regardless of free memory. Threads count, so a Java or Go service with many threads reaches the limit sooner than its process count suggests.
The first two numbers are kernel.pid_max and kernel.threads-max. systemd's DefaultTasksMax, the TasksMax= every service gets unless it sets its own, is 15% of the smallest of kernel.pid_max, kernel.threads-max and the root cgroup's task limit, which is unlimited: 15% of 27382 is 4107 on this machine. (Login sessions have their own limit, set on the slice, the cgroup group, that holds each user's processes.) RHEL 10 changes the default to 80%, as its systemd-system.conf man page says:
Now run a service into a limit of four tasks: a bash loop that tries to start six background sleep processes.
Tasks: 4 (limit: 4) is the symptom in one line: bash and three sleeps. bash reports the failed fork and retries with increasing delays, which is why the error says retry; most programs just fail. pids.events counts the forks the cgroup refused (max), and pids.peak shows the highest count reached. On a real service, first find out whether the task count is legitimate (more workers, more threads) or a leak (children never reaped, threads never joined); raising the limit on a leak only delays the failure.
Try this
Start the k-lifecycle-tasks service again and, within about five seconds while bash is still retrying, raise its limit without restarting it: sudo systemctl set-property --runtime k-lifecycle-tasks TasksMax=8. Wait a few seconds and run systemctl status k-lifecycle-tasks --lines=3: expect Tasks: 7 (limit: 8), bash and all six sleeps, because the next retry succeeded. --runtime makes the change last only until the next reboot. Stop the service with sudo systemctl stop k-lifecycle-tasks.
Takeaway
Before you escalate to kill -9 or a reboot, read the task's state, stack and signal masks in /proc: a zombie needs its parent fixed, a D task needs its I/O to finish, and a fork failure needs the task count explained before the limit is raised.