Processes: fork, exec, wait and signals

How processes are born, run and die.

Advanced16 min · lesson 3 of 21

Three symptoms bring you to the process lifecycle: a process that will not die even with kill -9, zombies piling up under a service, and fork: Resource temporarily unavailable on a machine with free memory. Each has its cause in how the kernel creates, runs and ends processes. In this lesson you follow a process from creation to reaping under strace, read its state and signal masks in /proc, make a task that SIGKILL cannot end and see why, and run a service into its task limit. Linux essentials covered signals, kill and job control from the shell; this is what the kernel does underneath.

Tasks, threads and PIDs

The kernel schedules tasks. A single-threaded process is one task; a multithreaded process is a group of tasks that share one address space, one file descriptor table and one set of signal handlers. Each task has its own ID. The ID of the first task is the thread group ID (TGID), and that is what ps, kill and /proc call the process ID.

deploy@web01 · Ubuntu 26.04 LTS
$ ps -L -o pid,lwp,comm -p $(pgrep -x rsyslogd)
PID LWP COMMAND 251171 251171 rsyslogd 251171 251172 in:imuxsock 251171 251173 in:imklog 251171 251174 rs:main Q:Reg

ps -L prints one line per thread. The rsyslog daemon is one process with four tasks: the LWP (light-weight process, the thread ID) of the first equals the PID, and the others have IDs of their own and names that rsyslog gave them. Every thread also appears under /proc/PID/task/TID. Threads count against the same task limits as processes, which matters at the end of this lesson.

fork, exec and wait

A new process starts as a copy of its parent. fork() creates a child that runs the same program with a new PID. Its memory is shared copy-on-write: parent and child see the same pages, and a page is copied only when one side writes to it. glibc implements fork() with the clone system call, the same call that creates threads with different flags.

The child then usually calls execve(), which replaces the program in its memory with a new one. The PID, the parent and the open files stay, except files marked close-on-exec, which the kernel closes at this point.

When a program ends it calls exit_group(), which ends all its threads, and the parent collects the exit status with wait4() or waitid(). strace, which the system-call lesson covers in depth, prints the system calls a program makes; -f follows its children, and -e trace=process shows only these calls.

deploy@web01 · Ubuntu 26.04 LTS
$ strace -f -q -e trace=process bash -c 'ls /etc | wc -l > /dev/null'
execve("/usr/bin/bash", ["bash", "-c", "ls /etc | wc -l > /dev/null"], 0xfffff33de3d0 /* 12 vars */) = 0 clone(child_stack=NULL, flags=CLONE_CHILD_CLEARTID|CLONE_CHILD_SETTID|SIGCHLD, child_tidptr=0xef8d173ff0f0) = 322720 [pid 322719] clone(child_stack=NULL, flags=CLONE_CHILD_CLEARTID|CLONE_CHILD_SETTID|SIGCHLD, child_tidptr=0xef8d173ff0f0) = 322721 [pid 322719] wait4(-1 <unfinished ...> [pid 322720] execve("/usr/bin/ls", ["ls", "/etc"], 0xba2401910860 /* 12 vars */) = 0 [pid 322721] execve("/usr/bin/wc", ["wc", "-l"], 0xba2401910860 /* 12 vars */) = 0 [pid 322720] exit_group(0) = ? [pid 322720] +++ exited with 0 +++ [pid 322719] <... wait4 resumed>, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, NULL) = 322720 [pid 322719] wait4(-1 <unfinished ...> [pid 322721] exit_group(0) = ? [pid 322721] +++ exited with 0 +++ <... wait4 resumed>, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, NULL) = 322721 --- SIGCHLD {si_signo=SIGCHLD, si_code=CLD_EXITED, si_pid=322720, si_uid=1001, si_status=0, si_utime=0, si_stime=0} --- wait4(-1, 0xfffff4208ad4, WNOHANG, NULL) = -1 ECHILD (No child processes) exit_group(0) = ? +++ exited with 0 +++

Read it from the top. execve starts bash. bash calls clone twice, once for each side of the pipe, and each call returns the new child's PID to the parent. Lines tagged [pid N] come from the children: each calls execve for ls and wc, then exit_group(0), and strace reports +++ exited with 0 +++. The parent's wait4(-1, ...) returns each child's PID with its status. The kernel also sends the parent SIGCHLD when a child exits, but two children exited and only one SIGCHLD arrived: standard signals do not queue, which is why a parent must keep calling wait until it reports ECHILD (no children left), as bash does last. The = ? after exit_group means the call never returns.

deploy@web01 · Ubuntu 26.04 LTS
$ sh -c 'echo "shell PID $$"; exec ps -o pid,comm -p $$'
shell PID 322730 PID COMMAND 322730 ps

The shell printed its PID and then replaced itself with ps, which reports the same PID with a different command. This is why a wrapper script that starts a service should end with exec /path/to/program: the program takes over the PID systemd is watching, so stop and reload signals reach the program rather than the shell.

A process from creation to reaping
1fork() / clone()
child is a copy of the parent, new PID
2execve()
new program replaces the image, same PID
3Runs and sleeps
states R, S, D and T
4exit()
memory and files freed; zombie until reaped
5Parent calls wait()
reads the exit status; the PID is released
6Parent exited first?
child moves to a subreaper or to PID 1

Exit, zombies and orphans

When a process exits, the kernel frees its memory and closes its files at once, but keeps a small record, its PID and exit status, until the parent asks for it with wait(). A task in that state is a zombie, Z in ps. A parent that never waits leaves zombies behind. To see one, start a shell that runs sleep 1 in the background and then replaces itself with sleep 600, a program that never calls wait(). pgrep -u $USER -x sleep finds your own sleep processes and prints their PIDs, which ps then shows; limiting it to your user keeps other people's sleep processes on a shared machine out of the picture.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-lifecycle-zombie --uid=$USER sh -c 'sleep 1 & exec sleep 600'
Running as unit: k-lifecycle-zombie.service; invocation ID: 200a841a129445aa95567648a17662b4
$ ps -o pid,ppid,stat,wchan:20,comm $(pgrep -u $USER -x sleep)
PID PPID STAT WCHAN COMMAND 322770 1 Ss hrtimer_nanosleep sleep 322780 322770 Z - sleep

The parent sleeps in S state; its child exited after a second and is now Z, with no wait channel because it is not waiting for anything (ps columns that show the command line mark it <defunct>). A zombie uses no CPU and no memory beyond that record, but it keeps its PID. Signals cannot remove it, because it is already dead.

deploy@web01 · Ubuntu 26.04 LTS
$ kill -KILL $(pgrep -u $USER -r Z -x sleep) ps -o pid,ppid,stat,comm $(pgrep -u $USER -x sleep)
PID PPID STAT COMMAND 322770 1 Ss sleep 322780 322770 Z sleep
$ kill $(pgrep -u $USER -r S -x sleep)
$ pgrep -u $USER -x sleep

pgrep -r selects by state, a quick way to count zombies. kill -KILL changed nothing. Ending the parent did: its children were handed to PID 1, and systemd, like every init system, reaps whatever it inherits. The fix for piling zombies is therefore always in the parent: fix its code so it waits, or restart it. The last pgrep printed nothing because no sleep of yours is left, and exited with status 1.

A child whose parent exits while the child is still running is an orphan, and the kernel gives it a new parent straight away.

deploy@web01 · Ubuntu 26.04 LTS
$ sh -c 'sleep 5 & echo "started sleep as PID $!"' ps -o pid,ppid,stat,comm $(pgrep -u $USER -x sleep)
started sleep as PID 322931 PID PPID STAT COMMAND 322931 1 S sleep

The shell exited as soon as it had started sleep, and the orphan's parent is now PID 1. A process can ask to adopt its orphaned descendants instead with prctl(PR_SET_CHILD_SUBREAPER). The systemd user manager does this, and so do container runtime shims such as containerd-shim and conmon, so orphans are reaped by the manager that started them. Inside a PID namespace, the separate PID numbering a container gets (the namespaces lesson builds one), orphans go to that namespace's PID 1, which is why a container's first process must reap children or be a small init such as tini.

Signals: caught, ignored, blocked and pending

Every signal has a default action: terminate, terminate with a core dump, ignore, stop or continue. A process can change that per signal by catching it with a handler or ignoring it, and each thread can block signals, which holds them pending until they are unblocked. SIGKILL and SIGSTOP cannot be caught, ignored or blocked. The kernel publishes all of this as bitmasks in status: SigCgt (caught), SigIgn (ignored), SigBlk (blocked), and two pending sets, SigPnd for signals sent to one thread and ShdPnd for signals sent to the whole process. This small program catches SIGTERM, blocks SIGUSR1 and sleeps; save it as /var/tmp/sigdemo, make it executable and start it as a service.

/var/tmp/sigdemo
#!/usr/bin/python3
# sigdemo: catches SIGTERM, blocks SIGUSR1, then sleeps.
import signal, time
signal.signal(signal.SIGTERM, lambda signum, frame: print('SIGTERM caught, carrying on', flush=True))
signal.pthread_sigmask(signal.SIG_BLOCK, {signal.SIGUSR1})
while True:
time.sleep(3600)
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-lifecycle-sigdemo --uid=$USER /var/tmp/sigdemo
Running as unit: k-lifecycle-sigdemo.service; invocation ID: d367ff0e6f2e4ccf91b4b04718675e26
$ grep -E '^(State|SigPnd|ShdPnd|SigBlk|SigIgn|SigCgt)' /proc/$(pgrep -x sigdemo)/status
State: S (sleeping) SigPnd: 0000000000000000 ShdPnd: 0000000000000000 SigBlk: 0000000000000200 SigIgn: 0000000001001000 SigCgt: 0000000000004002
$ kill -l 2 10 13 15 25
INT USR1 PIPE TERM XFSZ

Each mask is 64 bits in hexadecimal, and bit n-1 stands for signal n. SigCgt 0x4002 has bits 1 and 14 set: signals 2 (SIGINT, which Python handles by default to raise KeyboardInterrupt) and 15 (SIGTERM, our handler). SigIgn 0x1001000 is signals 13 and 25, SIGPIPE and SIGXFSZ, which the Python runtime ignores at startup. SigBlk 0x200 is signal 10, SIGUSR1. kill -l translates numbers to names.

deploy@web01 · Ubuntu 26.04 LTS
$ kill -USR1 $(pgrep -x sigdemo) grep -E '^(State|ShdPnd|SigBlk)' /proc/$(pgrep -x sigdemo)/status
State: S (sleeping) ShdPnd: 0000000000000200 SigBlk: 0000000000000200
$ kill -TERM $(pgrep -x sigdemo) sleep 1 journalctl -u k-lifecycle-sigdemo -o cat -n 1
SIGTERM caught, carrying on
$ kill -STOP $(pgrep -x sigdemo) ps -o pid,stat,wchan:20,comm -p $(pgrep -x sigdemo) kill -CONT $(pgrep -x sigdemo)
PID STAT WCHAN COMMAND 322968 Ts do_signal_stop sigdemo
$ kill -KILL $(pgrep -x sigdemo) sleep 1 systemctl is-active k-lifecycle-sigdemo
failed

SIGUSR1 did nothing visible: it is blocked, so it sits in ShdPnd until the program unblocks it. SIGTERM ran the handler, which logged a line to the journal, and the process carried on. SIGSTOP put it in state T, waiting in do_signal_stop, and cannot be caught; SIGCONT resumes a stopped process even if it has a handler. SIGKILL ended it, and systemd reports the unit as failed because its main process was killed. When a process ignores your SIGTERM, these masks tell you whether it caught, ignored or blocked the signal before you reach for SIGKILL.

Process states and the unkillable task

The essentials lesson "Processes and /proc" introduced the state letters (R, S, D, T, Z, I) and the zombie count. Two details matter for debugging. A lowercase t is a task in a ptrace stop, held by a debugger or tracer (gdb at a breakpoint, for example) rather than by a stop signal; it runs again when the tracer resumes it or detaches. And D, uninterruptible sleep, means the task is waiting inside the kernel for something that must finish before it can safely return, usually I/O. It counts in the load average although it uses no CPU, and whether a signal can end it depends on which kind of kernel sleep it is in.

To make one on purpose, the lab builds a 64 MiB disk from a file, puts a device-mapper device in front of it and suspends that device. A suspended device-mapper device holds every I/O sent to it, just like a storage path that stopped answering: a dead NFS server, a SAN path that is down, a failing disk. truncate creates the 64 MiB file, and losetup --find --show attaches it to a free loop device (a block device backed by a file) and prints its name. The dmsetup table 0 131072 linear $loop 0 says: the new device's sectors 0 to 131071 (512 bytes each, so 64 MiB) map one to one onto the loop device, starting at its sector 0. The block-layer lesson explains loop devices and device-mapper properly; here the device only has to stop answering. Two dd readers are then started as transient services: one reads with iflag=direct (O_DIRECT, bypassing the page cache), the other reads normally through the page cache.

deploy@web01 · Ubuntu 26.04 LTS
$ truncate -s 64M /var/tmp/k-lifecycle.img loop=$(sudo losetup --find --show /var/tmp/k-lifecycle.img) sudo dmsetup create k-lifecycle-slow --table "0 131072 linear $loop 0" sudo dmsetup suspend k-lifecycle-slow
$ sudo systemd-run --unit=k-lifecycle-direct dd if=/dev/mapper/k-lifecycle-slow of=/dev/null bs=4k count=1 iflag=direct sudo systemd-run --unit=k-lifecycle-cached dd if=/dev/mapper/k-lifecycle-slow of=/dev/null bs=4k count=1 skip=100
Running as unit: k-lifecycle-direct.service; invocation ID: 3696b20e265b4753a63ca93d6210dc6a Running as unit: k-lifecycle-cached.service; invocation ID: b769c46cd9604e06ba35a6b9c42647f1
$ sudo ps -o pid,stat,wchan:24,args -C dd
PID STAT WCHAN COMMAND 323206 Dsl submit_bio_wait /usr/bin/dd if=/dev/mapper/k-lifecycle-slow of=/dev/null bs=4k count=1 iflag=direct 323211 Dsl folio_wait_bit_common /usr/bin/dd if=/dev/mapper/k-lifecycle-slow of=/dev/null bs=4k count=1 skip=100

The readers run as root, and another user's wait channel is behind the kernel's ptrace read-access check (same user or root), so this ps runs with sudo. Both are in D. The direct reader waits in submit_bio_wait for its block I/O to complete; the buffered one waits in folio_wait_bit_common for a page of the page cache to be filled. In Dsl, s marks a session leader (each is its service's main process) and l a multithreaded process: Ubuntu's dd is the Rust uutils version, which starts a helper thread. Now send both SIGKILL.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo kill -KILL $(pgrep -x dd) sleep 1 sudo ps -o pid,stat,wchan:24,args -C dd
PID STAT WCHAN COMMAND 323206 Ds submit_bio_wait /usr/bin/dd if=/dev/mapper/k-lifecycle-slow of=/dev/null bs=4k count=1 iflag=direct
$ grep -E '^(State|ShdPnd)' /proc/$(pgrep -x dd)/status
State: D (disk sleep) ShdPnd: 0000000000000100
$ sudo cat /proc/$(pgrep -x dd)/stack
[<0>] submit_bio_wait+0x88/0xf8 [<0>] __blkdev_direct_IO_simple+0x144/0x2c0 [<0>] blkdev_direct_IO+0x124/0x1e8 [<0>] blkdev_read_iter+0xfc/0x1a8 [<0>] vfs_read+0x278/0x380 [<0>] ksys_read+0x70/0x130 [<0>] __arm64_sys_read+0x24/0x50 [<0>] invoke_syscall.constprop.0+0x60/0xf0 [<0>] el0_svc_common.constprop.0+0x114/0x140 [<0>] do_el0_svc+0x28/0x58 [<0>] el0_svc+0x40/0x1e0 [<0>] el0t_64_sync_handler+0xc0/0x110 [<0>] el0t_64_sync+0x1b8/0x1c0

The buffered reader is gone and the direct one is still there. D in ps covers two kernel sleep types. A task in TASK_KILLABLE sleep ignores ordinary signals but wakes for a fatal one; the page-cache read waits that way. A task in plain TASK_UNINTERRUPTIBLE sleep, like this direct I/O wait, wakes only when the I/O completes. The signal is not lost: ShdPnd has bit 8 set, which is signal 9, SIGKILL, pending until the task returns from the kernel. The stack shows where it waits, which on a real server names the subsystem to investigate (NFS, a device driver, a filesystem).

deploy@web01 · Ubuntu 26.04 LTS
$ sudo dmsetup resume k-lifecycle-slow sleep 1 ps -o pid,stat,args -C dd
PID STAT COMMAND

Resuming the device let the I/O complete, and the pending SIGKILL ended the process the moment it left the kernel. On a real server the equivalent is repairing the storage path or waiting for the I/O to time out; when it never completes, only a reboot clears the task. Clean up the lab device when you are done; losetup --associated finds the loop device that backs the file.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo dmsetup remove k-lifecycle-slow sudo losetup --detach "$(sudo losetup --noheadings --output NAME --associated /var/tmp/k-lifecycle.img)" rm /var/tmp/k-lifecycle.img

When fork fails: task limits

Every task needs a PID, and several limits cap how many can exist. kernel.pid_max is the largest PID, kernel.threads-max the most tasks the kernel will create (sized from RAM), each cgroup can have a pids.max, which systemd sets from a unit's TasksMax=, and ulimit -u (RLIMIT_NPROC) caps the tasks of one user. When a limit is reached, fork() and clone() fail with EAGAIN, "Resource temporarily unavailable", regardless of free memory. Threads count, so a Java or Go service with many threads reaches the limit sooner than its process count suggests.

deploy@web01 · Ubuntu 26.04 LTS
$ cat /proc/sys/kernel/pid_max /proc/sys/kernel/threads-max systemctl show --property=DefaultTasksMax
4194304 27382 DefaultTasksMax=4107

The first two numbers are kernel.pid_max and kernel.threads-max. systemd's DefaultTasksMax, the TasksMax= every service gets unless it sets its own, is 15% of the smallest of kernel.pid_max, kernel.threads-max and the root cgroup's task limit, which is unlimited: 15% of 27382 is 4107 on this machine. (Login sessions have their own limit, set on the slice, the cgroup group, that holds each user's processes.) RHEL 10 changes the default to 80%, as its systemd-system.conf man page says:

deploy@rocky10 · Rocky Linux 10.2
$ cat /proc/sys/kernel/pid_max /proc/sys/kernel/threads-max systemctl show --property=DefaultTasksMax
4194304 28730 DefaultTasksMax=22984

Now run a service into a limit of four tasks: a bash loop that tries to start six background sleep processes.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=k-lifecycle-tasks -p TasksMax=4 --uid=$USER bash -c 'for i in 1 2 3 4 5 6; do sleep 300 & done; wait'
Running as unit: k-lifecycle-tasks.service; invocation ID: 7e8ab6c47de84dda85a900569d64e0a0
$ systemctl status k-lifecycle-tasks
● k-lifecycle-tasks.service - [systemd-run] /usr/bin/bash -c "for i in 1 2 3 4 5 6; do sleep 300 & done; wait" Loaded: loaded (/run/systemd/transient/k-lifecycle-tasks.service; transient) Transient: yes Active: active (running) since Sun 2026-09-27 11:09:54 UTC; 4s ago Invocation: 7e8ab6c47de84dda85a900569d64e0a0 Main PID: 323452 (bash) Tasks: 4 (limit: 4) Memory: 4.3M (peak: 4.6M) CPU: 15ms CGroup: /system.slice/k-lifecycle-tasks.service ├─323452 /usr/bin/bash -c "for i in 1 2 3 4 5 6; do sleep 300 & done; wait" ├─323461 sleep 300 ├─323462 sleep 300 └─323463 sleep 300 Sep 27 11:09:54 web01 systemd[1]: Started k-lifecycle-tasks.service - [systemd-run] /usr/bin/bash -c "for i in 1 2 3 4 5 6; do sleep 300 & done; wait". Sep 27 11:09:54 web01 bash[323452]: /usr/bin/bash: fork: retry: Resource temporarily unavailable Sep 27 11:09:55 web01 bash[323452]: /usr/bin/bash: fork: retry: Resource temporarily unavailable Sep 27 11:09:57 web01 bash[323452]: /usr/bin/bash: fork: retry: Resource temporarily unavailable
$ grep . /sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.*
/sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.current:4 /sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.events:max 3 /sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.events.local:max 3 /sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.max:4 /sys/fs/cgroup/system.slice/k-lifecycle-tasks.service/pids.peak:4

Tasks: 4 (limit: 4) is the symptom in one line: bash and three sleeps. bash reports the failed fork and retries with increasing delays, which is why the error says retry; most programs just fail. pids.events counts the forks the cgroup refused (max), and pids.peak shows the highest count reached. On a real service, first find out whether the task count is legitimate (more workers, more threads) or a leak (children never reaped, threads never joined); raising the limit on a leak only delays the failure.

Try this

Start the k-lifecycle-tasks service again and, within about five seconds while bash is still retrying, raise its limit without restarting it: sudo systemctl set-property --runtime k-lifecycle-tasks TasksMax=8. Wait a few seconds and run systemctl status k-lifecycle-tasks --lines=3: expect Tasks: 7 (limit: 8), bash and all six sleeps, because the next retry succeeded. --runtime makes the change last only until the next reboot. Stop the service with sudo systemctl stop k-lifecycle-tasks.

Takeaway

Before you escalate to kill -9 or a reboot, read the task's state, stack and signal masks in /proc: a zombie needs its parent fixed, a D task needs its I/O to finish, and a fork failure needs the task count explained before the limit is raised.

Quick check
01Monitoring reports 3,000 processes in state Z, all children of worker PID 812, and the number keeps growing. kill -9 on the zombies changes nothing. What stops the growth?
Incorrect — Permission is not the issue: a zombie is already dead, so no signal can affect it, even from root.
Incorrect — That buys time but the zombies keep accumulating, because nothing collects their exit status.
Incorrect — The kernel keeps a zombie until its parent waits for it; there is no timeout.
Correct — Only the parent's wait() releases a zombie. If the parent exits, the zombies pass to PID 1 or a subreaper, which reaps them.
02A service's ExecStart is a shell script that sets two variables and then starts /opt/app/server without exec. The unit has ExecReload=/bin/kill -HUP $MAINPID, and systemctl reload never makes the server re-read its configuration. Why?
Incorrect — ExecReload runs exactly the command given, and kill -HUP $MAINPID signals one PID only.
Incorrect — Being a child of a shell does not change signal delivery; a SIGHUP sent to the server's PID would reach it.
Correct — Without exec, the shell remains the main process. With exec the server would take over the shell's PID and receive the signal.
Incorrect — systemd does not block SIGHUP for services; the signal mask in /proc/PID/status would show it if something did.
03After sudo kill -KILL, one blocked reader disappears but another stays in state D, and its /proc/PID/status shows ShdPnd: 0000000000000100. What does that line tell you?
Incorrect — A permission failure is reported by kill and queues nothing. A pending bit means the signal was accepted.
Correct — Bit 8 is signal 9. A task in uninterruptible sleep acts on it only when it leaves that sleep, when the I/O completes.
Incorrect — Bit n-1 represents signal n, so 0x100 (bit 8) is signal 9, SIGKILL, not signal 8.
Incorrect — SIGKILL cannot be blocked, caught or ignored. The delay comes from the kind of sleep, not from a mask.

Related