Processes, pipes and file descriptors in real scripts
Where Bash forks, file descriptors, process substitution status, process groups and wait -n.
tar -xzf scr-bash-internals.tar.gz, which creates scr-bash-internals/. SHA-256: a753f819b1807c4cd0c1023aa4f6f1c4b1ad0411a61fb74a3c2aa3b16b11e60cThis lesson shows where a Bash script creates child processes and what that does to its state, its exit statuses and its open files. By the end you can keep a count out of a subshell, collect the exit status of a process substitution so a security gate fails closed, see which file descriptors a background helper inherited, and stop a whole tree of processes without leaving orphans. These are the bugs that let a script report success while it checked nothing.
This course assumes you have finished "Bash for ops, done safely" and "Python for security automation", and that you write automation that runs unattended. Quick refresher: "Loops and reading input safely" (bash-ops) showed while IFS= read -r and the < <(cmd) form, and "Strict mode, honestly" showed what set -euo pipefail and inherit_errexit do. Here we look at what those tools do once a script starts children, holds locks and runs work in parallel.
sudo, with Bash 5.3 and Python 3.14 as shipped. Each lesson has a "Lesson files" box at the top: download it, unpack it in your home directory, and work in the directory it creates (here ~/scr-bash-internals). The terminals show a user called deploy; run everything as your own normal user, and where a command needs your user name it uses $USER. Each lesson shows the packages it needs when it first uses them; to install the archive ones in one go: sudo apt install jq shellcheck shfmt bats bats-assert bats-support tini python3-venv podman. Python packages are always installed into a virtual environment the lesson creates, never into the system Python.Where Bash forks
A subshell is a copy of the running shell in a new process. It starts with copies of all variables, and nothing it changes comes back to the parent. Bash creates one for ( ... ), for every $(...), for every command started with &, and for each stage of a pipeline. $BASHPID holds the PID of the process that is actually running (unlike $$, which always names the main script), so you can watch it happen:
Four different numbers: only the first line ran in the shell you typed into. The last one is the case "Loops and reading input safely" showed losing a count; here it hides inside a security gate that counts failed SSH logins and blocks when there are more than three:
#!/usr/bin/env bash# Deploy gate: blocks when failed SSH logins exceed a limit. BROKEN (the count is always 0).set -euo pipefaillog=${1:?usage: gate-pipe.sh LOGFILE [MAX]}max=${2:-3}failed=0grep -F 'Failed password' -- "$log" | while IFS= read -r _; dofailed=$((failed + 1))doneecho "failed logins: $failed (max $max)"if (( failed > max )); then echo "BLOCK"; exit 1; fiecho "PASS"
The log has five failed logins and the gate says zero and passes. The while loop is the last stage of the pipeline, so it ran in a subshell: it counted to five in its own copy of failed, and that copy was thrown away when the stage exited. The echo after the loop reads the parent's variable, which is still 0. ShellCheck 0.11 prints nothing for either gate, so a clean lint run tells you nothing about this class of bug.
One fix is the lastpipe option: when job control is off (the default in scripts), Bash runs the last stage of a pipeline in the current shell. Here it is switched on from the command line with bash -O lastpipe; in a script you would write shopt -s lastpipe near the top.
Process substitution hides the producer's exit status
The fix most people reach for instead is process substitution: <(cmd) runs cmd in the background, connects its output to a pipe and hands the loop a path such as /dev/fd/63 to read from. The loop now runs in the current shell, so the count survives:
#!/usr/bin/env bash# Deploy gate, second try: the loop now runs in this shell. Still unsafe (see the lesson).set -euo pipefaillog=${1:?usage: gate-procsub.sh LOGFILE [MAX]}max=${2:-3}failed=0while IFS= read -r _; dofailed=$((failed + 1))done < <(grep -F 'Failed password' -- "$log")echo "failed logins: $failed (max $max)"if (( failed > max )); then echo "BLOCK"; exit 1; fiecho "PASS"
The second run is the dangerous one. We made a copy of the log that nobody can read, which is what a gate meets in real life when a log is rotated, has the wrong owner or sits on a disk that failed. grep printed Permission denied and exited 2. Nothing noticed: set -e and pipefail only look at commands and pipelines, and the process that feeds < <(...) is neither. The loop read nothing, the count was 0 and the gate printed PASS with exit status 0. The gate failed open.
Bash 4.4 and later set $! to the PID of the most recent process substitution, and wait "$!" returns its exit status. The fixed gate collects that status after the loop and decides what it means. For grep, 0 means lines matched, 1 means no match (a valid answer: nothing failed) and 2 means an error. Only an error makes the gate refuse to decide, with its own exit code:
#!/usr/bin/env bash# Deploy gate: blocks when failed SSH logins exceed a limit, and fails closed# when the log cannot be read. Exit: 0 pass, 1 block, 2 cannot decide.set -euo pipefaillog=${1:?usage: gate.sh LOGFILE [MAX]}max=${2:-3}failed=0while IFS= read -r _; dofailed=$((failed + 1))done < <(grep -F 'Failed password' -- "$log")# $! is the PID of the last process substitution; wait returns its exit status.# grep: 0 = lines matched, 1 = no match (a valid answer), 2 or more = error.rc=0wait "$!" || rc=$?if (( rc > 1 )); thenecho "gate: cannot read $log (grep exit $rc); refusing to pass" >&2exit 2fiecho "failed logins: $failed (max $max)"if (( failed > max )); then echo "BLOCK"; exit 1; fiecho "PASS"
Three distinct results: 1 blocks, 0 passes, 2 says "I could not check", which a pipeline calling this gate must treat as a failure. lastpipe also fails closed, because grep is then still a pipeline stage that pipefail and set -e can see. The script stops at the pipeline with grep's status and never prints a verdict:
$! names only the most recent process substitution. A command with two of them, such as a drift check with diff <(...) <(...), loses the first status. Here one input is missing and the other has no failed logins, and diff reports "identical" with exit 0. Opening each substitution on its own file descriptor with exec {fd}< <(...) gives you one PID per producer to wait for:
Producer a exited 2 (missing file) and b exited 1 (no match), so this check can report that its first input was never read. Process substitution is a Bash feature, not POSIX sh; dash has none of this.
File descriptors you open yourself
A file descriptor (fd) is a small number that refers to an open file, pipe or socket in one process. exec with only redirections changes the current shell's descriptors for the rest of the script, and the {name}>file form lets Bash pick a free number (10 or above) and store it in $name, so you never collide with an fd some other code already uses. This sync job takes a lock with flock on one descriptor, writes its log through another, and starts a helper that keeps running after the job exits:
#!/usr/bin/env bash# Nightly sync: one run at a time, guarded by flock on a lock file descriptor.set -euo pipefailstate=${XDG_STATE_HOME:-$HOME/.local/state}/nightly-syncmkdir -p -- "$state"exec {lock}>"$state/lock" # Bash picks a free fd (10 or above) and stores it in $lockif ! flock -n "$lock"; thenecho "nightly-sync: another run holds the lock" >&2exit 75 # EX_TEMPFAIL: try again laterfiexec {log}>>"$state/run.log"printf 'run %s: started\n' "$$" >&"$log"# A helper that keeps running after this run exits../warm-cache.sh &printf 'run %s: helper %s started\n' "$$" "$!" >&"$log"echo "nightly-sync: run $$ done"
#!/usr/bin/env bash# Stand-in for a slow background helper (a cache warmer, a log shipper).sleep 20
The helper runs for 20 seconds, and that is the window for the next three commands. Run the job twice, the second straight after the first (paste both lines together):
The first run finished and exited. The second run is refused with 75, the "try again later" code, although nothing is running the sync any more. Within those 20 seconds, look at the helper's open descriptors in /proc (if the helper has already finished, find complains that the directory does not exist: run the two sync commands again and repeat):
A child process inherits every open descriptor of its parent that is not marked close-on-exec ("Files, descriptors and the VFS" in Advanced Linux internals covers the descriptor table behind this). The helper holds fd 10 on the lock file and fd 11 on the log. A flock lock belongs to the open file, not to the process that asked for it, so it is released only when every descriptor that shares it is closed. The helper keeps it for as long as it runs. (Fd 255 is Bash keeping its own script open so it can read the next line.) The fix closes both descriptors for that one command only:
Before you try the fixed script, wait until the first helper has finished, because its inherited descriptor still holds the lock. This loop waits (at most 40 seconds) until no helper is left:
Two runs back to back now both get the lock, while their helpers keep running. Close lock and log descriptors on any command you start in the background, and on long-running commands that must not keep the lock alive after the script is gone.
The fix also means two helpers can now run at the same time, and nothing waits for their exit status. If the helper touches shared state, give it its own lock (flock -n helper.lock ./warm-cache.sh). Under systemd a helper started with & does not outlive the job at all: when the main process of a Type=oneshot unit exits, systemd stops the unit and kills everything left in its control group. A helper that must outlive the job belongs in its own unit or its own scheduled job.
Process groups: stop the whole tree
Killing a process does not kill its children. check-host.sh stands in for a slow remote check; it runs sleep as a child. Stopping it by PID leaves that child behind:
#!/usr/bin/env bash# Stand-in for a slow remote check: sleeps, then exits with the given status.host=$1 secs=$2 status=$3sleep "$secs"echo "$host: checked (exit $status)"exit "$status"
The script died and its sleep 30 kept running, an orphan that nothing would ever wait for, until pkill removed it (-x -f matches the whole command line exactly, so it cannot hit anything else). With real work in place of sleep this is a scan that keeps hitting a host after you cancelled it.
Every process belongs to a process group, and kill with a negative number signals the whole group. setsid starts a program in a new session whose process group ID is its own PID. How you learn that ID depends on job control. In a script (job control off), setsid cmd & does not need to fork, so $! is the group ID; sweep.sh below relies on that. At an interactive prompt (job control on), & already puts the job in its own group, setsid forks, and $! names a parent that has already exited, so kill -- "-$!" misses. The form below works in both: setsid -f always forks, and the new group leader writes its own PID to a file before it becomes the check:
ps shows the check and its sleep with the same PGID (process group) and SID (session), both equal to the PID from check.pid. One kill -TERM -- "-$pgid" reached both; the -- stops kill from reading the negative number as an option. The last check looks for the child by its command line rather than by the group ID you assumed, so it would catch a survivor even if the group ID were wrong.
The same idea makes parallel work safe. wait -n (Bash 5.1 and later with a list of PIDs and -p) returns as soon as any one job finishes, with that job's status, and stores its PID in the variable you name. This sweep runs three checks at once, stops at the first failure and cleans up the rest through their process groups from a single EXIT trap:
#!/usr/bin/env bash# Runs host checks in parallel. The first failure stops the rest, and nothing is left running.set -euo pipefaildeclare -A host_of=() # PID of each running check -> host namestart() {# setsid: the check leads its own process group, so one kill reaches everything it startedsetsid ./check-host.sh "$@" &host_of[$!]=$1}# shellcheck disable=SC2329 # called by the EXIT trap; ShellCheck 0.11 misses thatstop_all() {local pid(( ${#host_of[@]} )) || return 0echo "sweep: stopping ${#host_of[@]} unfinished check(s)"for pid in "${!host_of[@]}"; dokill -TERM -- "-$pid" 2>/dev/null || truedonewait # reap them before this script exits}trap stop_all EXITstart web01 1 0start db01 2 3start cache01 30 0worst=0while (( ${#host_of[@]} )); dorc=0wait -n -p done_pid -- "${!host_of[@]}" || rc=$?echo "sweep: ${host_of[$done_pid]} finished with exit $rc"unset "host_of[$done_pid]"if (( rc != 0 )); then worst=$rc; break; fidoneexit "$worst"
db01 failed with 3, the loop stopped, and the EXIT trap killed cache01's group and reaped it with wait. The sweep exits with the failing check's status and pgrep finds no sleep left. The # shellcheck disable=SC2329 line is there because ShellCheck 0.11 reports a function called only from an EXIT trap as "never invoked" when the script ends with an explicit exit; a directive with its reason is the honest fix. One side effect of setsid to keep in mind: the checks are no longer in your terminal's process group, so Ctrl-C reaches only the sweep, and the sweep's trap has to stop them.
coproc starts one background command with a pipe to its input and one from its output, so a script can hold a conversation with it. It works for a short exchange with a helper that flushes every line (fflush() in awk, stdbuf -oL for many tools). The limits arrive quickly: a helper that buffers its output deadlocks your read (use read -t to bound it), and nothing checks the helper's exit status unless you wait "$NAME_PID". If a script needs a long-lived helper it talks to over pipes, it has outgrown Bash; "When a script has outgrown Bash" later in this course covers that handoff.Try this
Run ./gate.sh rotated.log for a log that does not exist and predict the exit status before you look. Then break the fix on purpose: copy nightly-sync2.sh to try-sync.sh, remove only {lock}>&- from the helper line, wait until no helper is left (the loop above), and run ./try-sync.sh && ./try-sync.sh. Confirm the cause with the find /proc/... command from above. Finally, once the helpers have finished, run pgrep -a -u "$USER" -x sleep.
Expected: the gate exits 2, with grep's "No such file or directory" above the gate's refusal. The second try-sync.sh exits 75 again, and find shows fd 10 on the lock file back in the helper. pgrep prints nothing and exits 1.
Takeaway
Every place Bash starts a process is a place where a status, a variable or a descriptor can go missing: collect producer statuses with wait "$!", close lock descriptors on background commands, and stop work by process group, then prove with pgrep that nothing is left.
while read -r l; do n=$((n+1)); done < <(trivy-to-lines report.json) under set -euo pipefail and blocks when n is above 0. One night the report file is missing, the converter exits 1, and the gate passes. What is the smallest correct fix?$(...). The converter runs in a process substitution, and errexit never looks at its status whatever the option says.$! to the last process substitution and wait returns its exit status, so the gate can tell "no findings" from "never read the report".n is lost, so the gate always sees 0; this swaps one fail-open bug for another.flock -n on exec {lock}>/run/backup.lock, then starts ./upload.sh & and exits. The next scheduled run fails with "lock held" although no backup script is running. What is holding the lock?{lock}>&-../scan.sh & (which runs nmap as a child) and on timeout runs kill "$scan_pid". Operators report scans still hitting hosts after the timeout. Which change fixes it?nmap child is orphaned exactly as before.nmap child, which keeps running after the wrapper returns.kill 0 signals the caller's whole process group. Started with & from a script, scan.sh shares the wrapper's group, so this also kills the wrapper.$! is the scan's own group, and a negative PID signals that whole group.