Processes, pipes and file descriptors in real scripts

Where Bash forks, file descriptors, process substitution status, process groups and wait -n.

Advanced35 min · lesson 1 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (10 files, 2 KB): scr-bash-internals.tar.gz. Unpack it with tar -xzf scr-bash-internals.tar.gz, which creates scr-bash-internals/. SHA-256: a753f819b1807c4cd0c1023aa4f6f1c4b1ad0411a61fb74a3c2aa3b16b11e60c

This lesson shows where a Bash script creates child processes and what that does to its state, its exit statuses and its open files. By the end you can keep a count out of a subshell, collect the exit status of a process substitution so a security gate fails closed, see which file descriptors a background helper inherited, and stop a whole tree of processes without leaving orphans. These are the bugs that let a script report success while it checked nothing.

This course assumes you have finished "Bash for ops, done safely" and "Python for security automation", and that you write automation that runs unattended. Quick refresher: "Loops and reading input safely" (bash-ops) showed while IFS= read -r and the < <(cmd) form, and "Strict mode, honestly" showed what set -euo pipefail and inherit_errexit do. Here we look at what those tools do once a script starts children, holds locks and runs work in parallel.

What you need for this course
An Ubuntu 26.04 server or VM where you can use sudo, with Bash 5.3 and Python 3.14 as shipped. Each lesson has a "Lesson files" box at the top: download it, unpack it in your home directory, and work in the directory it creates (here ~/scr-bash-internals). The terminals show a user called deploy; run everything as your own normal user, and where a command needs your user name it uses $USER. Each lesson shows the packages it needs when it first uses them; to install the archive ones in one go: sudo apt install jq shellcheck shfmt bats bats-assert bats-support tini python3-venv podman. Python packages are always installed into a virtual environment the lesson creates, never into the system Python.

Where Bash forks

A subshell is a copy of the running shell in a new process. It starts with copies of all variables, and nothing it changes comes back to the parent. Bash creates one for ( ... ), for every $(...), for every command started with &, and for each stage of a pipeline. $BASHPID holds the PID of the process that is actually running (unlike $$, which always names the main script), so you can watch it happen:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ echo "this shell: $BASHPID" ( echo "( ) subshell: $BASHPID" ) echo "command subst: $(echo "$BASHPID")" true | echo "pipe stage: $BASHPID"
this shell: 529259 ( ) subshell: 529274 command subst: 529275 pipe stage: 529277

Four different numbers: only the first line ran in the shell you typed into. The last one is the case "Loops and reading input safely" showed losing a count; here it hides inside a security gate that counts failed SSH logins and blocks when there are more than three:

gate-pipe.sh
#!/usr/bin/env bash
# Deploy gate: blocks when failed SSH logins exceed a limit. BROKEN (the count is always 0).
set -euo pipefail
log=${1:?usage: gate-pipe.sh LOGFILE [MAX]}
max=${2:-3}
failed=0
grep -F 'Failed password' -- "$log" | while IFS= read -r _; do
failed=$((failed + 1))
done
echo "failed logins: $failed (max $max)"
if (( failed > max )); then echo "BLOCK"; exit 1; fi
echo "PASS"
deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./gate-pipe.sh auth.log
failed logins: 0 (max 3) PASS
$ grep -c "Failed password" auth.log
5
$ shellcheck gate-pipe.sh gate-procsub.sh

The log has five failed logins and the gate says zero and passes. The while loop is the last stage of the pipeline, so it ran in a subshell: it counted to five in its own copy of failed, and that copy was thrown away when the stage exited. The echo after the loop reads the parent's variable, which is still 0. ShellCheck 0.11 prints nothing for either gate, so a clean lint run tells you nothing about this class of bug.

One fix is the lastpipe option: when job control is off (the default in scripts), Bash runs the last stage of a pipeline in the current shell. Here it is switched on from the command line with bash -O lastpipe; in a script you would write shopt -s lastpipe near the top.

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ bash -O lastpipe gate-pipe.sh auth.log
failed logins: 5 (max 3) BLOCK
Where does this code run?
A line in your script
and where its variable changes end up
{ ...; } or a plain command
Current shell
changes stick
( ... ), $(...), cmd &
Subshell
changes are lost when it exits
a | b | while ...
Every stage is a subshell
unless lastpipe, and only for the last stage
done < <(cmd)
Loop in current shell
cmd runs in a separate process

Process substitution hides the producer's exit status

The fix most people reach for instead is process substitution: <(cmd) runs cmd in the background, connects its output to a pipe and hands the loop a path such as /dev/fd/63 to read from. The loop now runs in the current shell, so the count survives:

gate-procsub.sh
#!/usr/bin/env bash
# Deploy gate, second try: the loop now runs in this shell. Still unsafe (see the lesson).
set -euo pipefail
log=${1:?usage: gate-procsub.sh LOGFILE [MAX]}
max=${2:-3}
failed=0
while IFS= read -r _; do
failed=$((failed + 1))
done < <(grep -F 'Failed password' -- "$log")
echo "failed logins: $failed (max $max)"
if (( failed > max )); then echo "BLOCK"; exit 1; fi
echo "PASS"
deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./gate-procsub.sh auth.log
failed logins: 5 (max 3) BLOCK
$ cp auth.log locked.log && chmod 000 locked.log
$ ./gate-procsub.sh locked.log
grep: locked.log: Permission denied failed logins: 0 (max 3) PASS

The second run is the dangerous one. We made a copy of the log that nobody can read, which is what a gate meets in real life when a log is rotated, has the wrong owner or sits on a disk that failed. grep printed Permission denied and exited 2. Nothing noticed: set -e and pipefail only look at commands and pipelines, and the process that feeds < <(...) is neither. The loop read nothing, the count was 0 and the gate printed PASS with exit status 0. The gate failed open.

Bash 4.4 and later set $! to the PID of the most recent process substitution, and wait "$!" returns its exit status. The fixed gate collects that status after the loop and decides what it means. For grep, 0 means lines matched, 1 means no match (a valid answer: nothing failed) and 2 means an error. Only an error makes the gate refuse to decide, with its own exit code:

gate.sh
#!/usr/bin/env bash
# Deploy gate: blocks when failed SSH logins exceed a limit, and fails closed
# when the log cannot be read. Exit: 0 pass, 1 block, 2 cannot decide.
set -euo pipefail
log=${1:?usage: gate.sh LOGFILE [MAX]}
max=${2:-3}
failed=0
while IFS= read -r _; do
failed=$((failed + 1))
done < <(grep -F 'Failed password' -- "$log")
# $! is the PID of the last process substitution; wait returns its exit status.
# grep: 0 = lines matched, 1 = no match (a valid answer), 2 or more = error.
rc=0
wait "$!" || rc=$?
if (( rc > 1 )); then
echo "gate: cannot read $log (grep exit $rc); refusing to pass" >&2
exit 2
fi
echo "failed logins: $failed (max $max)"
if (( failed > max )); then echo "BLOCK"; exit 1; fi
echo "PASS"
deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./gate.sh auth.log
failed logins: 5 (max 3) BLOCK
$ ./gate.sh clean.log
failed logins: 0 (max 3) PASS
$ ./gate.sh locked.log
grep: locked.log: Permission denied gate: cannot read locked.log (grep exit 2); refusing to pass

Three distinct results: 1 blocks, 0 passes, 2 says "I could not check", which a pipeline calling this gate must treat as a failure. lastpipe also fails closed, because grep is then still a pipeline stage that pipefail and set -e can see. The script stops at the pipeline with grep's status and never prints a verdict:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ bash -O lastpipe gate-pipe.sh locked.log
grep: locked.log: Permission denied

$! names only the most recent process substitution. A command with two of them, such as a drift check with diff <(...) <(...), loses the first status. Here one input is missing and the other has no failed logins, and diff reports "identical" with exit 0. Opening each substitution on its own file descriptor with exec {fd}< <(...) gives you one PID per producer to wait for:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ diff <(grep -F Failed -- rotated.log) <(grep -F Failed -- clean.log) echo "diff exit: $?"
grep: rotated.log: No such file or directory diff exit: 0
$ exec {a}< <(grep -F Failed -- rotated.log); pa=$! exec {b}< <(grep -F Failed -- clean.log); pb=$! diff "/dev/fd/$a" "/dev/fd/$b"; echo "diff exit: $?" wait "$pa"; echo "producer a: $?" wait "$pb"; echo "producer b: $?"
grep: rotated.log: No such file or directory diff exit: 0 producer a: 2 producer b: 1

Producer a exited 2 (missing file) and b exited 1 (no match), so this check can report that its first input was never read. Process substitution is a Bash feature, not POSIX sh; dash has none of this.

File descriptors you open yourself

A file descriptor (fd) is a small number that refers to an open file, pipe or socket in one process. exec with only redirections changes the current shell's descriptors for the rest of the script, and the {name}>file form lets Bash pick a free number (10 or above) and store it in $name, so you never collide with an fd some other code already uses. This sync job takes a lock with flock on one descriptor, writes its log through another, and starts a helper that keeps running after the job exits:

nightly-sync.sh
#!/usr/bin/env bash
# Nightly sync: one run at a time, guarded by flock on a lock file descriptor.
set -euo pipefail
state=${XDG_STATE_HOME:-$HOME/.local/state}/nightly-sync
mkdir -p -- "$state"
exec {lock}>"$state/lock" # Bash picks a free fd (10 or above) and stores it in $lock
if ! flock -n "$lock"; then
echo "nightly-sync: another run holds the lock" >&2
exit 75 # EX_TEMPFAIL: try again later
fi
exec {log}>>"$state/run.log"
printf 'run %s: started\n' "$$" >&"$log"
# A helper that keeps running after this run exits.
./warm-cache.sh &
printf 'run %s: helper %s started\n' "$$" "$!" >&"$log"
echo "nightly-sync: run $$ done"
warm-cache.sh
#!/usr/bin/env bash
# Stand-in for a slow background helper (a cache warmer, a log shipper).
sleep 20

The helper runs for 20 seconds, and that is the window for the next three commands. Run the job twice, the second straight after the first (paste both lines together):

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./nightly-sync.sh ./nightly-sync.sh
nightly-sync: run 529728 done nightly-sync: another run holds the lock
# exit status 75

The first run finished and exited. The second run is refused with 75, the "try again later" code, although nothing is running the sync any more. Within those 20 seconds, look at the helper's open descriptors in /proc (if the helper has already finished, find complains that the directory does not exist: run the two sync commands again and repeat):

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ find /proc/"$(pgrep -u "$USER" -x -f "bash ./warm-cache.sh")"/fd -mindepth 1 -printf "%f -> %l\n"
0 -> /dev/null 1 -> /dev/pts/0 (deleted) 2 -> /dev/pts/0 (deleted) 10 -> /home/deploy/.local/state/nightly-sync/lock 11 -> /home/deploy/.local/state/nightly-sync/run.log 255 -> /home/deploy/scr-bash-internals/warm-cache.sh
$ cat ~/.local/state/nightly-sync/run.log
run 529728: started run 529728: helper 529731 started

A child process inherits every open descriptor of its parent that is not marked close-on-exec ("Files, descriptors and the VFS" in Advanced Linux internals covers the descriptor table behind this). The helper holds fd 10 on the lock file and fd 11 on the log. A flock lock belongs to the open file, not to the process that asked for it, so it is released only when every descriptor that shares it is closed. The helper keeps it for as long as it runs. (Fd 255 is Bash keeping its own script open so it can read the next line.) The fix closes both descriptors for that one command only:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ diff nightly-sync.sh nightly-sync2.sh
16c16 < ./warm-cache.sh & --- > ./warm-cache.sh {lock}>&- {log}>&- & # the helper does not get the lock or log fds

Before you try the fixed script, wait until the first helper has finished, because its inherited descriptor still holds the lock. This loop waits (at most 40 seconds) until no helper is left:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ timeout 40 bash -c 'while pgrep -u "$USER" -x -f "bash ./warm-cache.sh" > /dev/null; do sleep 1; done' echo "no helper left"
no helper left
$ ./nightly-sync2.sh ./nightly-sync2.sh
nightly-sync: run 529923 done nightly-sync: run 529927 done

Two runs back to back now both get the lock, while their helpers keep running. Close lock and log descriptors on any command you start in the background, and on long-running commands that must not keep the lock alive after the script is gone.

The fix also means two helpers can now run at the same time, and nothing waits for their exit status. If the helper touches shared state, give it its own lock (flock -n helper.lock ./warm-cache.sh). Under systemd a helper started with & does not outlive the job at all: when the main process of a Type=oneshot unit exits, systemd stops the unit and kills everything left in its control group. A helper that must outlive the job belongs in its own unit or its own scheduled job.

Process groups: stop the whole tree

Killing a process does not kill its children. check-host.sh stands in for a slow remote check; it runs sleep as a child. Stopping it by PID leaves that child behind:

check-host.sh
#!/usr/bin/env bash
# Stand-in for a slow remote check: sleeps, then exits with the given status.
host=$1 secs=$2 status=$3
sleep "$secs"
echo "$host: checked (exit $status)"
exit "$status"
deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./check-host.sh demo 30 0 >/dev/null 2>&1 & sleep 0.5 kill "$!" sleep 0.5 pgrep -a -u "$USER" -x sleep
530176 sleep 30
$ pkill -u "$USER" -x -f "sleep 30"; sleep 0.3; pgrep -u "$USER" -x -f "sleep 30" || echo "orphan removed"
orphan removed

The script died and its sleep 30 kept running, an orphan that nothing would ever wait for, until pkill removed it (-x -f matches the whole command line exactly, so it cannot hit anything else). With real work in place of sleep this is a scan that keeps hitting a host after you cancelled it.

Every process belongs to a process group, and kill with a negative number signals the whole group. setsid starts a program in a new session whose process group ID is its own PID. How you learn that ID depends on job control. In a script (job control off), setsid cmd & does not need to fork, so $! is the group ID; sweep.sh below relies on that. At an interactive prompt (job control on), & already puts the job in its own group, setsid forks, and $! names a parent that has already exited, so kill -- "-$!" misses. The form below works in both: setsid -f always forks, and the new group leader writes its own PID to a file before it becomes the check:

deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ setsid -f bash -c 'echo "$$" > check.pid; exec ./check-host.sh demo 30 0' > /dev/null 2>&1 timeout 5 bash -c 'until [ -s check.pid ]; do sleep 0.1; done' pgid=$(cat check.pid) ps -o pid,pgid,sid,comm -s "$pgid" kill -TERM -- "-$pgid" sleep 0.5 pgrep -u "$USER" -a -x -f 'sleep 30' || echo "no sleep 30 left" rm check.pid
PID PGID SID COMMAND 530239 530239 530239 bash 530243 530239 530239 sleep no sleep 30 left
# COMMAND shows bash for check-host.sh: the interpreter running the script

ps shows the check and its sleep with the same PGID (process group) and SID (session), both equal to the PID from check.pid. One kill -TERM -- "-$pgid" reached both; the -- stops kill from reading the negative number as an option. The last check looks for the child by its command line rather than by the group ID you assumed, so it would catch a survivor even if the group ID were wrong.

The same idea makes parallel work safe. wait -n (Bash 5.1 and later with a list of PIDs and -p) returns as soon as any one job finishes, with that job's status, and stores its PID in the variable you name. This sweep runs three checks at once, stops at the first failure and cleans up the rest through their process groups from a single EXIT trap:

sweep.sh
#!/usr/bin/env bash
# Runs host checks in parallel. The first failure stops the rest, and nothing is left running.
set -euo pipefail
declare -A host_of=() # PID of each running check -> host name
start() {
# setsid: the check leads its own process group, so one kill reaches everything it started
setsid ./check-host.sh "$@" &
host_of[$!]=$1
}
# shellcheck disable=SC2329 # called by the EXIT trap; ShellCheck 0.11 misses that
stop_all() {
local pid
(( ${#host_of[@]} )) || return 0
echo "sweep: stopping ${#host_of[@]} unfinished check(s)"
for pid in "${!host_of[@]}"; do
kill -TERM -- "-$pid" 2>/dev/null || true
done
wait # reap them before this script exits
}
trap stop_all EXIT
start web01 1 0
start db01 2 3
start cache01 30 0
worst=0
while (( ${#host_of[@]} )); do
rc=0
wait -n -p done_pid -- "${!host_of[@]}" || rc=$?
echo "sweep: ${host_of[$done_pid]} finished with exit $rc"
unset "host_of[$done_pid]"
if (( rc != 0 )); then worst=$rc; break; fi
done
exit "$worst"
deploy@web01:~/scr-bash-internals · Ubuntu 26.04 LTS
$ ./sweep.sh
web01: checked (exit 0) sweep: web01 finished with exit 0 db01: checked (exit 3) sweep: db01 finished with exit 3 sweep: stopping 1 unfinished check(s)
$ pgrep -a -u "$USER" -x sleep

db01 failed with 3, the loop stopped, and the EXIT trap killed cache01's group and reaped it with wait. The sweep exits with the failing check's status and pgrep finds no sleep left. The # shellcheck disable=SC2329 line is there because ShellCheck 0.11 reports a function called only from an EXIT trap as "never invoked" when the script ends with an explicit exit; a directive with its reason is the honest fix. One side effect of setsid to keep in mind: the checks are no longer in your terminal's process group, so Ctrl-C reaches only the sweep, and the sweep's trap has to stop them.

Coprocesses: know when not to
coproc starts one background command with a pipe to its input and one from its output, so a script can hold a conversation with it. It works for a short exchange with a helper that flushes every line (fflush() in awk, stdbuf -oL for many tools). The limits arrive quickly: a helper that buffers its output deadlocks your read (use read -t to bound it), and nothing checks the helper's exit status unless you wait "$NAME_PID". If a script needs a long-lived helper it talks to over pipes, it has outgrown Bash; "When a script has outgrown Bash" later in this course covers that handoff.

Try this

Run ./gate.sh rotated.log for a log that does not exist and predict the exit status before you look. Then break the fix on purpose: copy nightly-sync2.sh to try-sync.sh, remove only {lock}>&- from the helper line, wait until no helper is left (the loop above), and run ./try-sync.sh && ./try-sync.sh. Confirm the cause with the find /proc/... command from above. Finally, once the helpers have finished, run pgrep -a -u "$USER" -x sleep.

Expected: the gate exits 2, with grep's "No such file or directory" above the gate's refusal. The second try-sync.sh exits 75 again, and find shows fd 10 on the lock file back in the helper. pgrep prints nothing and exits 1.

Takeaway

Every place Bash starts a process is a place where a status, a variable or a descriptor can go missing: collect producer statuses with wait "$!", close lock descriptors on background commands, and stop work by process group, then prove with pgrep that nothing is left.

Quick check
01A CI gate runs while read -r l; do n=$((n+1)); done < <(trivy-to-lines report.json) under set -euo pipefail and blocks when n is above 0. One night the report file is missing, the converter exits 1, and the gate passes. What is the smallest correct fix?
Incorrect — inherit_errexit only affects command substitution $(...). The converter runs in a process substitution, and errexit never looks at its status whatever the option says.
Incorrect — Even if the trap runs inside the substitution, it runs in that child process. The parent shell never receives the producer's status, so the gate still passes.
Correct — Bash 4.4+ sets $! to the last process substitution and wait returns its exit status, so the gate can tell "no findings" from "never read the report".
Incorrect — Without lastpipe the loop then runs in a subshell and n is lost, so the gate always sees 0; this swaps one fail-open bug for another.
02A backup script takes flock -n on exec {lock}>/run/backup.lock, then starts ./upload.sh & and exits. The next scheduled run fails with "lock held" although no backup script is running. What is holding the lock?
Correct — the lock belongs to the open file, and the child's inherited fd keeps it alive. Start the child with {lock}>&-.
Incorrect — There is no grace period: the lock goes away the moment the last descriptor sharing it is closed.
Incorrect — Truncating the file has no effect on locking; flock locks are not stored in the file's content.
Incorrect — An explicit release is not required: closing all descriptors releases the lock, which is why the script exiting would have been enough without the child.
03A wrapper starts ./scan.sh & (which runs nmap as a child) and on timeout runs kill "$scan_pid". Operators report scans still hitting hosts after the timeout. Which change fixes it?
Incorrect — SIGKILL still goes to one PID only. The script dies at once and its nmap child is orphaned exactly as before.
Incorrect — Waiting reaps the script you killed; it does nothing to the nmap child, which keeps running after the wrapper returns.
Incorrect — kill 0 signals the caller's whole process group. Started with & from a script, scan.sh shares the wrapper's group, so this also kills the wrapper.
Correct — in a script (job control off) setsid does not fork, so $! is the scan's own group, and a negative PID signals that whole group.

Related