Error propagation you can trust

What errexit cannot see, stack traces from ERR traps, and gates that fail closed.

Advanced20 min · lesson 2 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (8 files, 2 KB): scr-bash-errors.tar.gz. Unpack it with tar -xzf scr-bash-errors.tar.gz, which creates scr-bash-errors/. SHA-256: ecd524ee2d055b66ac00e818b97c9bdd81049f56590c6b2925e42faaf41d4ee6

This lesson is about making a failure in a Bash script reach the decision that depends on it. You will watch a strict-mode blocklist refresh install an empty list and report success, rebuild it so that every failure you can name is checked explicitly and leaves the old list in place, and add an ERR trap that prints a call stack for the failures you did not expect. The result is a job that either does its work or exits non-zero with a message that says why.

Refresher: "Strict mode, honestly" in bash-ops owns the rules this lesson builds on: what -e, -u, pipefail and inherit_errexit catch, the contexts where -e is switched off, PIPESTATUS, local x=$(...) masking, and the ERR trap with set -E. Here they meet a program with functions, substitutions and more than one failure path.

Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-errors. The examples download from a small local web server that serves the feed/ directory. Start it in the background and keep its process ID in a file, so you can stop it at the end of the lesson:

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ python3 -m http.server 18702 --bind 127.0.0.1 --directory feed > server.log 2>&1 & echo $! > server.pid timeout 10 bash -c 'until curl -fs -o /dev/null http://127.0.0.1:18702/blocklist.txt; do sleep 0.2; done' && echo "feed server is up"
feed server is up

A strict script that still fails open

Here is a blocklist refresh written the way many production scripts are: strict mode on, inherit_errexit on, small functions, an error message if the install fails. It passes ShellCheck's default checks with no findings.

refresh-naive.sh
#!/usr/bin/env bash
# Replaces the installed blocklist with a fresh download. Strict mode, and it still fails open.
set -Eeuo pipefail
shopt -s inherit_errexit
url=${1:?usage: refresh-naive.sh URL LIST}
list=${2:?usage: refresh-naive.sh URL LIST}
fetch() {
curl -fsS --max-time 10 "$url"
}
install_list() {
printf '%s\n' "$(fetch)" > "$list.new"
mv -- "$list.new" "$list"
echo "installed $(grep -c . "$list") entries"
}
if ! install_list; then
echo "refresh failed, keeping the old list" >&2
exit 1
fi

Point it at a URL that returns 404 (the feed was moved), with a copy of the installed list so we can see what happens to it:

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ cp lists/blocklist.txt naive-list.txt
$ ./refresh-naive.sh http://127.0.0.1:18702/missing.txt naive-list.txt
curl: (22) The requested URL returned error: 404 installed 0 entries
$ wc -c < naive-list.txt
1

curl reported the 404 and exited 22. The script then installed a list that is a single newline, printed "installed 0 entries" and exited 0. A firewall loading that file now blocks nothing, and the job that ran the refresh recorded a success.

Both causes are rules from "Strict mode, honestly". printf '%s\n' "$(fetch)" uses the substitution as an argument, so its status is thrown away even with inherit_errexit; only a plain assignment such as body=$(fetch) returns it. And if ! install_list calls the function as a condition, which switches -e off for every command inside it, so nothing in the install could stop the run. Either one alone would have been enough. ShellCheck's default checks accept both on purpose, because both forms are legal and sometimes intended; two optional checks from "Check your script" in bash-ops name them:

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ shellcheck -f gcc -o check-set-e-suppressed,check-extra-masked-returns refresh-naive.sh
refresh-naive.sh:12:20: note: Consider invoking this command separately to avoid masking its return value (or use '|| true' to ignore). [SC2312] refresh-naive.sh:14:21: note: Consider invoking this command separately to avoid masking its return value (or use '|| true' to ignore). [SC2312] refresh-naive.sh:17:6: note: This function is invoked in a ! condition so set -e will be disabled. Invoke separately if failures should cause the script to exit. [SC2310] refresh-naive.sh:17:6: note: This function is invoked in an 'if' condition so set -e will be disabled. Invoke separately if failures should cause the script to exit. [SC2310]

SC2312 marks each substitution whose status is masked and SC2310 the function called as a condition. Turn them on for scripts that depend on errexit, and treat each hit as a question: is this failure meant to be ignored here?

The second cause is the one that grows with a program. Any function that does real work and relies on errexit to stop early is correct only until someone writes if my_function; then, my_function || rollback or validate && apply. Call such functions as plain statements, or write them so each step that can fail checks itself (cmd || return, or || refuse with a function that exits), so the function never depends on errexit. The same goes for done < <(cmd): "Processes, pipes and file descriptors in real scripts" showed how to collect cmd's status with wait "$!".

Explicit checks: a refresh that fails closed

Strict mode is a backstop for mistakes. The failures you can name in advance (download failed, empty download, garbage in the download) deserve explicit checks with their own messages, placed where no condition context can disable them. Failing closed means the job keeps the last known-good list and exits non-zero whenever it cannot prove the new one is good.

Refresh that fails closed
1mktemp next to the list
same filesystem, so mv is a rename
2curl -f with timeouts
HTTP error or hang: refuse
3Count entries and bad lines
none or any bad: refuse
4mv over the old list
readers see old or new, never half
Every refuse exits 1 and leaves the installed list untouched; the EXIT trap removes the temp file.
refresh.sh
#!/usr/bin/env bash
# Refreshes the installed blocklist and fails closed: the old list stays unless the
# download succeeds and every line is a comment or an IPv4 address/CIDR.
# Exit: 0 installed; non-zero: nothing installed (a trace on stderr marks an unplanned failure).
set -Eeuo pipefail
shopt -s inherit_errexit
# shellcheck source=trace.sh
source "${BASH_SOURCE[0]%/*}/trace.sh"
url=${1:?usage: refresh.sh URL LIST}
list=${2:?usage: refresh.sh URL LIST}
entry_re='^([0-9]{1,3}\.){3}[0-9]{1,3}(/[0-9]{1,2})?$'
tmp=
trap '[[ -z $tmp ]] || rm -f -- "$tmp"' EXIT
refuse() { echo "refresh: $*; keeping $list" >&2; exit 1; }
count() { # count GREP-ARGS...: prints the number of selected lines; "none" (grep 1) is a valid 0
local rc=0
grep -c "$@" || rc=$?
(( rc <= 1 )) || return "$rc"
}
refresh_list() {
local entries invalid
tmp=$(mktemp -- "$list.XXXXXX") # same directory as the list: mv is a rename
curl -fsS --connect-timeout 3 --max-time 10 -o "$tmp" "$url" || refuse "download failed"
entries=$(count -E -e "$entry_re" -- "$tmp")
invalid=$(count -Ev -e '^#' -e "$entry_re" -- "$tmp")
(( entries > 0 )) || refuse "no entries in download"
(( invalid == 0 )) || refuse "$invalid invalid line(s) in download"
chmod 0644 -- "$tmp"
mv -f -- "$tmp" "$list"
echo "refresh: installed $entries entries"
}
refresh_list

Three details carry the design. count wraps grep -c so that "no lines matched" (grep's 1) is a valid count of 0, and any status above 1 is returned as an error. Both counts are plain assignments, so an error inside count stops the script. The download goes to a temporary file in the list's own directory, so the final mv replaces the list in one step; "Temp files, locks and timeouts" in bash-ops covers why. The curl ... || refuse form is safe because refuse exits: the || handles the failure instead of hiding it.

Two smaller choices matter as much. --connect-timeout and --max-time turn a feed server that accepts the connection and never answers into a failure after at most ten seconds, instead of a job that hangs until something else kills it; a hang is a silent failure, because nothing reports it. And the exit status is the contract with whatever runs the job: cron, a systemd timer or a CI step only see 0 or non-zero, so every refusal exits 1 and writes its reason to stderr, where the journal or the CI log keeps it.

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ ./refresh.sh http://127.0.0.1:18702/blocklist.txt lists/blocklist.txt
refresh: installed 4 entries
$ cat lists/blocklist.txt
# lab blocklist (documentation address ranges only) 203.0.113.7 203.0.113.9 198.51.100.0/24 192.0.2.44
$ ./refresh.sh http://127.0.0.1:18702/missing.txt lists/blocklist.txt
curl: (22) The requested URL returned error: 404 refresh: download failed; keeping lists/blocklist.txt
$ ./refresh.sh http://127.0.0.1:18702/mixed.txt lists/blocklist.txt
refresh: 1 invalid line(s) in download; keeping lists/blocklist.txt
$ sha256sum feed/blocklist.txt lists/blocklist.txt; ls lists/
05f32db66239168ac1b08a47aaf7cf08edbd256841873da0eb6aceb941aab7ca feed/blocklist.txt 05f32db66239168ac1b08a47aaf7cf08edbd256841873da0eb6aceb941aab7ca lists/blocklist.txt blocklist.txt

The good feed installs four entries. The 404 and the download with a proxy error page in the middle are both refused with a reason, and the two identical hashes show that the installed list is still the good one. ls shows that no temporary file was left behind.

Be exact about what these checks prove: the download is well-formed. They do not prove it is sensible. The pattern accepts 0.0.0.0/0 (which would block everything) and 999.999.999.999, and a feed that shrinks from ten thousand entries to one passes. A production job adds plausibility gates: a minimum entry count or a maximum shrink against the current list ("Designing a CLI tool" builds that check into blsync as max_remove_fraction), a floor on prefix length, an allowlist of your own ranges that may never be blocked, a real address parser such as Python's ipaddress, and a feed fetched over HTTPS with a checksum or signature.

Stack traces for the failures you did not plan for

An explicit check covers what you predicted. For everything else, an ERR trap can report where the script died. With set -E the trap is inherited by functions and subshells, which has a side effect worth seeing once:

dup-trace.sh
#!/usr/bin/env bash
# A plain ERR trap under set -E: count how often it fires for one failure.
set -Eeuo pipefail
shopt -s inherit_errexit
trap 'echo "ERR (subshell level $BASH_SUBSHELL): $BASH_COMMAND" >&2' ERR
list=$(ls /nonexistent)
echo "never printed: $list"
deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ ./dup-trace.sh
ls: cannot access '/nonexistent': No such file or directory ERR (subshell level 1): ls /nonexistent ERR (subshell level 0): list=$(ls /nonexistent)

One failure, two reports: once inside the $(...) subshell (level 1), where ls failed, and again in the main shell, where the assignment failed. In a script with nested substitutions this multiplies. The trap library below reports only from the main shell, and walks the call stack with caller, which prints the line, function and file of each active call (frame 0 is the closest one):

trace.sh
# shellcheck shell=bash
# trace.sh: source from a script that uses set -Eeuo pipefail. On an unexpected failure it
# prints the failing command and the call stack, then exits with the failing status.
trace_err() {
local rc=$? i=0 frame line func file
# set -E runs this trap inside $(...) too; report once, from the main shell.
(( BASH_SUBSHELL == 0 )) || exit "$rc"
printf 'error: %s exited %d\n' "$BASH_COMMAND" "$rc" >&2
while frame=$(caller "$i"); do
read -r line func file <<<"$frame"
printf ' at %s:%s in %s\n' "$file" "$line" "$func" >&2
i=$((i + 1))
done
exit "$rc"
}
trap trace_err ERR

refresh.sh sources it near the top. To trigger a failure nobody planned for, make the list's directory read-only so that mktemp cannot create its file:

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ chmod 555 lists
$ ./refresh.sh http://127.0.0.1:18702/blocklist.txt lists/blocklist.txt
mktemp: Permission denied (os error 13) at path "/home/deploy/scr-bash-errors/lists/blocklist.txt.y6gMmx" error: tmp=$(mktemp -- "$list.XXXXXX") exited 1 at ./refresh.sh:26 in refresh_list at ./refresh.sh:36 in main
$ chmod 755 lists

The first line is the error from uutils mktemp (the default mktemp on Ubuntu 26.04), the second names the command that failed and its status, and the stack reads from the failure outward: line 26 inside refresh_list, which was called from line 36 of the script body. Bash calls the top level main, which is why the refresh function is not named that. Nothing was installed, and the exit status is non-zero.

Try this

Serve an empty file (feed/empty.txt is already there) and run ./refresh.sh http://127.0.0.1:18702/empty.txt lists/blocklist.txt: it must exit 1 with "no entries in download". Then point it at a port where nothing listens, http://127.0.0.1:18703/blocklist.txt: curl fails with "(7) Failed to connect" and the refresh exits 1. For the timeout path you need a server that accepts connections and never answers. A Python socket that listens and never reads is enough, because the kernel completes the connection for it: python3 -c 'import socket, time; s = socket.create_server(("127.0.0.1", 18704)); time.sleep(60)' & echo $! > hang.pid. Refresh from port 18704: after ten seconds curl exits with "(28) Operation timed out" and the refresh refuses. Stop the listener with kill "$(cat hang.pid)". After all three, sha256sum lists/blocklist.txt must print the same hash as before and ls lists/ must show only blocklist.txt.

When you are done with the lesson, stop the feed server:

deploy@web01:~/scr-bash-errors · Ubuntu 26.04 LTS
$ kill "$(cat server.pid)" && rm server.pid

Takeaway

Check every failure you can name explicitly, at a place where the check cannot be switched off by a condition, and keep set -e plus an ERR stack trace for the ones you cannot; a job that cannot prove its result is good must leave the old state and exit non-zero.

Quick check
01A script has set -Eeuo pipefail, shopt -s inherit_errexit and trap 'echo "failed: $BASH_COMMAND" >&2' ERR. It runs cfg=$(cat /etc/app/missing.conf) and the file is missing. How many "failed:" lines appear?
Incorrect — That is the behaviour without set -E. With -E the trap is inherited by command substitutions, so the subshell reports too.
Incorrect — A plain assignment is not an ignored context; its status is the substitution's, and a failure there triggers ERR.
Correct — one failure is reported once per shell level it passes through, which is why trace.sh reports only when BASH_SUBSHELL is 0.
Incorrect — The assignment takes the substitution's status, and a failing assignment in the main shell triggers ERR there as well.
02refresh_list checks its download with (( entries > 0 )) || refuse "no entries" (refuse exits 1) and relies on errexit for everything else, such as a failing mktemp. A teammate now calls it as if refresh_list; then notify_ok; fi. What still protects the run?
Incorrect — Called as an if condition, the function runs with errexit off from start to finish; the header does not change that.
Correct — refuse exits whatever the context, while errexit and the ERR trap are both switched off inside a tested function.
Incorrect — exit ends the shell even inside a condition (it would only end a subshell), so the explicit checks keep working.
Incorrect — set -E controls inheritance, not the ignored contexts; in a tested function the ERR trap does not fire at all.
03A blocklist refresh downloads with curl -fsS -o "$tmp" "$url" and has explicit checks for a failed download, an empty file and malformed lines. One night the feed server accepts the connection and never sends a byte. What does the job do?
Correct — curl waits on an open connection for as long as it stays open, and a job that never exits reports nothing to cron, the timer or CI.
Incorrect — curl has no default limit on a transfer; only the connection phase has one, and here the connection succeeded.
Incorrect — -f acts on a status code of 400 or more; with no answer there is no status to judge, so it never fires.
Incorrect — Bash has no command timeout; errexit reacts to an exit status, and a command that never returns never produces one.

Related