Bash strict mode: writing ops scripts that fail loudly

set -euo pipefail is the start, not the whole story. Traps, quoting, and safe temp files for scripts that run as root.

Oct 28, 2025·Updated ·6 min readBeginner·By SecOpsLog · command-tested
bash — observed: the same typo, with and without strict mode (bash 5.2, in a container holding 20 files under /home/deploy)observed
bash deploy-loose.sh # cd /srv/relese; rm -rf ./*; echo "deploy finished"
deploy-loose.sh: line 3: cd: /srv/relese: No such file or directory
deploy finished
exit status 0; 0 files left in /home/deploy
cd failed, the script kept going, and rm -rf ./* ran in the directory it started in
bash deploy-strict.sh # the same three lines after set -euo pipefail
deploy-strict.sh: line 3: cd: /srv/relese: No such file or directory
exit status 1; 20 files left in /home/deploy

Bash's defaults were designed for an interactive shell, where carrying on after a failed command is the friendly choice. In a script that runs as root, or in a pipeline that deploys, those defaults are a way of failing quietly: a failed cd is ignored, an unset variable expands to an empty string, and a pipeline's status is whatever its last stage returned. Three flags reverse each of those, and understanding where they do not apply matters as much as setting them.

The header, and what each flag decides

strict.sh
#!/usr/bin/env bash
set -euo pipefail
shopt -s inherit_errexit # bash 4.4+: -e also applies inside $(...) substitutions
IFS=$'\n\t' # word-split on newline and tab only, never on spaces
# -e a command that fails (non-zero) ends the script
# -u an unset variable is an error, not an empty string
# -o pipefail a pipeline fails if any stage fails, not only the last one

-u is the flag that catches rm -rf "$BUILD_DIR/" when BUILD_DIR was never exported: without it the expression is rm -rf /. -o pipefail is what makes curl … | tar xz fail when the download fails rather than when the archive turns out to be a 404 page. -e is the one with exceptions, and the exceptions follow a pattern:

Where set -e does not fire

ContextWhyWhat to do instead
the condition of if, while, untilthe exit status is the testnothing; this is the intended use
any command except the last in a && or || listthe list itself is the testwrite cmd || exit 1 deliberately, or split the list
a command whose status is negated with !same rulecheck explicitly
inside $(…) unless inherit_errexit is set-e is off in the subshell by defaultshopt -s inherit_errexit; still check the assignment’s result
a function called from any of the contexts abovethe whole function inherits the exemptionavoid running critical functions inside conditions
local x=$(cmd) or export x=$(cmd)the status of local masks the command’sdeclare first, assign on the next line

The last row is the one that surprises people who already know the others: local out=$(risky) succeeds because local succeeded, whatever risky returned. For the handful of commands whose failure the script cannot tolerate, an explicit check reads better than trusting the flag, and set -E with a trap … ERR gives a single place to log the failing line before the script exits.

bash — observed: one probe per table row, each in a child bash with set -euo pipefailobserved
bash exceptions.sh # prints whether the line after the failure was reached
plain failing command exit=1
condition of if exit=0 reached-next-line
not-last in && list exit=0 reached-next-line
negated with ! exit=0 reached-next-line
inside $(...) without inherit_errexit exit=0 reached-next-line
inside $(...) with inherit_errexit exit=1
function called in an if condition exit=0 inside-f-after-false reached-next-line
local x=$(false) exit=0 after-local reached-next-line
local x; x=$(false) exit=1
pipeline, last stage ok, no pipefail exit=0 reached-next-line
pipeline, last stage ok, pipefail exit=1
every row in the table behaves as described on 5.2.37; the function row shows the whole body running past its own failure

Cleanup that runs however the script leaves

With -e the script can exit from any line, so cleanup at the bottom is cleanup that is skipped. A trap on EXIT runs on normal completion, on an error exit and on Ctrl-C, which makes it the place for removing temporary files and releasing locks. Use mktemp for the paths it removes, so two runs never collide on a hard-coded /tmp/myapp and an attacker cannot pre-create the path.

strict.sh (continued)
tmp="$(mktemp -d)"
trap 'rm -rf -- "$tmp"' EXIT # runs on success, error and interrupt
trap 'echo "failed at line $LINENO: $BASH_COMMAND" >&2' ERR
set -E # ERR trap fires in functions and subshells too
curl -fsSL "$URL" -o "$tmp/archive.tgz"
tar -xzf "$tmp/archive.tgz" -C "$tmp"

Quoting and arrays

An unquoted expansion is split on IFS and glob-expanded before the command sees it, so a path with a space becomes two arguments and a value of * becomes every file in the directory. Quote every expansion, including the ones inside $(…), and build argument lists as arrays, not space-joined strings, because an array is the only Bash structure that can carry an argument containing a space. None of this validates input: strict mode stops your typos, and it does nothing about a value that arrives from outside containing ; rm -rf / if some later line passes it to eval or an unquoted sh -c.

quoting.sh
# splits and globs on the contents of $path
rm -rf $path/*
# quoted scalar, array for the argument list
rm -rf -- "$path"/*
args=(--config "app prod.conf" --force)
mycmd "${args[@]}"
# read lines safely, whatever they contain
while IFS= read -r line; do
printf '%s\n' "$line"
done < "$input"

Let a machine review it

ShellCheck finds the local masking a status, the unquoted expansion, the [ ] test that should be [[ ]] and the cd without a check, at review time rather than at 03:00, but not all at the same level, which decides what a gate sees. On a nine-line deploy script with all four mistakes, 0.11.0 at --severity=warning reported only SC2155 (the local); the unquoted cd $1 is SC2086 at info level, [ ] versus [[ ]] is the optional check require-double-brackets (off unless -o enables it), and SC2164 for the unchecked cd is suppressed when the script sets -e, because the shell already handles it. Running it in CI on every script under bin/ and ci/ is the only enforcement mechanism that survives staff turnover; run it at --severity=info (or the default, style) so the quoting findings count, and add -o require-double-brackets if that rule matters to the team. It is a static check, so it catches the mistakes in the table above that strict mode cannot see at runtime; logic errors are still the runtime flags' job.

bash — observed: ShellCheck 0.11.0 on the deploy script, at three severitiesobserved
shellcheck --severity=warning ci/deploy.sh; echo exit=$?
In ci/deploy.sh line 5:
local out=$(build)
^-^ SC2155 (warning): Declare and assign separately to avoid masking return values.
exit=1
shellcheck --severity=error ci/deploy.sh; echo exit=$?
exit=0
shellcheck --severity=info -f gcc ci/deploy.sh
ci/deploy.sh:5:9: warning: Declare and assign separately to avoid masking return values. [SC2155]
ci/deploy.sh:6:6: note: Double quote to prevent globbing and word splitting. [SC2086]
shellcheck -s sh deploy-strict.sh
^------^ SC3040 (warning): In POSIX sh, set option pipefail is undefined.
the review comment nobody would have written, on every push; a gate at warning level would have let the unquoted cd $1 through

When a strict script fails in production: recovery in order of preference

SituationDoDo not
a script that used to work now exits at a line that always returned non-zero (grep with no match, diff, a status probe)make the exception explicit on that line: grep … || true, or if grep …; then, and keep -e onremove set -e from the script; every other line loses its check
-u stops a script on a variable that is legitimately optional"${VAR:-}" or "${VAR:-default}" at the point of useset +u at the top; the rm -rf "$BUILD_DIR/"* case comes back
the EXIT trap removed something it should have kept (a log, a partial artifact)move the path out of the trap and into a deliberate rm at the end of the success path, so a failed run keeps it for the post-mortemremove the trap; temp directories then survive every failure
What was run for this article
GNU bash 5.2.37 (the bash:5.2 image) and ShellCheck 0.11.0 (koalaman/shellcheck:v0.11.0), Docker Engine 28.5.2, linux/arm64. The terminal blocks marked observed are copied from that run: the cd-then-rm scenario in a container directory holding twenty files, one probe per row of the blind-spot table, the unset BUILD_DIR glob (printed, not executed: it expands to the root directory’s entries), the EXIT and ERR traps, and ShellCheck at warning, error, info, with the optional double-bracket check and as POSIX sh. Twelve exit codes are asserted. The deploy script, the CI gate and the 03:00 incident are representative, and the recovery table follows Bash’s documented semantics rather than a run.
Strict mode is Bash-only
A script that starts with #!/bin/sh runs under dash on Debian and Ubuntu, where pipefail does not exist (it is in POSIX 2024 but not in every shell yet) and inherit_errexit is meaningless. Use the bash shebang when you rely on these flags, and let ShellCheck know which shell it is checking for.
Go deeper in a courseBash for ops, done safelyStrict mode, traps, quoting, testing and the scripts that run as root.View course

Related posts

Quick reference