Testing shell automation

Bats with fakes, testing signals and timeouts, and ShellCheck and Bats as CI gates.

Advanced35 min · lesson 5 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (11 files, 5 KB): scr-bash-test.tar.gz. Unpack it with tar -xzf scr-bash-test.tar.gz, which creates scr-bash-test/. SHA-256: 82a57d667f50565882c69dc520cdaa9bb7c2da05b36422276deed888b22f3690

This lesson tests the shell automation you have written in the rest of this course: the failure paths, the signal handling, the timeouts and the cleanup. You will make a script testable so a test can load it without running it, write assertions with bats-assert, fake the command that touches the network (and check the fake against the real thing), prove that a job cleans up and returns 143 on SIGTERM, read a failing test, and gate a change in CI with pinned, checksum-verified tools and a workflow that actionlint has checked.

Refresher: "Check your script" in bash-ops owns ShellCheck (severities, directives, exit status 1), shfmt with .editorconfig, and your first Bats file (run, $status, $output, BATS_TEST_TMPDIR). This lesson is only the delta. It uses the archive's bats 1.13 with bats-assert and bats-support (sudo apt install bats bats-assert bats-support if bash-ops did not install them already). Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-test; they include status-server.py, a tiny local HTTP server that answers each request with the status code named in its path, used once below.

Design for the test first

A test cannot exercise one function of a script if loading the file runs the whole thing, and it cannot load the file at all if loading it turns on options or connects to a service. So put the runtime work, and the set -Eeuo pipefail that goes with it, inside main, and guard main so it runs only on direct execution:

healthcheck.sh
#!/usr/bin/env bash
# Check one HTTP endpoint and classify it. Written to be testable: the logic is in
# small functions, and the file does its work (and turns on strict mode) only when
# executed directly, so a test can source it without running main or changing the
# test shell's options. Exit: 0 up, 1 down, 2 unknown or unreachable.
http_status() { # http_status URL -> the HTTP status code, or 000 if nothing answered
# No -f: with --fail, curl exits 22 on any 4xx/5xx, and a 503 would look like
# "could not connect". Without it, curl prints the code and exits 0.
curl -sS -o /dev/null -w '%{http_code}' --max-time 5 -- "$1"
}
classify() { # classify CODE -> up | down | unknown
local code=$1
if ((code >= 200 && code < 300)); then
echo up
elif ((code >= 500)); then
echo down
else
echo unknown
fi
}
main() {
set -Eeuo pipefail
local url=${1:?usage: healthcheck.sh URL} code state
code=$(http_status "$url") || code=000
state=$(classify "$code")
printf '%s %s\n' "$state" "$code"
case $state in
up) return 0 ;;
down) return 1 ;;
*) return 2 ;;
esac
}
# BASH_SOURCE[0] is this file; $0 is what the shell is running. They match on direct
# execution and differ when a test sources the file, so main runs only for real.
if [[ ${BASH_SOURCE[0]} == "${0}" ]]; then
main "$@"
fi

The guard [[ ${BASH_SOURCE[0]} == "${0}" ]] compares the file the code lives in with the file the shell was told to run: they match on ./healthcheck.sh and differ when a test sources it, so main runs only for real. classify is a pure function with no I/O. The exit status carries the result (0 up, 1 down, 2 unknown or unreachable), so cron or a CI step can act on it. The comment in http_status about -f is there because an earlier version of this script had it, and the tests passed anyway; the contract check below shows why.

Assertions and fakes with bats-assert

bats-assert turns raw [ "$status" -eq 0 ] checks into readable assertions with useful failure messages: assert_success, assert_failure 1, assert_output, assert_equal. bats_load_library finds it in /usr/lib/bats, where bats 1.13 looks by default. The command that reaches the network, curl, is faked two ways:

checks.bats
#!/usr/bin/env bats
# Tests for healthcheck.sh. Run with: bats checks.bats
setup() {
bats_load_library bats-support
bats_load_library bats-assert
# Sourcing is safe: healthcheck.sh turns on strict mode only inside main().
source "$BATS_TEST_DIRNAME/healthcheck.sh"
}
@test "classify maps codes to up/down/unknown" {
run classify 204
assert_output "up"
run classify 503
assert_output "down"
run classify 404
assert_output "unknown"
}
@test "http_status passes the code through (fake curl as a function)" {
# A function beats a command on PATH, and exists only in this test shell.
curl() { "$BATS_TEST_DIRNAME/fakebin/curl" "$@"; }
FAKE_CODE=503 run http_status http://svc.invalid/
assert_success
assert_output "503"
}
@test "main: a healthy endpoint is up, exit 0 (fake curl on PATH)" {
PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_CODE=200 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/
assert_success
assert_output "up 200"
}
@test "main: a 503 is down, exit 1" {
PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_CODE=503 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/
assert_failure 1
assert_output "down 503"
}
@test "main: no answer is unknown 000, exit 2" {
PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_FAIL=1 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/
assert_failure 2
assert_output "unknown 000"
}
fake curl (fakebin/curl)
#!/usr/bin/env bash
# Test double for curl. It ignores the request and "answers" with FAKE_CODE, and it
# keeps real curl's exit rules so a test cannot pass on behaviour curl does not have:
# FAKE_FAIL=1 is a connection failure (prints 000, exit 7); with -f or --fail an
# answer of 400 or more exits 22 after printing the code.
fail=0
for arg in "$@"; do
case $arg in
--fail) fail=1 ;;
--*) ;;
-*f*) fail=1 ;; # -f alone or in a group such as -fsS
esac
done
if [[ ${FAKE_FAIL:-0} == 1 ]]; then
printf '000'
exit 7
fi
code=${FAKE_CODE:-200}
printf '%s' "$code"
if ((fail && code >= 400)); then exit 22; fi

The second test defines a curl function after sourcing: a function wins over a command on PATH, and it exists only in the test shell, so production has no hook for it. The last three run healthcheck.sh as a separate process, which cannot see the test's functions, so they put fakebin first on PATH. A common alternative, CURL=${CURL:-curl} in the script, works too, but then whoever controls the job's environment (a CI variable, an EnvironmentFile=) picks the binary, which is the PATH problem from "Trust boundaries". Run the suite the way CI does, without a terminal, so bats prints TAP:

deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ bats --tap checks.bats
1..5 ok 1 classify maps codes to up/down/unknown ok 2 http_status passes the code through (fake curl as a function) ok 3 main: a healthy endpoint is up, exit 0 (fake curl on PATH) ok 4 main: a 503 is down, exit 1 ok 5 main: no answer is unknown 000, exit 2

1..5 then one ok line per test is the TAP output a CI log shows; the ✓ form appears only on an interactive terminal.

A fake is only as good as its likeness to the real command. Real curl with -f (--fail) exits 22 on any status of 400 or more, after printing the code. An earlier fake printed the code and exited 0 whatever the flags, so a healthcheck.sh that used -fsS passed its "down 503" test while in production every 503 became unknown 000. That is why this fake implements the -f rule. Check a fake against the real thing once, with a local server that answers with the status in the path. Start it in the background from ~/scr-bash-test:

deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ python3 status-server.py 18731 > server.log 2>&1 & echo $! > server.pid timeout 10 bash -c 'until grep -q listening server.log; do sleep 0.2; done'; cat server.log
status-server: listening on 127.0.0.1:18731
$ curl -fsS -o /dev/null -w '%{http_code}' http://127.0.0.1:18731/503; echo " exit $?" curl -sS -o /dev/null -w '%{http_code}' http://127.0.0.1:18731/503; echo " exit $?"
curl: (22) The requested URL returned error: 503 503 exit 22 503 exit 0
$ ./healthcheck.sh http://127.0.0.1:18731/200
up 200
# exit status 0
$ ./healthcheck.sh http://127.0.0.1:18731/503
down 503
# exit status 1
$ ./healthcheck.sh http://127.0.0.1:18731/404
unknown 404
# exit status 2
$ ./healthcheck.sh http://127.0.0.1:18732/
curl: (7) Failed to connect to 127.0.0.1 port 18732 after 0 ms: Could not connect to server unknown 000
# exit status 2
$ kill "$(cat server.pid)" && rm server.pid

curl -f printed 503 and exited 22; without -f it printed 503 and exited 0. Against the real server the script reports up, down, unknown 404 and, where nothing listens, unknown 000, with the matching exit statuses, the same results the tests get from the fake. The last command stopped the server.

Testing signals, timeouts and cleanup

The behaviour lessons 1 to 4 added is the behaviour most shell test suites skip: does the job clean up on SIGTERM, does it return an honest status, does a whole-run timeout stop it. This job removes its scratch directory on any exit, and on SIGTERM cleans up and re-raises so its status is 143. It writes the path of its scratch directory to $MARKER, so a test can prove the directory is gone:

guarded.sh
#!/usr/bin/env bash
# A long-running job that cleans up its scratch directory on any exit and, on SIGTERM,
# cleans up and then dies of the signal so its exit status is an honest 143.
set -Eeuo pipefail
workdir=$(mktemp -d)
# shellcheck disable=SC2329 # called from the EXIT and TERM traps
cleanup() { rm -rf -- "$workdir"; }
on_term() {
cleanup
trap - TERM EXIT
kill -s TERM "$$"
}
trap cleanup EXIT
trap on_term TERM
# Record where the scratch directory is, so a test can prove it was removed.
printf '%s\n' "$workdir" >"${MARKER:-/dev/null}"
while :; do sleep 0.2; done
signals.bats
#!/usr/bin/env bats
# Tests for guarded.sh: cleanup and exit status on SIGTERM and under a timeout.
setup() {
bats_load_library bats-support
bats_load_library bats-assert
job_pid=
}
# Runs after every test, also after a failed assertion: never leave the job running.
teardown() {
if [[ -n $job_pid ]]; then kill -TERM "$job_pid" 2>/dev/null || true; fi
}
@test "guarded.sh cleans up and re-raises on SIGTERM (exit 143)" {
local marker="$BATS_TEST_TMPDIR/marker" wd st=0
# 3>&-: bats waits for every process that holds its fd 3, so a background job
# that kept it would hang the whole run if a test failed before the kill.
MARKER="$marker" "$BATS_TEST_DIRNAME/guarded.sh" 3>&- &
job_pid=$!
for _ in $(seq 1 50); do
[ -s "$marker" ] && break
sleep 0.1
done
wd=$(cat "$marker")
assert [ -d "$wd" ]
kill -TERM "$job_pid"
wait "$job_pid" || st=$?
job_pid=
assert_equal "$st" 143
assert [ ! -d "$wd" ]
}
@test "guarded.sh stopped by a whole-run timeout returns 124" {
run timeout -s TERM 1 "$BATS_TEST_DIRNAME/guarded.sh"
assert_equal "$status" 124
}
deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ bats --tap signals.bats
1..2 ok 1 guarded.sh cleans up and re-raises on SIGTERM (exit 143) ok 2 guarded.sh stopped by a whole-run timeout returns 124

The signal test starts the job in the background, waits (at most five seconds) for the marker, sends SIGTERM from the test (never by hand, so it repeats), and captures the status with wait "$job_pid" || st=$?, the form "Capstone: a backup job you can leave running" used for its SIGTERM test. The timeout test asserts 124, the status timeout returns when it had to stop a command.

Two lines protect the test run itself. Bats waits for every process that still holds its file descriptor 3, so a background job that inherited it keeps the whole suite running if an assertion fails before the kill: lesson 1's inherited-descriptor bug, inside the test harness. 3>&- closes it for the job, and teardown, which runs after every test (also a failed one), stops the job if it is still running. The lab broke the first assertion on purpose to check this: the test failed, the suite ended, and no guarded.sh was left.

A failing test, and the local gates

A test suite earns trust only when you have seen it go red. A deliberately wrong assertion shows what a failure reads like under TAP:

deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ cat > red.bats <<'EOF' setup() { bats_load_library bats-support; bats_load_library bats-assert; source "$BATS_TEST_DIRNAME/healthcheck.sh"; } @test "classify 500 is up (wrong on purpose)" { run classify 500; assert_output "up"; } EOF bats --tap red.bats; rc=$?; rm -f red.bats; exit $rc
1..1 not ok 1 classify 500 is up (wrong on purpose) # (in test file red.bats, line 2) # `@test "classify 500 is up (wrong on purpose)" { run classify 500; assert_output "up"; }' failed # # -- output differs -- # expected : up # actual : down # -- #

not ok 1 names the test, the file and line, and bats-assert prints expected against actual, so the cause is visible without rerunning.

CI must run the same tool versions you run, not whatever the runner image ships: GitHub's ubuntu-26.04 image has ShellCheck 0.11.0 today, its ubuntu-24.04 image 0.9.0, neither has shfmt, and an image update can change any of them. A version number in a download URL is not enough: release files can be replaced, so each download is checked against a SHA-256 recorded next to the version. One script does that for ShellCheck, shfmt and actionlint, and both CI and you run it:

ci/install-tools.sh
#!/usr/bin/env bash
# Install the pinned lint tools into DIR (default tools/bin), each one verified
# against its SHA-256 before it is installed. CI runs this; so can you.
# A version and its checksums change together, in this file only.
set -euo pipefail
dest=${1:-tools/bin}
sc_ver=0.11.0 fmt_ver=3.12.0 al_ver=1.7.12
case $(uname -m) in
x86_64)
sc_arch=x86_64 go_arch=amd64
sc_sha=8c3be12b05d5c177a04c29e3c78ce89ac86f1595681cab149b65b97c4e227198
fmt_sha=d9fbb2a9c33d13f47e7618cf362a914d029d02a6df124064fff04fd688a745ea
al_sha=8aca8db96f1b94770f1b0d72b6dddcb1ebb8123cb3712530b08cc387b349a3d8
;;
aarch64)
sc_arch=aarch64 go_arch=arm64
sc_sha=12b331c1d2db6b9eb13cfca64306b1b157a86eb69db83023e261eaa7e7c14588
fmt_sha=5f3fe3fa6a9f766e6a182ba79a94bef8afedafc57db0b1ad32b0f67fae971ba4
al_sha=325e971b6ba9bfa504672e29be93c24981eeb1c07576d730e9f7c8805afff0c6
;;
*)
echo "install-tools: no pinned checksums for $(uname -m)" >&2
exit 1
;;
esac
tmp=$(mktemp -d)
trap 'rm -rf -- "$tmp"' EXIT
gh=https://github.com
# fetch URL FILE SHA256: download, then refuse the file unless the checksum matches.
fetch() {
curl -fsSL --retry 3 --connect-timeout 10 --max-time 120 -o "$tmp/$2" -- "$1"
echo "$3 $tmp/$2" | sha256sum -c --quiet -
}
fetch "$gh/koalaman/shellcheck/releases/download/v$sc_ver/shellcheck-v$sc_ver.linux.$sc_arch.tar.xz" sc.tar.xz "$sc_sha"
fetch "$gh/mvdan/sh/releases/download/v$fmt_ver/shfmt_v${fmt_ver}_linux_$go_arch" shfmt "$fmt_sha"
fetch "$gh/rhysd/actionlint/releases/download/v$al_ver/actionlint_${al_ver}_linux_$go_arch.tar.gz" al.tar.gz "$al_sha"
mkdir -p -- "$dest"
tar -xJf "$tmp/sc.tar.xz" -C "$tmp"
tar -xzf "$tmp/al.tar.gz" -C "$tmp" actionlint
install -m 0755 "$tmp/shellcheck-v$sc_ver/shellcheck" "$tmp/shfmt" "$tmp/actionlint" "$dest/"
printf 'installed in %s: shellcheck %s, shfmt %s, actionlint %s\n' "$dest" "$sc_ver" "$fmt_ver" "$al_ver"
deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ ./ci/install-tools.sh bin
installed in bin: shellcheck 0.11.0, shfmt 3.12.0, actionlint 1.7.12

sha256sum -c --quiet prints nothing when a checksum matches and fails the script when one does not, so a replaced file never gets installed. shfmt reads its settings from .editorconfig (a command-line flag such as -i 2 would make it ignore the file, as "Check your script" showed), so a plain shfmt -d . formats the same on every machine:

.editorconfig
# shfmt reads its settings from here, so a plain `shfmt -d .` works the same
# on a laptop and in CI (command-line flags would make it ignore this file).
root = true
[*]
indent_style = space
indent_size = 2
switch_case_indent = true

The workflow, verified by actionlint

A CI workflow is code, and a broken one fails silently by never running the check you thought it did. The job runs on ubuntu-26.04, the same release as your machine (GitHub's ubuntu-latest still means 24.04 until it moves in November 2026), and the one third-party action is pinned to a full commit SHA with its tag in a comment, because a tag like @v7 can be repointed at new code while a SHA cannot; "Release gates" later in this course checks each pin against upstream:

.github/workflows/ci.yml
name: shell-ci
on:
push:
pull_request:
permissions:
contents: read
jobs:
lint-and-test:
runs-on: ubuntu-26.04
steps:
# Third-party actions are pinned to a commit SHA, not a moving tag: a tag can be
# repointed at new code, a SHA cannot. The comment records the human-readable tag.
- name: Check out the code
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Install the pinned lint tools (checksums verified)
run: |
bash ci/install-tools.sh "$RUNNER_TEMP/bin"
echo "$RUNNER_TEMP/bin" >> "$GITHUB_PATH"
- name: Install Bats from the runner's Ubuntu archive (not pinned)
run: |
sudo apt-get update
sudo apt-get install -y bats bats-assert bats-support
- name: Format check
run: shfmt -d .
- name: ShellCheck
run: shellcheck ./*.sh fakebin/curl ci/install-tools.sh
- name: Tests
run: bats checks.bats signals.bats
- name: Workflow lint
run: actionlint -config-file .github/actionlint.yaml .github/workflows/ci.yml

actionlint 1.7.12, the current release, predates the ubuntu-26.04 label and would reject it as unknown, so the repository declares it in actionlint's own configuration file. actionlint reads .github/actionlint.yaml by itself only inside a Git repository, which is why the gate step names it with -config-file: the command then works the same in a CI checkout and in an unpacked directory. persist-credentials: false in the workflow keeps the checkout token out of the working tree.

.github/actionlint.yaml
# actionlint 1.7.12 predates GitHub's ubuntu-26.04 runner label and reports it as unknown.
# Declaring it here (the documented place for runner labels actionlint does not know) keeps
# the label check on for every other label. Remove the entry once actionlint knows it.
self-hosted-runner:
labels:
- ubuntu-26.04

Bats comes from the runner's Ubuntu archive and is not pinned; the workflow says so in the step name, and its tests use nothing newer than that version has. Here are the four gate steps, run exactly as the workflow writes them, with the pinned tools first on PATH:

deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ export PATH="$PWD/bin:$PATH" shfmt -d . shellcheck ./*.sh fakebin/curl ci/install-tools.sh bats checks.bats signals.bats actionlint -config-file .github/actionlint.yaml .github/workflows/ci.yml echo "all four gates passed"
1..7 ok 1 classify maps codes to up/down/unknown ok 2 http_status passes the code through (fake curl as a function) ok 3 main: a healthy endpoint is up, exit 0 (fake curl on PATH) ok 4 main: a 503 is down, exit 1 ok 5 main: no answer is unknown 000, exit 2 ok 6 guarded.sh cleans up and re-raises on SIGTERM (exit 143) ok 7 guarded.sh stopped by a whole-run timeout returns 124 all four gates passed
$ printf "f() {\n echo hi\n}\n" > bad.sh ./bin/shfmt -d bad.sh; rc=$?; rm -f bad.sh; exit $rc
diff bad.sh.orig bad.sh --- bad.sh.orig +++ bad.sh @@ -1,3 +1,3 @@ f() { - echo hi + echo hi }

All four passed: shfmt found nothing to change, ShellCheck nothing to report, the seven tests passed and actionlint accepted the workflow. The second command shows the format gate failing: shfmt -d prints the diff it wants and exits 1. A ShellCheck finding fails its gate the same way (bash-ops showed that one). Now let actionlint catch two real workflow mistakes:

deploy@web01:~/scr-bash-test · Ubuntu 26.04 LTS
$ sed "s/ubuntu-26.04/ubunto-26.04/" .github/workflows/ci.yml > bad-ci.yml ./bin/actionlint -no-color -config-file .github/actionlint.yaml bad-ci.yml; rc=$?; rm -f bad-ci.yml; exit $rc
bad-ci.yml:9:14: label "ubunto-26.04" is unknown. available labels are "windows-latest", "windows-latest-8-cores", "windows-2025", "windows-2025-vs2026", "windows-2022", "windows-11-arm", "ubuntu-slim", "ubuntu-latest", "ubuntu-latest-4-cores", "ubuntu-latest-8-cores", "ubuntu-latest-16-cores", "ubuntu-24.04", "ubuntu-24.04-arm", "ubuntu-22.04", "ubuntu-22.04-arm", "macos-latest", "macos-latest-xlarge", "macos-latest-large", "macos-26-intel", "macos-26-xlarge", "macos-26-large", "macos-26", "macos-15-intel", "macos-15-xlarge", "macos-15-large", "macos-15", "macos-14-xlarge", "macos-14-large", "macos-14", "self-hosted", "x64", "arm", "arm64", "linux", "macos", "windows", "ubuntu-26.04". if it is a custom label for self-hosted runner, set list of labels in actionlint.yaml config file [runner-label] | 9 | runs-on: ubunto-26.04 | ^~~~~~~~~~~~
$ cat > inj.yml <<'EOF' name: inj on: [pull_request] permissions: { contents: read } jobs: greet: runs-on: ubuntu-26.04 steps: - run: echo "title is ${{ github.event.pull_request.title }}" EOF ./bin/actionlint -no-color -config-file .github/actionlint.yaml inj.yml; rc=$?; rm -f inj.yml; exit $rc
inj.yml:8:33: "github.event.pull_request.title" is potentially untrusted. avoid using it directly in inline scripts. instead, pass it through an environment variable. see https://docs.github.com/en/actions/reference/security/secure-use#good-practices-for-mitigating-script-injection-attacks for more details [expression] | 8 | - run: echo "title is ${{ github.event.pull_request.title }}" | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

It caught a mistyped runner label (ubunto-26.04) and listed the valid ones, ending with the ubuntu-26.04 entry from the config file, and it flagged ${{ github.event.pull_request.title }} used directly in a run: step as script injection: a pull request title is attacker-controlled text, and it would be pasted into the shell script before the shell runs.

The merge gate for a shell change
1shfmt -d .
formatting drift fails
2ShellCheck (pinned, verified)
any finding fails
3bats (TAP)
status, output, signals, cleanup
4actionlint
the workflow itself is valid
All four run locally and in CI; a red result blocks the merge.

Try this

First, prove the fake earns its keep: in a copy of healthcheck.sh, put -f back (curl -fsS ...), point a copy of checks.bats at it, and run the suite. Predict which tests fail before you look.

Second, record arguments instead of faking an answer. remote.sh is a four-line function from "Trust boundaries" that runs one command on another host:

remote.sh
#!/usr/bin/env bash
# restart_remote HOST UNIT: restart a unit on another host. UNIT comes from outside,
# so it is quoted for the remote shell with ${var@Q} (see "Trust boundaries").
restart_remote() {
ssh -o BatchMode=yes -- "$1" "sudo systemctl restart -- ${2@Q}"
}

Write fakebin/ssh so that it appends $* as one line to the file named in $RECORD and exits 0. Then write remote.bats: source remote.sh, run restart_remote web01 'app; rm -rf /tmp/x' with fakebin first on PATH and RECORD pointing into BATS_TEST_TMPDIR, and assert on the recorded line. No network is involved, and the test proves the remote command was built safely. Your fake is a shell script in the repository, so CI's format gate (shfmt -d .) checks it too: run ./bin/shfmt -d fakebin/ssh before you commit and expect no output. With the lesson's .editorconfig, shfmt wants no space between a redirection and its target (>>"$RECORD"), and the lab's reference fake written as >> "$RECORD" failed the gate.

Expected: with -f back, two tests fail. "http_status passes the code through" fails because the fake now exits 22 for a 503, and "a 503 is down" fails because the script turns that 22 into unknown 000: exactly the production bug, caught by a test. The other three still pass. The recorded line must be -o BatchMode=yes -- web01 sudo systemctl restart -- 'app; rm -rf /tmp/x', with the unit in single quotes. Both results are verified in this lab.

Takeaway

Make a script sourceable, fake its outside calls in a way that production cannot be steered by, and keep each fake honest with one check against the real tool. Test what suites usually skip (exit status, SIGTERM cleanup, timeouts), and gate every change on pinned, checksum-verified tools and an actionlint-checked workflow whose actions are pinned by SHA.

Quick check
01A library script begins with set -Eeuo pipefail at file scope and ends with main "$@" (no guard). A Bats file does source ./lib.sh in setup to test one pure function. What goes wrong, and what is the fix?
Incorrect — Sourcing runs every top-level line, so both set -e and main execute in the test shell; it is not inert.
Correct — a source guard stops main, and moving set into main keeps the test shell's options unchanged.
Incorrect — Bats can source such a file; the problem is the side effects at load time, not set itself, and you keep strict mode inside main.
Incorrect — The real problem is that main runs and errexit is enabled on load; a subshell does not add the missing source guard.
02A health check calls curl -fsS -w '%{http_code}' URL and falls back to 000 when curl fails. Its Bats suite, using a fake curl that prints FAKE_CODE and always exits 0, passes a "503 is down" test. What does the check report in production when the service returns 503?
Incorrect — The fake prints the same code but not the same exit status: real curl -f exits 22 on a 503, which the script treats as a failure.
Incorrect — A plain assignment from $(...) does carry the exit status, which is how the || code=000 fallback fires.
Incorrect — -w '%{http_code}' still prints the code with -f; the problem is the exit status that follows it.
Correct — the fake did not model -f. Drop -f (or keep the printed code), make the fake follow curl's exit rules, and check it once against a real server.
03A shell CI workflow uses uses: actions/checkout@v7 and downloads ShellCheck with curl ... shellcheck-v0.11.0.tar.xz | tar -xJ. Which change most improves it?
Correct — a tag can move, and a version in a URL does not fix the bytes; a checked SHA and a checksum make the gate reproducible.
Incorrect — The trigger does not fix tag mutability or an unverified download; pinning the action and checking the file does.
Incorrect — That makes the gate advisory and lets real findings merge, the opposite of what a gate is for.
Incorrect — The runner's version changes with the image (0.9.0 on ubuntu-24.04, 0.11.0 on ubuntu-26.04 today), so laptop and CI can drift apart; pin and verify instead.

Related