Testing shell automation
Bats with fakes, testing signals and timeouts, and ShellCheck and Bats as CI gates.
tar -xzf scr-bash-test.tar.gz, which creates scr-bash-test/. SHA-256: 82a57d667f50565882c69dc520cdaa9bb7c2da05b36422276deed888b22f3690This lesson tests the shell automation you have written in the rest of this course: the failure paths, the signal handling, the timeouts and the cleanup. You will make a script testable so a test can load it without running it, write assertions with bats-assert, fake the command that touches the network (and check the fake against the real thing), prove that a job cleans up and returns 143 on SIGTERM, read a failing test, and gate a change in CI with pinned, checksum-verified tools and a workflow that actionlint has checked.
Refresher: "Check your script" in bash-ops owns ShellCheck (severities, directives, exit status 1), shfmt with .editorconfig, and your first Bats file (run, $status, $output, BATS_TEST_TMPDIR). This lesson is only the delta. It uses the archive's bats 1.13 with bats-assert and bats-support (sudo apt install bats bats-assert bats-support if bash-ops did not install them already). Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-test; they include status-server.py, a tiny local HTTP server that answers each request with the status code named in its path, used once below.
Design for the test first
A test cannot exercise one function of a script if loading the file runs the whole thing, and it cannot load the file at all if loading it turns on options or connects to a service. So put the runtime work, and the set -Eeuo pipefail that goes with it, inside main, and guard main so it runs only on direct execution:
#!/usr/bin/env bash# Check one HTTP endpoint and classify it. Written to be testable: the logic is in# small functions, and the file does its work (and turns on strict mode) only when# executed directly, so a test can source it without running main or changing the# test shell's options. Exit: 0 up, 1 down, 2 unknown or unreachable.http_status() { # http_status URL -> the HTTP status code, or 000 if nothing answered# No -f: with --fail, curl exits 22 on any 4xx/5xx, and a 503 would look like# "could not connect". Without it, curl prints the code and exits 0.curl -sS -o /dev/null -w '%{http_code}' --max-time 5 -- "$1"}classify() { # classify CODE -> up | down | unknownlocal code=$1if ((code >= 200 && code < 300)); thenecho upelif ((code >= 500)); thenecho downelseecho unknownfi}main() {set -Eeuo pipefaillocal url=${1:?usage: healthcheck.sh URL} code statecode=$(http_status "$url") || code=000state=$(classify "$code")printf '%s %s\n' "$state" "$code"case $state inup) return 0 ;;down) return 1 ;;*) return 2 ;;esac}# BASH_SOURCE[0] is this file; $0 is what the shell is running. They match on direct# execution and differ when a test sources the file, so main runs only for real.if [[ ${BASH_SOURCE[0]} == "${0}" ]]; thenmain "$@"fi
The guard [[ ${BASH_SOURCE[0]} == "${0}" ]] compares the file the code lives in with the file the shell was told to run: they match on ./healthcheck.sh and differ when a test sources it, so main runs only for real. classify is a pure function with no I/O. The exit status carries the result (0 up, 1 down, 2 unknown or unreachable), so cron or a CI step can act on it. The comment in http_status about -f is there because an earlier version of this script had it, and the tests passed anyway; the contract check below shows why.
Assertions and fakes with bats-assert
bats-assert turns raw [ "$status" -eq 0 ] checks into readable assertions with useful failure messages: assert_success, assert_failure 1, assert_output, assert_equal. bats_load_library finds it in /usr/lib/bats, where bats 1.13 looks by default. The command that reaches the network, curl, is faked two ways:
#!/usr/bin/env bats# Tests for healthcheck.sh. Run with: bats checks.batssetup() {bats_load_library bats-supportbats_load_library bats-assert# Sourcing is safe: healthcheck.sh turns on strict mode only inside main().source "$BATS_TEST_DIRNAME/healthcheck.sh"}@test "classify maps codes to up/down/unknown" {run classify 204assert_output "up"run classify 503assert_output "down"run classify 404assert_output "unknown"}@test "http_status passes the code through (fake curl as a function)" {# A function beats a command on PATH, and exists only in this test shell.curl() { "$BATS_TEST_DIRNAME/fakebin/curl" "$@"; }FAKE_CODE=503 run http_status http://svc.invalid/assert_successassert_output "503"}@test "main: a healthy endpoint is up, exit 0 (fake curl on PATH)" {PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_CODE=200 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/assert_successassert_output "up 200"}@test "main: a 503 is down, exit 1" {PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_CODE=503 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/assert_failure 1assert_output "down 503"}@test "main: no answer is unknown 000, exit 2" {PATH="$BATS_TEST_DIRNAME/fakebin:$PATH" FAKE_FAIL=1 run "$BATS_TEST_DIRNAME/healthcheck.sh" http://svc.invalid/assert_failure 2assert_output "unknown 000"}
#!/usr/bin/env bash# Test double for curl. It ignores the request and "answers" with FAKE_CODE, and it# keeps real curl's exit rules so a test cannot pass on behaviour curl does not have:# FAKE_FAIL=1 is a connection failure (prints 000, exit 7); with -f or --fail an# answer of 400 or more exits 22 after printing the code.fail=0for arg in "$@"; docase $arg in--fail) fail=1 ;;--*) ;;-*f*) fail=1 ;; # -f alone or in a group such as -fsSesacdoneif [[ ${FAKE_FAIL:-0} == 1 ]]; thenprintf '000'exit 7ficode=${FAKE_CODE:-200}printf '%s' "$code"if ((fail && code >= 400)); then exit 22; fi
The second test defines a curl function after sourcing: a function wins over a command on PATH, and it exists only in the test shell, so production has no hook for it. The last three run healthcheck.sh as a separate process, which cannot see the test's functions, so they put fakebin first on PATH. A common alternative, CURL=${CURL:-curl} in the script, works too, but then whoever controls the job's environment (a CI variable, an EnvironmentFile=) picks the binary, which is the PATH problem from "Trust boundaries". Run the suite the way CI does, without a terminal, so bats prints TAP:
1..5 then one ok line per test is the TAP output a CI log shows; the ✓ form appears only on an interactive terminal.
A fake is only as good as its likeness to the real command. Real curl with -f (--fail) exits 22 on any status of 400 or more, after printing the code. An earlier fake printed the code and exited 0 whatever the flags, so a healthcheck.sh that used -fsS passed its "down 503" test while in production every 503 became unknown 000. That is why this fake implements the -f rule. Check a fake against the real thing once, with a local server that answers with the status in the path. Start it in the background from ~/scr-bash-test:
curl -f printed 503 and exited 22; without -f it printed 503 and exited 0. Against the real server the script reports up, down, unknown 404 and, where nothing listens, unknown 000, with the matching exit statuses, the same results the tests get from the fake. The last command stopped the server.
Testing signals, timeouts and cleanup
The behaviour lessons 1 to 4 added is the behaviour most shell test suites skip: does the job clean up on SIGTERM, does it return an honest status, does a whole-run timeout stop it. This job removes its scratch directory on any exit, and on SIGTERM cleans up and re-raises so its status is 143. It writes the path of its scratch directory to $MARKER, so a test can prove the directory is gone:
#!/usr/bin/env bash# A long-running job that cleans up its scratch directory on any exit and, on SIGTERM,# cleans up and then dies of the signal so its exit status is an honest 143.set -Eeuo pipefailworkdir=$(mktemp -d)# shellcheck disable=SC2329 # called from the EXIT and TERM trapscleanup() { rm -rf -- "$workdir"; }on_term() {cleanuptrap - TERM EXITkill -s TERM "$$"}trap cleanup EXITtrap on_term TERM# Record where the scratch directory is, so a test can prove it was removed.printf '%s\n' "$workdir" >"${MARKER:-/dev/null}"while :; do sleep 0.2; done
#!/usr/bin/env bats# Tests for guarded.sh: cleanup and exit status on SIGTERM and under a timeout.setup() {bats_load_library bats-supportbats_load_library bats-assertjob_pid=}# Runs after every test, also after a failed assertion: never leave the job running.teardown() {if [[ -n $job_pid ]]; then kill -TERM "$job_pid" 2>/dev/null || true; fi}@test "guarded.sh cleans up and re-raises on SIGTERM (exit 143)" {local marker="$BATS_TEST_TMPDIR/marker" wd st=0# 3>&-: bats waits for every process that holds its fd 3, so a background job# that kept it would hang the whole run if a test failed before the kill.MARKER="$marker" "$BATS_TEST_DIRNAME/guarded.sh" 3>&- &job_pid=$!for _ in $(seq 1 50); do[ -s "$marker" ] && breaksleep 0.1donewd=$(cat "$marker")assert [ -d "$wd" ]kill -TERM "$job_pid"wait "$job_pid" || st=$?job_pid=assert_equal "$st" 143assert [ ! -d "$wd" ]}@test "guarded.sh stopped by a whole-run timeout returns 124" {run timeout -s TERM 1 "$BATS_TEST_DIRNAME/guarded.sh"assert_equal "$status" 124}
The signal test starts the job in the background, waits (at most five seconds) for the marker, sends SIGTERM from the test (never by hand, so it repeats), and captures the status with wait "$job_pid" || st=$?, the form "Capstone: a backup job you can leave running" used for its SIGTERM test. The timeout test asserts 124, the status timeout returns when it had to stop a command.
Two lines protect the test run itself. Bats waits for every process that still holds its file descriptor 3, so a background job that inherited it keeps the whole suite running if an assertion fails before the kill: lesson 1's inherited-descriptor bug, inside the test harness. 3>&- closes it for the job, and teardown, which runs after every test (also a failed one), stops the job if it is still running. The lab broke the first assertion on purpose to check this: the test failed, the suite ended, and no guarded.sh was left.
A failing test, and the local gates
A test suite earns trust only when you have seen it go red. A deliberately wrong assertion shows what a failure reads like under TAP:
not ok 1 names the test, the file and line, and bats-assert prints expected against actual, so the cause is visible without rerunning.
CI must run the same tool versions you run, not whatever the runner image ships: GitHub's ubuntu-26.04 image has ShellCheck 0.11.0 today, its ubuntu-24.04 image 0.9.0, neither has shfmt, and an image update can change any of them. A version number in a download URL is not enough: release files can be replaced, so each download is checked against a SHA-256 recorded next to the version. One script does that for ShellCheck, shfmt and actionlint, and both CI and you run it:
#!/usr/bin/env bash# Install the pinned lint tools into DIR (default tools/bin), each one verified# against its SHA-256 before it is installed. CI runs this; so can you.# A version and its checksums change together, in this file only.set -euo pipefaildest=${1:-tools/bin}sc_ver=0.11.0 fmt_ver=3.12.0 al_ver=1.7.12case $(uname -m) inx86_64)sc_arch=x86_64 go_arch=amd64sc_sha=8c3be12b05d5c177a04c29e3c78ce89ac86f1595681cab149b65b97c4e227198fmt_sha=d9fbb2a9c33d13f47e7618cf362a914d029d02a6df124064fff04fd688a745eaal_sha=8aca8db96f1b94770f1b0d72b6dddcb1ebb8123cb3712530b08cc387b349a3d8;;aarch64)sc_arch=aarch64 go_arch=arm64sc_sha=12b331c1d2db6b9eb13cfca64306b1b157a86eb69db83023e261eaa7e7c14588fmt_sha=5f3fe3fa6a9f766e6a182ba79a94bef8afedafc57db0b1ad32b0f67fae971ba4al_sha=325e971b6ba9bfa504672e29be93c24981eeb1c07576d730e9f7c8805afff0c6;;*)echo "install-tools: no pinned checksums for $(uname -m)" >&2exit 1;;esactmp=$(mktemp -d)trap 'rm -rf -- "$tmp"' EXITgh=https://github.com# fetch URL FILE SHA256: download, then refuse the file unless the checksum matches.fetch() {curl -fsSL --retry 3 --connect-timeout 10 --max-time 120 -o "$tmp/$2" -- "$1"echo "$3 $tmp/$2" | sha256sum -c --quiet -}fetch "$gh/koalaman/shellcheck/releases/download/v$sc_ver/shellcheck-v$sc_ver.linux.$sc_arch.tar.xz" sc.tar.xz "$sc_sha"fetch "$gh/mvdan/sh/releases/download/v$fmt_ver/shfmt_v${fmt_ver}_linux_$go_arch" shfmt "$fmt_sha"fetch "$gh/rhysd/actionlint/releases/download/v$al_ver/actionlint_${al_ver}_linux_$go_arch.tar.gz" al.tar.gz "$al_sha"mkdir -p -- "$dest"tar -xJf "$tmp/sc.tar.xz" -C "$tmp"tar -xzf "$tmp/al.tar.gz" -C "$tmp" actionlintinstall -m 0755 "$tmp/shellcheck-v$sc_ver/shellcheck" "$tmp/shfmt" "$tmp/actionlint" "$dest/"printf 'installed in %s: shellcheck %s, shfmt %s, actionlint %s\n' "$dest" "$sc_ver" "$fmt_ver" "$al_ver"
sha256sum -c --quiet prints nothing when a checksum matches and fails the script when one does not, so a replaced file never gets installed. shfmt reads its settings from .editorconfig (a command-line flag such as -i 2 would make it ignore the file, as "Check your script" showed), so a plain shfmt -d . formats the same on every machine:
# shfmt reads its settings from here, so a plain `shfmt -d .` works the same# on a laptop and in CI (command-line flags would make it ignore this file).root = true[*]indent_style = spaceindent_size = 2switch_case_indent = true
The workflow, verified by actionlint
A CI workflow is code, and a broken one fails silently by never running the check you thought it did. The job runs on ubuntu-26.04, the same release as your machine (GitHub's ubuntu-latest still means 24.04 until it moves in November 2026), and the one third-party action is pinned to a full commit SHA with its tag in a comment, because a tag like @v7 can be repointed at new code while a SHA cannot; "Release gates" later in this course checks each pin against upstream:
name: shell-cion:push:pull_request:permissions:contents: readjobs:lint-and-test:runs-on: ubuntu-26.04steps:# Third-party actions are pinned to a commit SHA, not a moving tag: a tag can be# repointed at new code, a SHA cannot. The comment records the human-readable tag.- name: Check out the codeuses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1with:persist-credentials: false- name: Install the pinned lint tools (checksums verified)run: |bash ci/install-tools.sh "$RUNNER_TEMP/bin"echo "$RUNNER_TEMP/bin" >> "$GITHUB_PATH"- name: Install Bats from the runner's Ubuntu archive (not pinned)run: |sudo apt-get updatesudo apt-get install -y bats bats-assert bats-support- name: Format checkrun: shfmt -d .- name: ShellCheckrun: shellcheck ./*.sh fakebin/curl ci/install-tools.sh- name: Testsrun: bats checks.bats signals.bats- name: Workflow lintrun: actionlint -config-file .github/actionlint.yaml .github/workflows/ci.yml
actionlint 1.7.12, the current release, predates the ubuntu-26.04 label and would reject it as unknown, so the repository declares it in actionlint's own configuration file. actionlint reads .github/actionlint.yaml by itself only inside a Git repository, which is why the gate step names it with -config-file: the command then works the same in a CI checkout and in an unpacked directory. persist-credentials: false in the workflow keeps the checkout token out of the working tree.
# actionlint 1.7.12 predates GitHub's ubuntu-26.04 runner label and reports it as unknown.# Declaring it here (the documented place for runner labels actionlint does not know) keeps# the label check on for every other label. Remove the entry once actionlint knows it.self-hosted-runner:labels:- ubuntu-26.04
Bats comes from the runner's Ubuntu archive and is not pinned; the workflow says so in the step name, and its tests use nothing newer than that version has. Here are the four gate steps, run exactly as the workflow writes them, with the pinned tools first on PATH:
All four passed: shfmt found nothing to change, ShellCheck nothing to report, the seven tests passed and actionlint accepted the workflow. The second command shows the format gate failing: shfmt -d prints the diff it wants and exits 1. A ShellCheck finding fails its gate the same way (bash-ops showed that one). Now let actionlint catch two real workflow mistakes:
It caught a mistyped runner label (ubunto-26.04) and listed the valid ones, ending with the ubuntu-26.04 entry from the config file, and it flagged ${{ github.event.pull_request.title }} used directly in a run: step as script injection: a pull request title is attacker-controlled text, and it would be pasted into the shell script before the shell runs.
Try this
First, prove the fake earns its keep: in a copy of healthcheck.sh, put -f back (curl -fsS ...), point a copy of checks.bats at it, and run the suite. Predict which tests fail before you look.
Second, record arguments instead of faking an answer. remote.sh is a four-line function from "Trust boundaries" that runs one command on another host:
#!/usr/bin/env bash# restart_remote HOST UNIT: restart a unit on another host. UNIT comes from outside,# so it is quoted for the remote shell with ${var@Q} (see "Trust boundaries").restart_remote() {ssh -o BatchMode=yes -- "$1" "sudo systemctl restart -- ${2@Q}"}
Write fakebin/ssh so that it appends $* as one line to the file named in $RECORD and exits 0. Then write remote.bats: source remote.sh, run restart_remote web01 'app; rm -rf /tmp/x' with fakebin first on PATH and RECORD pointing into BATS_TEST_TMPDIR, and assert on the recorded line. No network is involved, and the test proves the remote command was built safely. Your fake is a shell script in the repository, so CI's format gate (shfmt -d .) checks it too: run ./bin/shfmt -d fakebin/ssh before you commit and expect no output. With the lesson's .editorconfig, shfmt wants no space between a redirection and its target (>>"$RECORD"), and the lab's reference fake written as >> "$RECORD" failed the gate.
Expected: with -f back, two tests fail. "http_status passes the code through" fails because the fake now exits 22 for a 503, and "a 503 is down" fails because the script turns that 22 into unknown 000: exactly the production bug, caught by a test. The other three still pass. The recorded line must be -o BatchMode=yes -- web01 sudo systemctl restart -- 'app; rm -rf /tmp/x', with the unit in single quotes. Both results are verified in this lab.
Takeaway
Make a script sourceable, fake its outside calls in a way that production cannot be steered by, and keep each fake honest with one check against the real tool. Test what suites usually skip (exit status, SIGTERM cleanup, timeouts), and gate every change on pinned, checksum-verified tools and an actionlint-checked workflow whose actions are pinned by SHA.
set -Eeuo pipefail at file scope and ends with main "$@" (no guard). A Bats file does source ./lib.sh in setup to test one pure function. What goes wrong, and what is the fix?set -e and main execute in the test shell; it is not inert.main, and moving set into main keeps the test shell's options unchanged.set itself, and you keep strict mode inside main.main runs and errexit is enabled on load; a subshell does not add the missing source guard.curl -fsS -w '%{http_code}' URL and falls back to 000 when curl fails. Its Bats suite, using a fake curl that prints FAKE_CODE and always exits 0, passes a "503 is down" test. What does the check report in production when the service returns 503?curl -f exits 22 on a 503, which the script treats as a failure.$(...) does carry the exit status, which is how the || code=000 fallback fires.-w '%{http_code}' still prints the code with -f; the problem is the exit status that follows it.-f. Drop -f (or keep the printed code), make the fake follow curl's exit rules, and check it once against a real server.uses: actions/checkout@v7 and downloads ShellCheck with curl ... shellcheck-v0.11.0.tar.xz | tar -xJ. Which change most improves it?ubuntu-24.04, 0.11.0 on ubuntu-26.04 today), so laptop and CI can drift apart; pin and verify instead.