When a script has outgrown Bash

Clear criteria, a Bash wrapper with a Python core, JSON between them and a safe migration.

Advanced30 min · lesson 6 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (6 files, 3 KB): scr-bash-handoff.tar.gz. Unpack it with tar -xzf scr-bash-handoff.tar.gz, which creates scr-bash-handoff/. SHA-256: 7233566dbf6a9ea97803220de4c9208d35fbb348f6684eebd89db9e419d6ff4a

This lesson is about deciding when a shell script should stay in Bash and when its core should become Python, and how to make that move without breaking what works. You will see awk do the job it is best at, watch a crafted log line fool it, fix that bug in awk (because a parser bug is not a reason to change language), then meet a requirement the shell really handles badly. From there you build a Bash wrapper and a Python core that exchange JSON under an exit-code contract, and migrate by diffing old and new output over the same fixture. The Python is stdlib only and runs the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.

Refresher: the py-sec course owns the Python itself, "Command-line options" in bash-ops owns the Bash exit-code contract (0 ok, 1 findings, 2 usage), and "Command-line tools with argparse" in py-sec owns the Python side of it. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-handoff.

awk is the last stop in the shell

awk reads a file once, splits each line into fields, and keeps a tally across lines, so a per-source count of failed logins is a single pass with no grep | cut | sort | uniq pipeline:

summarize.awk
# Tally failed SSH logins per source address, one pass over the log.
# It finds the source by splitting on "from " and taking the next word.
/Failed password/ {
n = split($0, a, "from ")
split(a[2], b, " ")
count[b[1]]++
}
END {
for (ip in count) print count[ip], ip
}
deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ awk -f summarize.awk auth-clean.log | sort
1 198.51.100.9 3 203.0.113.44

One read of the file, one count per source address. For a quick tally over a well-formed log this is the right tool, and Python would be overkill.

A parser bug is not a reason to leave

sshd logs the username a client tried exactly as it was sent, spaces included, so part of every "Failed password" line is text the attacker wrote. "Arrays and parameter expansion" in bash-ops and "Regular expressions for logs and tool output" in py-sec showed such a username fooling a parser that searches forward from the start of the line; here the same trick meets awk. The next command adds three crafted usernames to the fixture: x from 9.9.9.9, x from 9.9.9.9 port 22 and ad"min (with a double quote). It is a shown command rather than a file to edit, because the exact characters are the point.

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ cp auth-clean.log auth.log printf '%s\n' \ '2026-09-20T01:16:44.771002+00:00 web01 sshd-session[2480]: Failed password for invalid user x from 9.9.9.9 from 203.0.113.44 port 51377 ssh2' \ '2026-09-20T01:17:03.118420+00:00 web01 sshd-session[2481]: Failed password for invalid user x from 9.9.9.9 port 22 from 203.0.113.44 port 51390 ssh2' \ '2026-09-20T01:17:30.550113+00:00 web01 sshd-session[2484]: Failed password for invalid user ad"min from 198.51.100.9 port 40210 ssh2' \ >> auth.log wc -l auth.log
8 auth.log
deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ awk -f summarize.awk auth.log | sort
2 198.51.100.9 2 9.9.9.9 3 203.0.113.44

awk credited 9.9.9.9, an address the attacker typed, twice, and undercounted the real source 203.0.113.44. This is where people reach for another language, but a forward search with Python's re.search falls for the same line, as py-sec showed. The bug is in the parsing rule, not in the language. The rule that works is to read the field from the side of the line the writer controls. sshd always ends the line with from <address> port <number> ssh2, so the source is the fourth field from the end, whatever the username contains. The same rule applies at the other end: Failed password counts only as the first words of sshd's message (fields 4 and 5, right after sshd-session[pid]:), because a username can contain those words too:

summarize-end.awk
# Tally failed SSH logins per source, reading the source from the END of the line.
# sshd ends the line with "from <address> port <number> ssh2". Everything between
# "for" and that ending, including the username, can be attacker text, so never search
# forward for "from", and take "Failed password" only as the first words of sshd's
# message (fields 4 and 5, right after "sshd-session[pid]:"), never anywhere in the line.
$3 ~ /^sshd(-session)?\[[0-9]+\]:$/ && $4 == "Failed" && $5 == "password" &&
$NF == "ssh2" && $(NF-2) == "port" && $(NF-4) == "from" {
count[$(NF-3)]++
}
END {
for (ip in count) print count[ip], ip
}
deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ awk -f summarize-end.awk auth.log | sort
2 198.51.100.9 5 203.0.113.44

Every line is now counted against the address sshd wrote: five for 203.0.113.44, two for 198.51.100.9. If a tally is all the job needs, this fixed awk is the finished tool.

The signs a script has outgrown Bash

The honest criteria are about what the job now has to do, not about line count or one bug:

Structured output. A consumer wants nested JSON, and some values are untrusted strings. Retries and deadlines. A call must be retried with backoff and stop at a time limit, with the reason recorded. Concurrency with shared results. Parallel work that must merge per-item results or retry items one by one. Tests and types. Logic with enough branches that it needs unit tests and a type checker.

Concurrency needs a word, because lesson 1 ran a parallel sweep in Bash. sweep.sh started a handful of independent checks, waited with wait -n, and cleaned up process groups on exit. That is fine in Bash: each job is an external command, and its only result is an exit status. Move when the parallel jobs must share state, retry individually, or return structured results that have to be merged. Doing that in Bash means temp files, locks and hand-written parsers of those files.

Keep it in Bash, or move the core to Python?
What does the job have to do now?
glue, a tally, a few fan-out jobs
Stay in Bash
run tools, one awk pass, wait -n over exit statuses
a parser bug
Fix the parser
anchor on the fields the writer controls, in either language
JSON with untrusted strings
Move the core to Python
a real encoder and data structures
retries, merged results, tests, types
Move the core to Python
backoff, per-item results, pytest, mypy
The wrapper can stay in Bash; the logic that grew is what moves.

Now the job gets a new requirement: the security team's log pipeline wants JSON with the count per source and the usernames each source tried. A first attempt in awk builds each record with printf:

records.awk
# One JSON record per failed login, built by hand with printf.
# UNSAFE: the username is attacker text and goes into the JSON unescaped.
/Failed password/ {
user = $0
sub(/.*Failed password for (invalid user )?/, "", user)
sub(/ from [^ ]+ port [0-9]+ ssh2$/, "", user)
printf "{\"source\": \"%s\", \"user\": \"%s\"}\n", $(NF-3), user
}
deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ awk -f records.awk auth.log | jq -Mc .
{"source":"203.0.113.44","user":"admin"} {"source":"203.0.113.44","user":"admin"} {"source":"198.51.100.9","user":"root"} {"source":"203.0.113.44","user":"test"} {"source":"203.0.113.44","user":"x from 9.9.9.9"} {"source":"203.0.113.44","user":"x from 9.9.9.9 port 22"} jq: parse error: Invalid numeric literal at line 7, column 43
# jq -M only turns off colour, so the output here is plain

jq read six records and stopped at the seventh with exit 5: the quote in ad"min ended the JSON string early, and the rest of the line is not JSON. Could awk do this? Yes, with an escape function for quotes, backslashes and control characters, or by piping each value through jq -n --arg. Either way you are writing a JSON encoder, the logic is split across awk, jq and Bash, and none of it has a unit test. That is the criterion firing, and it is where the core moves.

A wrapper in Bash, a core in Python

Moving to Python does not mean throwing the shell away. Bash is a good entry point for a scheduler or a systemd unit: it checks arguments, finds files, and turns a result into an exit status. Python is good at the logic. Give them one contract: the Python core prints JSON on stdout and returns 0, or 2 on a read error; the Bash wrapper reads that JSON and decides. The core uses the same anchoring rules as the fixed awk, expressed as a regex anchored at both ends: START must match sshd's prefix and Failed password for at the start of the line, and the full pattern must end at ssh2 and the end of the line:

summarize.py
#!/usr/bin/env python3
"""Tally failed SSH password logins per source address and emit JSON.
sshd logs the attempted username as sent, so everything between "for" and the source
can be attacker text, even " from 9.9.9.9 port 22" or "Failed password for root".
Only two parts of the line are sshd's own: the start of its message, right after the
"sshd-session[pid]: " prefix, and the end, "from <address> port <number> ssh2". The
pattern is anchored at both (^ and $), so no username can move the source or turn
another kind of failure into a counted one. Only valid IP addresses are counted.
Exit status: 0 on success, 2 on a read error.
"""
import ipaddress
import json
import re
import sys
from collections import Counter, defaultdict
from collections.abc import Iterable
# The line starts "<time> <host> sshd-session[pid]: " and the message must start right
# there: "Failed password" inside a username (a "Failed none for ..." line) is not sshd's.
START = re.compile(r"^\S+ \S+ sshd(?:-session)?\[\d+\]: Failed password for ")
# (?P<user>.*) is greedy, and the pattern must end at "ssh2" and the end of the line,
# so the match takes the LAST "from <addr> port <n>", the one sshd wrote.
FAILED = re.compile(
START.pattern + r"(?:invalid user )?(?P<user>.*)"
r" from (?P<addr>\S+) port \d+ ssh2$"
)
def summarize(lines: Iterable[str]) -> dict:
counts: Counter[str] = Counter()
users: defaultdict[str, set[str]] = defaultdict(set)
total = 0
for line in lines:
if not START.match(line):
continue
total += 1
m = FAILED.match(line)
if not m:
continue
try:
ip = str(ipaddress.ip_address(m["addr"]))
except ValueError:
continue # a hostname or malformed token is not a source we count
counts[ip] += 1
users[ip].add(m["user"])
return {
"total_failed": total,
"by_source": dict(sorted(counts.items())),
"users": {ip: sorted(names) for ip, names in sorted(users.items())},
}
def main(argv: list[str]) -> int:
try:
if len(argv) > 1:
# Iterate the open file: a multi-GB log is read line by line, never whole.
with open(argv[1], encoding="utf-8", errors="replace") as f:
report = summarize(f)
else:
report = summarize(sys.stdin)
except OSError as exc:
print(f"summarize: {exc}", file=sys.stderr)
return 2
json.dump(report, sys.stdout, indent=2)
sys.stdout.write("\n")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))
deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ python3 summarize.py auth.log
{ "total_failed": 7, "by_source": { "198.51.100.9": 2, "203.0.113.44": 5 }, "users": { "198.51.100.9": [ "ad\"min", "root" ], "203.0.113.44": [ "admin", "test", "x from 9.9.9.9", "x from 9.9.9.9 port 22" ] } }

The counts match the fixed awk, and the users lists hold the attacker's strings safely escaped: "ad\"min" is valid JSON because json.dump wrote it. ipaddress drops a hostname or malformed token, and a file is read line by line (iterating the open file) so a multi-gigabyte log never sits in memory. Iterable[str] in the signature says what summarize accepts, which is what mypy checks in the next lesson.

The start needs its anchor too. When a client tries the none authentication method, sshd logs a Failed none for ... line with the username as sent, and a username can hold the words Failed password for root from 6.6.6.6 port 1. With a pattern anchored only at the end, the lab counted such a line as a failed password from the attacker's own address, so an attacker could inflate their count at will. Anchored at both ends, neither program counts it:

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ printf '%s\n' '2026-09-20T01:18:02.204517+00:00 web01 sshd-session[2490]: Failed none for invalid user Failed password for root from 6.6.6.6 port 1 from 203.0.113.45 port 51400 ssh2' > none.log python3 summarize.py none.log | jq -Mc '{total_failed, by_source}' echo "summarize-end.awk: [$(awk -f summarize-end.awk none.log)]"
{"total_failed":0,"by_source":{}} summarize-end.awk: []

total_failed is 0 and by_source is empty; the awk printed nothing between the brackets. The line is still in the log for anyone who reads it; it just cannot pose as a password failure.

logwatch.sh
#!/usr/bin/env bash
# Bash wrapper, Python core. The wrapper handles arguments, locates the log and turns
# the core's JSON into a decision; the Python core does the parsing. They agree on one
# contract: the core prints JSON on stdout, and this script's exit status is
# 0 (no source over the threshold), 1 (findings), or 2 (usage or error).
set -euo pipefail
die() {
printf 'logwatch: %s\n' "$1" >&2
exit 2
}
[[ $# -eq 2 ]] || die "usage: logwatch.sh LOGFILE THRESHOLD"
log=$1 threshold=$2
[[ $threshold =~ ^[0-9]+$ ]] || die "THRESHOLD must be a whole number"
[[ -f $log ]] || die "no such file: $log"
command -v jq >/dev/null || die "jq is not installed"
here=$(dirname -- "${BASH_SOURCE[0]}")
report=$(python3 "$here/summarize.py" "$log") || die "summarize failed"
# Every jq result goes into a variable first, with its own status check. Written as
# printf ... "$(jq ...)" or mapfile < <(jq ...), a failing jq would be ignored and
# the run would end as "no findings" with exit 0.
total=$(jq -r '.total_failed' <<<"$report") || die "cannot read the report"
hot=$(jq -r --argjson t "$threshold" \
'.by_source | to_entries[] | select(.value >= $t) | "\(.value) \(.key)"' \
<<<"$report") || die "cannot read the report"
printf 'failed logins: %s\n' "$total"
if [[ -z $hot ]]; then
printf 'no source at or above %s\n' "$threshold"
exit 0
fi
mapfile -t lines <<<"$hot"
printf 'at or above %s:\n' "$threshold"
printf ' %s\n' "${lines[@]}"
exit 1

The jq filter reads like a pipeline: .by_source | to_entries[] turns the object into one {key, value} item per source, select(.value >= $t) keeps the items at or above the threshold (--argjson t passes the threshold in as a number), and "\(.value) \(.key)" prints each as count address.

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ ./logwatch.sh auth.log 3
failed logins: 7 at or above 3: 5 203.0.113.44
# exit status 1: a finding
$ ./logwatch.sh auth.log 99
failed logins: 7 no source at or above 99
# exit status 0
$ ./logwatch.sh
logwatch: usage: logwatch.sh LOGFILE THRESHOLD
# exit status 2
$ ./logwatch.sh nope.log 3
logwatch: no such file: nope.log
# exit status 2

logwatch.sh exits 1 when a source is at or above the threshold (a finding), 0 when none is, and 2 for a usage error or a missing file. That exit status is what cron or a CI step branches on. The core writes errors to stderr and data to stdout, so the wrapper's $(...) captures only data, and the wrapper checks the core's exit status rather than trusting that empty output means success.

The wrapper applies the same rule to jq. "Strict mode, honestly" in bash-ops and lesson 1 of this course showed two ways a status gets lost: a substitution used as an argument (printf '%s' "$(jq ...)") and a process substitution (mapfile < <(jq ...)). Written either way, a broken jq would print nothing and the run would end with "no source at or above 3" and exit 0, which the contract reads as "all clear". Each jq result is assigned to a variable first, with || die on it. Prove it with a fake jq that fails, put first on PATH for one command:

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ mkdir -p broken printf '#!/bin/sh\necho "jq: error (simulated)" >&2\nexit 5\n' > broken/jq chmod +x broken/jq PATH="$PWD/broken:$PATH" ./logwatch.sh auth.log 3
jq: error (simulated) logwatch: cannot read the report
# exit status 2
$ rm -r broken

The wrapper stopped with exit 2 and a reason, not a false all-clear.

Migrating with characterization tests

When you replace a working script, you need proof that the rewrite did not change behaviour by accident. A characterization test runs the old and new programs over the same fixture and compares the output. The old program here is the summarize.awk that has been running in cron; the jq filter turns the core's JSON into the same count address lines so diff can compare them:

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ diff <(awk -f summarize.awk auth-clean.log | sort) <(python3 summarize.py auth-clean.log | jq -r ".by_source|to_entries[]|\"\(.value) \(.key)\"" | sort) && echo "identical on the clean fixture"
identical on the clean fixture
$ diff <(awk -f summarize.awk auth.log | sort) <(python3 summarize.py auth.log | jq -r ".by_source|to_entries[]|\"\(.value) \(.key)\"" | sort)
2,3c2 < 2 9.9.9.9 < 3 203.0.113.44 --- > 5 203.0.113.44

On the clean fixture the outputs are identical, so the core preserves behaviour on well-formed input. On the crafted fixture they differ in exactly one place: the old awk counted 9.9.9.9 twice and 203.0.113.44 three times, the core counted 203.0.113.44 five times. A characterization test does not demand identical output; it shows every difference so you can decide whether each is a regression or an intended fix. This one is the parser fix, and you record it as intended. Had the awk been fixed first, the diff would be empty:

deploy@web01:~/scr-bash-handoff · Ubuntu 26.04 LTS
$ diff <(awk -f summarize-end.awk auth.log | sort) <(python3 summarize.py auth.log | jq -r ".by_source|to_entries[]|\"\(.value) \(.key)\"" | sort) && echo "identical to the fixed awk on the crafted fixture"
identical to the fixed awk on the crafted fixture

Keep the crafted lines in the fixture for good. They are the cases that forced the change, and the next edit to either program has to pass them too.

Try this

First, practise the decision. For each job below, choose "stay in Bash" or "wrapper plus Python core" and name the criterion: (a) every night, check free space on five mounts and mail one line when any is over 90 percent; (b) run the lesson 1 sweep over 200 hosts, retry each failed host twice with backoff, and write one JSON report with a status and error per host; (c) the fixed summarize-end.awk, now also asked to report a count per network.

Then build (c) in the core. Add a by_network key to summarize.py that groups sources by /24 for IPv4 and /64 for IPv6: ipaddress.ip_network(f"{addr}/24", strict=False) gives the network of an address. Run python3 summarize.py auth.log | jq -c .by_network, then feed one IPv6 line on stdin, Failed password for root from 2001:db8::44 port 22 ssh2.

Check your answers: (a) stays in Bash (glue, one line for a human); (b) moves (per-host retries and merged JSON results); (c) moves once the network grouping arrives, because address arithmetic with IPv6 is typed data that awk only sees as strings. The file should print {"198.51.100.0/24":2,"203.0.113.0/24":5} and the IPv6 line {"2001:db8::/64":1}. Both outputs are verified in this lab.

Takeaway

Fix a parser bug where it is, in whatever language it is in, by reading the fields the writer controls. Move the core to Python when the job needs structured output with untrusted strings, per-item retries or merged results, or tests and types; keep Bash as a thin wrapper with an exit-code contract, and migrate behind a characterization diff over a fixture that keeps the crafted lines.

Quick check
01A 90-line Bash script tails an application log, extracts JSON fields with grep and sed, retries a webhook three times, and has started to collect subtle bugs. A teammate says "it is under 150 lines, so keep it in Bash." What is the sounder call?
Incorrect — Strict mode does not turn grep/sed into a JSON parser or add retry logic; the work the script does is the signal, not its length.
Incorrect — A parsing bug is fixable in either language (the lesson fixed the awk by reading from the end of the line). It is not, on its own, a reason to switch.
Correct — structured data and retries are concrete signs the shell has been outgrown; a small line count does not make it the right tool for them.
Incorrect — Splitting the file leaves the JSON parsing and the retries in line-oriented tools; the hard work is still in the wrong place.
02A wrapper runs python3 core.py, then prints its verdict with printf 'hot: %s\n' "$(jq -r '.hot[]' <<<"$report")" under set -euo pipefail and exits 0 when that list is empty. What happens when jq is missing from the host?
Incorrect — errexit does not see a failure inside a command substitution used as an argument; only printf's own status counts, and it succeeds.
Incorrect — A here-string is a redirection, not a pipeline, so pipefail has nothing to report here.
Incorrect — Bash only looks up a command when it runs it; a missing tool is found at that line, not when the script starts.
Correct — the substitution's status is lost in an argument. A plain assignment carries jq's status, and an explicit check turns it into exit 2.
03You migrate the old awk report to Python and run a characterization test: diff is empty on the clean fixture but shows a difference on the lines with crafted usernames. Is the migration broken?
Incorrect — Forcing a match would bring back the awk parsing bug; a characterization test shows differences to judge, not to erase.
Correct — identical output on clean input shows behaviour is preserved, and the one difference is the bug the migration set out to fix.
Incorrect — A match on well-formed input is exactly the evidence that behaviour is preserved; differences are expected only where the rewrite intends them.
Incorrect — Both sides were normalized and sorted before diff; the difference is a real change in which source was counted.

Related