When a script has outgrown Bash
Clear criteria, a Bash wrapper with a Python core, JSON between them and a safe migration.
tar -xzf scr-bash-handoff.tar.gz, which creates scr-bash-handoff/. SHA-256: 7233566dbf6a9ea97803220de4c9208d35fbb348f6684eebd89db9e419d6ff4aThis lesson is about deciding when a shell script should stay in Bash and when its core should become Python, and how to make that move without breaking what works. You will see awk do the job it is best at, watch a crafted log line fool it, fix that bug in awk (because a parser bug is not a reason to change language), then meet a requirement the shell really handles badly. From there you build a Bash wrapper and a Python core that exchange JSON under an exit-code contract, and migrate by diffing old and new output over the same fixture. The Python is stdlib only and runs the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.
Refresher: the py-sec course owns the Python itself, "Command-line options" in bash-ops owns the Bash exit-code contract (0 ok, 1 findings, 2 usage), and "Command-line tools with argparse" in py-sec owns the Python side of it. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-handoff.
awk is the last stop in the shell
awk reads a file once, splits each line into fields, and keeps a tally across lines, so a per-source count of failed logins is a single pass with no grep | cut | sort | uniq pipeline:
# Tally failed SSH logins per source address, one pass over the log.# It finds the source by splitting on "from " and taking the next word./Failed password/ {n = split($0, a, "from ")split(a[2], b, " ")count[b[1]]++}END {for (ip in count) print count[ip], ip}
One read of the file, one count per source address. For a quick tally over a well-formed log this is the right tool, and Python would be overkill.
A parser bug is not a reason to leave
sshd logs the username a client tried exactly as it was sent, spaces included, so part of every "Failed password" line is text the attacker wrote. "Arrays and parameter expansion" in bash-ops and "Regular expressions for logs and tool output" in py-sec showed such a username fooling a parser that searches forward from the start of the line; here the same trick meets awk. The next command adds three crafted usernames to the fixture: x from 9.9.9.9, x from 9.9.9.9 port 22 and ad"min (with a double quote). It is a shown command rather than a file to edit, because the exact characters are the point.
awk credited 9.9.9.9, an address the attacker typed, twice, and undercounted the real source 203.0.113.44. This is where people reach for another language, but a forward search with Python's re.search falls for the same line, as py-sec showed. The bug is in the parsing rule, not in the language. The rule that works is to read the field from the side of the line the writer controls. sshd always ends the line with from <address> port <number> ssh2, so the source is the fourth field from the end, whatever the username contains. The same rule applies at the other end: Failed password counts only as the first words of sshd's message (fields 4 and 5, right after sshd-session[pid]:), because a username can contain those words too:
# Tally failed SSH logins per source, reading the source from the END of the line.# sshd ends the line with "from <address> port <number> ssh2". Everything between# "for" and that ending, including the username, can be attacker text, so never search# forward for "from", and take "Failed password" only as the first words of sshd's# message (fields 4 and 5, right after "sshd-session[pid]:"), never anywhere in the line.$3 ~ /^sshd(-session)?\[[0-9]+\]:$/ && $4 == "Failed" && $5 == "password" &&$NF == "ssh2" && $(NF-2) == "port" && $(NF-4) == "from" {count[$(NF-3)]++}END {for (ip in count) print count[ip], ip}
Every line is now counted against the address sshd wrote: five for 203.0.113.44, two for 198.51.100.9. If a tally is all the job needs, this fixed awk is the finished tool.
The signs a script has outgrown Bash
The honest criteria are about what the job now has to do, not about line count or one bug:
Structured output. A consumer wants nested JSON, and some values are untrusted strings. Retries and deadlines. A call must be retried with backoff and stop at a time limit, with the reason recorded. Concurrency with shared results. Parallel work that must merge per-item results or retry items one by one. Tests and types. Logic with enough branches that it needs unit tests and a type checker.
Concurrency needs a word, because lesson 1 ran a parallel sweep in Bash. sweep.sh started a handful of independent checks, waited with wait -n, and cleaned up process groups on exit. That is fine in Bash: each job is an external command, and its only result is an exit status. Move when the parallel jobs must share state, retry individually, or return structured results that have to be merged. Doing that in Bash means temp files, locks and hand-written parsers of those files.
Now the job gets a new requirement: the security team's log pipeline wants JSON with the count per source and the usernames each source tried. A first attempt in awk builds each record with printf:
# One JSON record per failed login, built by hand with printf.# UNSAFE: the username is attacker text and goes into the JSON unescaped./Failed password/ {user = $0sub(/.*Failed password for (invalid user )?/, "", user)sub(/ from [^ ]+ port [0-9]+ ssh2$/, "", user)printf "{\"source\": \"%s\", \"user\": \"%s\"}\n", $(NF-3), user}
jq read six records and stopped at the seventh with exit 5: the quote in ad"min ended the JSON string early, and the rest of the line is not JSON. Could awk do this? Yes, with an escape function for quotes, backslashes and control characters, or by piping each value through jq -n --arg. Either way you are writing a JSON encoder, the logic is split across awk, jq and Bash, and none of it has a unit test. That is the criterion firing, and it is where the core moves.
A wrapper in Bash, a core in Python
Moving to Python does not mean throwing the shell away. Bash is a good entry point for a scheduler or a systemd unit: it checks arguments, finds files, and turns a result into an exit status. Python is good at the logic. Give them one contract: the Python core prints JSON on stdout and returns 0, or 2 on a read error; the Bash wrapper reads that JSON and decides. The core uses the same anchoring rules as the fixed awk, expressed as a regex anchored at both ends: START must match sshd's prefix and Failed password for at the start of the line, and the full pattern must end at ssh2 and the end of the line:
#!/usr/bin/env python3"""Tally failed SSH password logins per source address and emit JSON.sshd logs the attempted username as sent, so everything between "for" and the sourcecan be attacker text, even " from 9.9.9.9 port 22" or "Failed password for root".Only two parts of the line are sshd's own: the start of its message, right after the"sshd-session[pid]: " prefix, and the end, "from <address> port <number> ssh2". Thepattern is anchored at both (^ and $), so no username can move the source or turnanother kind of failure into a counted one. Only valid IP addresses are counted.Exit status: 0 on success, 2 on a read error."""import ipaddressimport jsonimport reimport sysfrom collections import Counter, defaultdictfrom collections.abc import Iterable# The line starts "<time> <host> sshd-session[pid]: " and the message must start right# there: "Failed password" inside a username (a "Failed none for ..." line) is not sshd's.START = re.compile(r"^\S+ \S+ sshd(?:-session)?\[\d+\]: Failed password for ")# (?P<user>.*) is greedy, and the pattern must end at "ssh2" and the end of the line,# so the match takes the LAST "from <addr> port <n>", the one sshd wrote.FAILED = re.compile(START.pattern + r"(?:invalid user )?(?P<user>.*)"r" from (?P<addr>\S+) port \d+ ssh2$")def summarize(lines: Iterable[str]) -> dict:counts: Counter[str] = Counter()users: defaultdict[str, set[str]] = defaultdict(set)total = 0for line in lines:if not START.match(line):continuetotal += 1m = FAILED.match(line)if not m:continuetry:ip = str(ipaddress.ip_address(m["addr"]))except ValueError:continue # a hostname or malformed token is not a source we countcounts[ip] += 1users[ip].add(m["user"])return {"total_failed": total,"by_source": dict(sorted(counts.items())),"users": {ip: sorted(names) for ip, names in sorted(users.items())},}def main(argv: list[str]) -> int:try:if len(argv) > 1:# Iterate the open file: a multi-GB log is read line by line, never whole.with open(argv[1], encoding="utf-8", errors="replace") as f:report = summarize(f)else:report = summarize(sys.stdin)except OSError as exc:print(f"summarize: {exc}", file=sys.stderr)return 2json.dump(report, sys.stdout, indent=2)sys.stdout.write("\n")return 0if __name__ == "__main__":sys.exit(main(sys.argv))
The counts match the fixed awk, and the users lists hold the attacker's strings safely escaped: "ad\"min" is valid JSON because json.dump wrote it. ipaddress drops a hostname or malformed token, and a file is read line by line (iterating the open file) so a multi-gigabyte log never sits in memory. Iterable[str] in the signature says what summarize accepts, which is what mypy checks in the next lesson.
The start needs its anchor too. When a client tries the none authentication method, sshd logs a Failed none for ... line with the username as sent, and a username can hold the words Failed password for root from 6.6.6.6 port 1. With a pattern anchored only at the end, the lab counted such a line as a failed password from the attacker's own address, so an attacker could inflate their count at will. Anchored at both ends, neither program counts it:
total_failed is 0 and by_source is empty; the awk printed nothing between the brackets. The line is still in the log for anyone who reads it; it just cannot pose as a password failure.
#!/usr/bin/env bash# Bash wrapper, Python core. The wrapper handles arguments, locates the log and turns# the core's JSON into a decision; the Python core does the parsing. They agree on one# contract: the core prints JSON on stdout, and this script's exit status is# 0 (no source over the threshold), 1 (findings), or 2 (usage or error).set -euo pipefaildie() {printf 'logwatch: %s\n' "$1" >&2exit 2}[[ $# -eq 2 ]] || die "usage: logwatch.sh LOGFILE THRESHOLD"log=$1 threshold=$2[[ $threshold =~ ^[0-9]+$ ]] || die "THRESHOLD must be a whole number"[[ -f $log ]] || die "no such file: $log"command -v jq >/dev/null || die "jq is not installed"here=$(dirname -- "${BASH_SOURCE[0]}")report=$(python3 "$here/summarize.py" "$log") || die "summarize failed"# Every jq result goes into a variable first, with its own status check. Written as# printf ... "$(jq ...)" or mapfile < <(jq ...), a failing jq would be ignored and# the run would end as "no findings" with exit 0.total=$(jq -r '.total_failed' <<<"$report") || die "cannot read the report"hot=$(jq -r --argjson t "$threshold" \'.by_source | to_entries[] | select(.value >= $t) | "\(.value) \(.key)"' \<<<"$report") || die "cannot read the report"printf 'failed logins: %s\n' "$total"if [[ -z $hot ]]; thenprintf 'no source at or above %s\n' "$threshold"exit 0fimapfile -t lines <<<"$hot"printf 'at or above %s:\n' "$threshold"printf ' %s\n' "${lines[@]}"exit 1
The jq filter reads like a pipeline: .by_source | to_entries[] turns the object into one {key, value} item per source, select(.value >= $t) keeps the items at or above the threshold (--argjson t passes the threshold in as a number), and "\(.value) \(.key)" prints each as count address.
logwatch.sh exits 1 when a source is at or above the threshold (a finding), 0 when none is, and 2 for a usage error or a missing file. That exit status is what cron or a CI step branches on. The core writes errors to stderr and data to stdout, so the wrapper's $(...) captures only data, and the wrapper checks the core's exit status rather than trusting that empty output means success.
The wrapper applies the same rule to jq. "Strict mode, honestly" in bash-ops and lesson 1 of this course showed two ways a status gets lost: a substitution used as an argument (printf '%s' "$(jq ...)") and a process substitution (mapfile < <(jq ...)). Written either way, a broken jq would print nothing and the run would end with "no source at or above 3" and exit 0, which the contract reads as "all clear". Each jq result is assigned to a variable first, with || die on it. Prove it with a fake jq that fails, put first on PATH for one command:
The wrapper stopped with exit 2 and a reason, not a false all-clear.
Migrating with characterization tests
When you replace a working script, you need proof that the rewrite did not change behaviour by accident. A characterization test runs the old and new programs over the same fixture and compares the output. The old program here is the summarize.awk that has been running in cron; the jq filter turns the core's JSON into the same count address lines so diff can compare them:
On the clean fixture the outputs are identical, so the core preserves behaviour on well-formed input. On the crafted fixture they differ in exactly one place: the old awk counted 9.9.9.9 twice and 203.0.113.44 three times, the core counted 203.0.113.44 five times. A characterization test does not demand identical output; it shows every difference so you can decide whether each is a regression or an intended fix. This one is the parser fix, and you record it as intended. Had the awk been fixed first, the diff would be empty:
Keep the crafted lines in the fixture for good. They are the cases that forced the change, and the next edit to either program has to pass them too.
Try this
First, practise the decision. For each job below, choose "stay in Bash" or "wrapper plus Python core" and name the criterion: (a) every night, check free space on five mounts and mail one line when any is over 90 percent; (b) run the lesson 1 sweep over 200 hosts, retry each failed host twice with backoff, and write one JSON report with a status and error per host; (c) the fixed summarize-end.awk, now also asked to report a count per network.
Then build (c) in the core. Add a by_network key to summarize.py that groups sources by /24 for IPv4 and /64 for IPv6: ipaddress.ip_network(f"{addr}/24", strict=False) gives the network of an address. Run python3 summarize.py auth.log | jq -c .by_network, then feed one IPv6 line on stdin, Failed password for root from 2001:db8::44 port 22 ssh2.
Check your answers: (a) stays in Bash (glue, one line for a human); (b) moves (per-host retries and merged JSON results); (c) moves once the network grouping arrives, because address arithmetic with IPv6 is typed data that awk only sees as strings. The file should print {"198.51.100.0/24":2,"203.0.113.0/24":5} and the IPv6 line {"2001:db8::/64":1}. Both outputs are verified in this lab.
Takeaway
Fix a parser bug where it is, in whatever language it is in, by reading the fields the writer controls. Move the core to Python when the job needs structured output with untrusted strings, per-item retries or merged results, or tests and types; keep Bash as a thin wrapper with an exit-code contract, and migrate behind a characterization diff over a fixture that keeps the crafted lines.
grep and sed, retries a webhook three times, and has started to collect subtle bugs. A teammate says "it is under 150 lines, so keep it in Bash." What is the sounder call?grep/sed into a JSON parser or add retry logic; the work the script does is the signal, not its length.python3 core.py, then prints its verdict with printf 'hot: %s\n' "$(jq -r '.hot[]' <<<"$report")" under set -euo pipefail and exits 0 when that list is empty. What happens when jq is missing from the host?printf's own status counts, and it succeeds.jq's status, and an explicit check turns it into exit 2.diff is empty on the clean fixture but shows a difference on the lines with crafted usernames. Is the migration broken?diff; the difference is a real change in which source was counted.