Failure design: exceptions, deadlines, retries and cleanup

Exception taxonomy, deadline budgets, idempotent retries, SIGTERM and ExitStack.

Advanced40 min · lesson 8 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (18 files, 7 KB): scr-py-resilience.tar.gz. Unpack it with tar -xzf scr-py-resilience.tar.gz, which creates scr-py-resilience/. SHA-256: b6a2296242282fb86f152cf2264da477429562acff492438d299b1bd8a163cca

This lesson designs the failure behaviour of a whole tool, not of one call. You will give a tool an exception taxonomy that every layer shares, report many failures at once with ExceptionGroup and except*, spend one time budget across every call, wait and slow answer of a run, honour Retry-After in both of its forms without letting a server park your job for an hour, make a retried POST safe with an idempotency key, and stop cleanly on SIGTERM with contextlib.ExitStack and an fcntl lock. The tool is ticketer, which opens one ticket per security finding against a local API that fails on purpose. Stdlib only; it behaves the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.

Refresher: "Logging, retries and failure design" in py-sec owns temporary versus permanent errors, raise ... from, logging to stderr, and one retry loop with full jitter and a deadline; "Calling HTTP APIs safely" owns timeouts, 429, a capped Retry-After, and read_body(), which reads an answer within a time budget. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-py-resilience. Besides the modules shown here, they hold the findings*.json batches and ticketapi.py, the test API described below.

A taxonomy the whole tool agrees on

Each kind of failure needs a different response from whoever runs the tool, so ticketer sorts every failure into one of these classes, and the class carries the exit status:

errors.py
"""What can go wrong in ticketer, sorted by what the caller should do about it."""
class ToolError(Exception):
exit_status = 1
class PermanentError(ToolError):
"""Retrying cannot help: the request or the data is wrong. Exit 1."""
class TransientError(ToolError):
"""May work later. retry_after: the wait the server asked for, in seconds, if any."""
exit_status = 75 # EX_TEMPFAIL: the scheduler should run the job again later
def __init__(self, message: str, retry_after: float | None = None) -> None:
super().__init__(message)
self.retry_after = retry_after
class DeadlineExceeded(TransientError):
"""The run's time budget is spent."""
class Terminated(BaseException):
"""SIGTERM arrived. A BaseException, like KeyboardInterrupt, so "except Exception" in
library code cannot swallow a request to stop."""

PermanentError means a human must fix something, exit status 1. TransientError means "run me again later", exit status 75 (EX_TEMPFAIL from sysexits.h), and can carry the wait the server asked for. DeadlineExceeded is a kind of transient error. Terminated derives from BaseException, like KeyboardInterrupt and SystemExit, for one reason: a library's except Exception must not swallow a request to stop. Only the HTTP client knows HTTP, so it is the one place that translates status codes, socket errors and unusable answers into the taxonomy:

client.py
"""POST one finding to the ticket API; translate every failure into the taxonomy."""
import http.client
import json
import time
import urllib.error
import urllib.request
from errors import PermanentError, TransientError
from readbody import read_body
from retry_after import parse_retry_after
RETRYABLE = {429, 500, 502, 503, 504}
def post_ticket(api: str, finding: dict, key: str | None, timeout: float) -> dict:
"""timeout: the longest silence on the socket, and the time allowed until the body is in."""
fid = finding["id"]
headers = {"Content-Type": "application/json"}
if key:
headers["Idempotency-Key"] = key
req = urllib.request.Request(f"{api}/tickets", data=json.dumps(finding).encode(),
headers=headers, method="POST")
give_up_at = time.monotonic() + timeout
try:
with urllib.request.urlopen(req, timeout=timeout) as resp:
kind = resp.headers.get_content_type()
if kind != "application/json": # a proxy's or a captive portal's page, status 200
raise PermanentError(f"{fid}: expected JSON, got {kind}")
body = read_body(resp, fid, give_up_at)
except urllib.error.HTTPError as err: # 4xx and 5xx answers
what = f"{fid}: HTTP {err.code}"
if err.code in RETRYABLE:
raise TransientError(what, parse_retry_after(err.headers.get("Retry-After"))) from err
raise PermanentError(what) from err
except (OSError, http.client.HTTPException) as err: # timed out, refused, cut off mid-answer
raise TransientError(f"{fid}: {getattr(err, 'reason', err)}") from err
try:
return json.loads(body)
except ValueError as err: # garbled on the way: the same key makes a retry safe
raise TransientError(f"{fid}: invalid JSON: {err}") from err

The second except is py-sec's rule: OSError covers timeouts and refused connections, http.client.HTTPException an answer cut off halfway. A 200 is not proof of a ticket: a captive portal's page arrives as 200 text/html, and retrying gets the same page, so it is permanent. A body that fails to parse was garbled on the way; a retry with the same idempotency key is safe, so it is transient. read_body() belongs to the deadline section.

ticketapi.py chooses its failure from the finding id: busy- answers 429 once with Retry-After: 2, busydate- the same with an HTTP-date, slow- creates the ticket and then answers too late, bad- 422, down- always 503, later- asks for an hour, hang- never answers, html- returns a login page with status 200, and drip- sends its answer one byte every half second. Start it in the background, keeping its PID in api.pid, and run a clean batch and a mixed one:

deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ python3 ticketapi.py 2> api.log & echo $! > api.pid sleep 1; cat api.log
api: listening on http://127.0.0.1:18507
$ python3 ticketer.py findings.json; echo "exit status $?"
retry: busy-102: HTTP 429 (attempt 1), retrying in 2.4 s retry: busydate-103: HTTP 429 (attempt 1), retrying in 2.3 s retry: slow-104: timed out (attempt 1), retrying in 0.0 s ticketer: posted 4 tickets ticketer: cleanup done exit status 0
$ python3 ticketer.py findings-mixed.json; echo "exit status $?"
retry: down-203: HTTP 503 (attempt 1), retrying in 0.1 s retry: down-203: HTTP 503 (attempt 2), retrying in 0.2 s retry: down-203: HTTP 503 (attempt 3), retrying in 0.5 s ticketer: cleanup done ticketer: permanent: bad-202: HTTP 422 ticketer: permanent: html-205: expected JSON, got text/html ticketer: try later: down-203: HTTP 503 ticketer: try later: later-204: HTTP 429, Retry-After 3600 s is over our 30 s limit exit status 1

The first batch needed three retries (two 429s and a timeout) and ended with exit status 0. In the mixed batch, f-201 was posted, the 422 and the HTML page were reported as permanent without a retry, down-203 used its four attempts, and later-204 was not retried at all, because the server's hour is longer than this tool will ever wait. The exit status is 1: when a permanent failure and transient ones happen together, the one that needs a human wins.

ticketer uses the course's exit-code contract, which "Observable jobs" states in full later in this course: 0 success, 1 a human must look (a findings file it cannot read counts as that too), 2 a command line that argparse rejects, 70 a bug, and 75 only transient failures.

Many failures at once: ExceptionGroup and except*

A batch job should attempt every item and then report every failure, not stop at the first. post_all collects the failures and raises them together as an ExceptionGroup, a built-in exception (3.11+) that carries a list of exceptions:

job.py
"""The work: one ticket per finding, every finding attempted, failures reported together."""
import hashlib
from functools import partial
from client import post_ticket
from deadline import Deadline
from errors import DeadlineExceeded, ToolError
from retrypolicy import call_with_retries
def idempotency_key(finding: dict) -> str:
# From what identifies the finding, not the whole body: a field such as last_seen changes
# on every scan. Same finding, same key: in a retry, and in a re-run of the whole job.
return "ticketer-" + hashlib.sha256(finding["id"].encode()).hexdigest()[:32]
def post_all(findings: list[dict], api: str, deadline: Deadline) -> list[dict]:
posted, failed = [], []
for finding in findings:
if deadline.remaining() <= 0:
failed.append(DeadlineExceeded(f"{finding['id']}: not attempted, no time left"))
continue
send = partial(post_ticket, api, finding, idempotency_key(finding))
try:
posted.append(call_with_retries(send, deadline))
except ToolError as err:
failed.append(err) # keep going: one bad finding must not hide the others
if failed:
raise ExceptionGroup(f"{len(failed)} of {len(findings)} findings failed", failed)
return posted

partial(post_ticket, api, finding, key), from functools, builds a new function with those three arguments already filled in, so the retry loop only supplies the last one, the timeout. except* handles a group: each except* clause receives the sub-exceptions that match its class, as a smaller group, and several clauses can run for one group. This is the end of run() in ticketer.py (the whole file comes in the last section, once its other parts are explained):

ticketer.py (the end of run())
except* PermanentError as group:
for err in group.exceptions:
log.error("permanent: %s", err)
status = 1
except* TransientError as group:
for err in group.exceptions:
log.error("try later: %s", err)
status = status or 75
return status

Two rules come with except*. A clause cannot return, break or continue, so the code sets status and returns after the try. And except and except* cannot be mixed in one try. You will meet groups even where you did not create them, because asyncio.TaskGroup wraps every failure of its tasks in one. asyncio itself is the next lesson's subject; here you only need what the program prints:

tg_wrap.py
"""A TaskGroup wraps even a single failure in an ExceptionGroup."""
import asyncio
import sys
from errors import TransientError
async def check(name: str) -> str:
await asyncio.sleep(0.1)
if name == "api":
raise TransientError("api: HTTP 503")
return name
async def main() -> None:
async with asyncio.TaskGroup() as tg:
for name in ("dns", "api", "db"):
tg.create_task(check(name))
def plain() -> None:
try:
asyncio.run(main())
except TransientError as err: # never matches: the TaskGroup raises an ExceptionGroup
print("caught:", err)
def star() -> None:
try:
asyncio.run(main())
except* TransientError as group:
print("caught:", [str(err) for err in group.exceptions])
if __name__ == "__main__":
star() if "--star" in sys.argv else plain()
deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ python3 tg_wrap.py
+ Exception Group Traceback (most recent call last): … | ExceptionGroup: unhandled errors in a TaskGroup (1 sub-exception) +-+---------------- 1 ---------------- | Traceback (most recent call last): | File "/home/deploy/scr-py-resilience/tg_wrap.py", line 11, in check | raise TransientError("api: HTTP 503") | errors.TransientError: api: HTTP 503 +------------------------------------
$ python3 tg_wrap.py --star
caught: ['api: HTTP 503']

The plain except TransientError did not match: what arrived was an ExceptionGroup with one TransientError inside, and the program died with exit status 1. Code moved from a loop to a TaskGroup loses its error handling this way. except* matched inside the group; it also matches a single exception outside a group, which is why run() can use it for everything.

One deadline for the whole run, and Retry-After

A per-call timeout bounds one request, not a run: three findings, four attempts each, one second per attempt plus backoff is well over ten seconds, while a cron slot or a CI step has a fixed length. Deadline holds one budget for the run, and every call and every wait draws from it:

deadline.py
"""One time budget shared by every call and every wait in a run."""
import time
from errors import DeadlineExceeded
class Deadline:
def __init__(self, seconds: float) -> None:
self.end = time.monotonic() + seconds
def remaining(self) -> float:
return self.end - time.monotonic()
def timeout(self, cap: float) -> float:
"""Timeout for the next call: at most cap seconds, never past the end of the budget."""
left = self.remaining()
if left <= 0:
raise DeadlineExceeded("run deadline reached")
return min(cap, left)
retrypolicy.py
"""One retry loop that respects the run's deadline and the server's Retry-After."""
import logging
import random
import time
from collections.abc import Callable
from deadline import Deadline
from errors import DeadlineExceeded, TransientError
log = logging.getLogger("retry")
def call_with_retries(call: Callable[[float], dict], deadline: Deadline, attempts: int = 4,
per_call: float = 1.0, max_retry_after: float = 30.0) -> dict:
for attempt in range(1, attempts + 1):
try:
return call(deadline.timeout(per_call))
except TransientError as err:
if attempt == attempts:
raise
wait = random.uniform(0, min(8.0, 0.5 * 2 ** (attempt - 1))) # full jitter
if err.retry_after is not None:
if err.retry_after > max_retry_after:
raise TransientError(f"{err}, Retry-After {err.retry_after:.0f} s is over"
f" our {max_retry_after:.0f} s limit") from err
wait = err.retry_after + random.uniform(0, 0.5) # what it asked, plus jitter
if wait >= deadline.remaining():
raise DeadlineExceeded(f"{err}, no time left for a {wait:.1f} s wait") from err
log.warning("%s (attempt %d), retrying in %.1f s", err, attempt, wait)
time.sleep(wait)
raise DeadlineExceeded("no attempts allowed") # only reached with attempts < 1
deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ time python3 ticketer.py findings-hang.json --budget 3; echo "exit status $?"
retry: hang-301: timed out (attempt 1), retrying in 0.2 s retry: hang-301: timed out (attempt 2), retrying in 0.5 s ticketer: cleanup done ticketer: try later: hang-301: timed out, no time left for a 0.1 s wait ticketer: try later: hang-302: not attempted, no time left ticketer: try later: f-303: not attempted, no time left real 0m3.080s user 0m0.061s sys 0m0.021s exit status 75

With --budget 3 the run ended after 3.1 seconds. hang-301 used all of it: in this run two attempts timed out after a second each, the third got only what was left of the budget as its timeout, and the retry loop then gave up rather than start a 0.1-second wait it had no time for. hang-302 and f-303 are reported as not attempted. The random backoff decides whether a third attempt fits (other runs of the lab stopped after two) and how the budget is split between findings; the end time does not change. Exit status 75 tells the scheduler to try later.

That worked because the server went silent. "Calling HTTP APIs safely" in py-sec showed the other case: a socket timeout bounds silence, not duration, so a server that sends one byte every half second outlasts urlopen(timeout=1), and read_body() stops it by checking the clock after every chunk. post_ticket sets give_up_at before the request and reads the body with a copy of read_body() that raises ticketer's errors. The API's drip- findings answer that way:

readbody.py
"""Read a response body within a time limit (py-sec's read_body, raising ticketer's errors)."""
import time
from errors import PermanentError, TransientError
MAX_BODY = 1_000_000 # bytes; a larger answer is a bug or an attack, not a ticket
def read_body(resp, fid: str, give_up_at: float) -> bytes:
body = b""
while True:
chunk = resp.read1(65536) # returns as soon as some bytes have arrived
if not chunk:
return body
body += chunk
if len(body) > MAX_BODY:
raise PermanentError(f"{fid}: answer larger than {MAX_BODY} bytes")
if time.monotonic() > give_up_at: # checked per chunk: a drip cannot outlast it
raise TransientError(f"{fid}: answer still incomplete at the time limit")
deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ time python3 ticketer.py findings-drip.json --budget 3; echo "exit status $?"
retry: drip-501: answer still incomplete at the time limit (attempt 1), retrying in 0.1 s retry: drip-501: answer still incomplete at the time limit (attempt 2), retrying in 0.5 s ticketer: cleanup done ticketer: try later: drip-501: timed out, no time left for a 0.3 s wait real 0m3.101s user 0m0.078s sys 0m0.029s exit status 75

The first two attempts were cut off at their one-second limit although bytes kept arriving; the third had only the rest of the budget as its socket timeout and timed out, and the run ended after 3.1 seconds with a DeadlineExceeded and exit status 75. With other random waits the loop stops after one or two attempts, the last error is the read limit instead, and the run ends a little before the budget (2.3 seconds in one lab run); the exit status is 75 either way. Two limits remain. The check runs when a chunk arrives, so an attempt can overrun by one read timeout. And it does not cover the status line and headers, which urlopen() reads before it returns. A server that drips its headers (the API's driphead- mode) is stopped only from outside the process:

deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ time timeout 6 python3 ticketer.py findings-driphead.json --budget 2; echo "exit status $?"
ticketer: cleanup done ticketer: SIGTERM: stopped after cleanup real 0m6.047s user 0m0.047s sys 0m0.039s exit status 124

The 2-second budget did nothing; timeout sent SIGTERM after 6 seconds, ticketer cleaned up and died of the signal (the last section shows how), and timeout reported 124. Keep such a backstop around every scheduled job: timeout in the cron line, or TimeoutStartSec= on a systemd oneshot service ("Signals, timeouts and shutdown under systemd and Kubernetes" showed that RuntimeMaxSec= does nothing there).

The retry loop takes one more input: Retry-After. RFC 9110 allows two forms, a number of seconds or an HTTP-date, and a client that only parses the number silently ignores the date:

retry_after.py
"""Retry-After is either delay-seconds ("120") or an HTTP-date (RFC 9110, section 10.2.3)."""
import email.utils
from datetime import UTC, datetime
def parse_retry_after(value: str | None, now: datetime | None = None) -> float | None:
"""Seconds to wait, or None when the header is missing or unreadable."""
if value is None:
return None
value = value.strip()
if value.isascii() and value.isdigit(): # isdigit() alone accepts '²', which float() refuses
return float(value)
try:
when = email.utils.parsedate_to_datetime(value)
except (TypeError, ValueError):
return None # not a date either: use our own backoff
if when.tzinfo is None:
when = when.replace(tzinfo=UTC)
return max(0.0, (when - (now or datetime.now(UTC))).total_seconds())
deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ python3 - <<'EOF' from datetime import UTC, datetime from retry_after import parse_retry_after now = datetime(2026, 9, 27, 12, 0, 0, tzinfo=UTC) for value in ["120", "Sun, 27 Sep 2026 12:00:30 GMT", "Sun, 27 Sep 2026 11:59:00 GMT", "soon", "²", "٣"]: print(f"{value!r:34} -> {parse_retry_after(value, now)}") EOF
'120' -> 120.0 'Sun, 27 Sep 2026 12:00:30 GMT' -> 30.0 'Sun, 27 Sep 2026 11:59:00 GMT' -> 0.0 'soon' -> None '²' -> None '٣' -> None

Both forms became seconds, a past date 0, and text that is neither None ("use our own backoff"). The last two values are the trap "Calling HTTP APIs safely" showed: ² passes isdigit() and makes float() raise ValueError, which would escape post_ticket as a crash, and ٣ (Arabic-Indic three) is a digit float() accepts although RFC 9110 allows only ASCII digits. isascii() rejects both. email.utils.parsedate_to_datetime reads the HTTP-date, which is the email date format. In call_with_retries, a server's value replaces the random backoff, plus up to half a second of jitter so that clients told the same second do not return together. It is capped: over 30 seconds ends the retries at once, as later-204 showed, because exiting 75 beats sleeping an hour inside a job.

Idempotency keys: retrying a POST safely

Retrying a GET is harmless. A POST creates something, and a timeout does not say whether it did. The slow- findings reproduce the dangerous case: the server creates the ticket and then answers after the client has given up.

post_one.py
"""Post one finding with the retry policy, with or without an Idempotency-Key."""
import logging
import sys
from functools import partial
from client import post_ticket
from deadline import Deadline
from job import idempotency_key
from retrypolicy import call_with_retries
logging.basicConfig(format="%(name)s: %(message)s")
finding = {"id": sys.argv[1], "title": "ssh password spraying from 203.0.113.44"}
key = idempotency_key(finding) if "--key" in sys.argv else None
send = partial(post_ticket, "http://127.0.0.1:18507", finding, key)
print(call_with_retries(send, Deadline(10)))
deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ python3 post_one.py slow-1
retry: slow-1: timed out (attempt 1), retrying in 0.2 s {'ticket': 7, 'id': 'slow-1'}
$ python3 post_one.py slow-2 --key
retry: slow-2: timed out (attempt 1), retrying in 0.2 s {'ticket': 8, 'id': 'slow-2'}
$ curl -s http://127.0.0.1:18507/tickets | jq -M -c '.[] | select(.id | startswith("slow-"))'
{"ticket":4,"id":"slow-104"} {"ticket":6,"id":"slow-1"} {"ticket":7,"id":"slow-1"} {"ticket":8,"id":"slow-2"}

Without a key, the retry created a second ticket for slow-1 (tickets 6 and 7). With --key, the POST carried an Idempotency-Key header; the server had stored the answer for that key when it created ticket 8, so the retry got ticket 8 back and created nothing. The header comes from an IETF draft and your API must support it; never retry a non-idempotent call to an API that has no equivalent.

Look at how job.py builds the key: a hash of what identifies the finding, not a random UUID per attempt and not the whole body. Scanner output carries fields that change on every scan (last_seen, counts), and hashing those would give every run a new key and a new ticket. The key is also only as good as the server's memory: the draft lets servers expire keys (24 hours is common), so a re-run after a long outage is safe only inside that window, unless the server also refuses a second ticket for the same finding id.

Stopping cleanly: one instance, SIGTERM and ExitStack

By default Python does not handle SIGTERM: the process dies on the spot, and no finally or context manager runs. ticketer installs a handler, on_sigterm, that raises Terminated. Python runs signal handlers in the main thread between bytecodes, so the exception appears wherever the program is (in a socket read, in time.sleep) and unwinds the stack. ExitStack holds several context managers and callbacks and undoes them in reverse order however the block ends. The first thing on the stack is the lock that keeps a second copy from running at the same time:

lock.py
"""One run at a time: an exclusive, non-blocking flock on a file that is never deleted."""
import fcntl
from collections.abc import Iterator
from contextlib import contextmanager
from pathlib import Path
from errors import TransientError
@contextmanager
def single_instance(path: Path) -> Iterator[None]:
with path.open("a") as f:
try:
fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
except BlockingIOError:
raise TransientError(f"another run holds {path}") from None
yield # the lock goes when the file is closed, or when the process dies

@contextmanager turns a generator function into a context manager: the code before yield runs when the with block starts, the code after it (here, closing the file) when the block ends, even by an exception. Now the whole tool:

ticketer.py
"""ticketer FINDINGS.json: open one ticket per finding.
Exit status: 0 all posted, 1 a permanent failure (an unreadable findings file too), 2 usage,
70 a bug, 75 try again later; killed by SIGTERM after cleanup when asked to stop."""
import argparse
import json
import logging
import os
import signal
import tempfile
from contextlib import ExitStack
from pathlib import Path
from deadline import Deadline
from errors import PermanentError, Terminated, TransientError
from job import post_all
from lock import single_instance
log = logging.getLogger("ticketer")
def on_sigterm(signum: int, frame: object) -> None:
signal.signal(signal.SIGTERM, signal.SIG_IGN) # a second SIGTERM must not cut cleanup short
raise Terminated() # raised wherever the main thread is: in a read, in a sleep
def run(findings: list[dict], args: argparse.Namespace) -> int:
status = 0
try:
with ExitStack() as stack: # undone in reverse order, however the block ends
stack.callback(log.info, "cleanup done") # registered first, so it runs last
stack.enter_context(single_instance(Path("ticketer.lock")))
spool = Path(stack.enter_context(
tempfile.TemporaryDirectory(prefix="ticketer-", dir="."))) # same filesystem
posted = post_all(findings, args.api, Deadline(args.budget))
(spool / "report.json").write_text(json.dumps(posted))
os.replace(spool / "report.json", "report.json")
log.info("posted %d tickets", len(posted))
except* PermanentError as group:
for err in group.exceptions:
log.error("permanent: %s", err)
status = 1
except* TransientError as group:
for err in group.exceptions:
log.error("try later: %s", err)
status = status or 75
return status
def main() -> int:
p = argparse.ArgumentParser(prog="ticketer", allow_abbrev=False)
p.add_argument("findings", type=Path)
p.add_argument("--api", default="http://127.0.0.1:18507")
p.add_argument("--budget", type=float, default=10.0, help="seconds for the whole run")
args = p.parse_args() # a bad option or a missing argument exits 2 here
logging.basicConfig(level=logging.INFO, format="%(name)s: %(message)s")
try:
findings = json.loads(args.findings.read_text())
except (OSError, ValueError) as err: # bad input, not bad usage: a human fixes the file
log.error("cannot read %s: %s", args.findings, err)
return 1
signal.signal(signal.SIGTERM, on_sigterm)
try:
return run(findings, args)
except Terminated:
log.warning("SIGTERM: stopped after cleanup")
signal.signal(signal.SIGTERM, signal.SIG_DFL)
os.kill(os.getpid(), signal.SIGTERM) # die of the signal: a clean stop to systemd
raise
except Exception: # outside the taxonomy: a bug, so keep the traceback
log.exception("unexpected error")
return 70 # EX_SOFTWARE
if __name__ == "__main__":
raise SystemExit(main())

The first line of on_sigterm switches SIGTERM to SIG_IGN. Without it, a second SIGTERM (a second kill, or timeout and systemd both stopping the job) would raise Terminated inside the cleanup and skip the rest, the problem "Signals, timeouts and shutdown" solved for Bash by ignoring TERM in the handler. The Python docs warn that code cannot be made fully safe against exceptions from signal handlers; a handler that only sets a flag, checked between findings, is the sturdier design for long batch jobs. Run a job, try a second copy, and stop the first:

deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ python3 ticketer.py findings-hang.json --budget 20 & sleep 1.5 python3 ticketer.py findings.json; echo "second run: exit status $?" kill -TERM $!; wait $!; echo "first run: exit status $?" ls -d ticketer-* 2>/dev/null | wc -l
retry: hang-301: timed out (attempt 1), retrying in 0.4 s ticketer: cleanup done ticketer: try later: another run holds ticketer.lock second run: exit status 75 ticketer: cleanup done ticketer: SIGTERM: stopped after cleanup first run: exit status 143 0
$ python3 ticketer.py findings.json; echo "exit status $?"; tail -n 4 api.log
ticketer: posted 4 tickets ticketer: cleanup done exit status 0 api: POST /tickets -> 200 (replayed) api: POST /tickets -> 200 (replayed) api: POST /tickets -> 200 (replayed) api: POST /tickets -> 200 (replayed)

The second run could not take the lock and exited 75 at once. The first got SIGTERM in a hanging request: the stack released the lock and removed the spool directory, and main() logged the stop, reset SIGTERM to its default action and sent it to itself, so the process died of SIGTERM (143 in the shell) after its cleanup. systemd counts that as a clean stop and an exit 143 as a failure, as the Bash lesson showed. No spool directory was left. The last run posted all four findings again and the API answered every one from its idempotency store, so re-running the job created no duplicates.

What happens on SIGTERM
1SIGTERM arrives
the handler ignores further SIGTERMs
2raise Terminated
from inside the read or the sleep
3ExitStack unwinds
spool removed, lock released, in reverse order
4except Terminated in main()
log, restore the default action
5os.kill(own pid, SIGTERM)
death by signal: a clean stop

The lock follows the rules "Temp files, locks and timeouts" in bash-ops set for flock(1): non-blocking, on a file that is never deleted, and released by the kernel when the holder dies, even by SIGKILL. What is specific to Python is inheritance. Child processes started with subprocess do not inherit the descriptor, because Python opens files as non-inheritable (close-on-exec, which "Files, descriptors and the VFS" in Advanced Linux internals explains); a child made by os.fork() without an exec does share it, and the lock with it.

Try this

When the API is down, every finding spends its four attempts. Run time python3 ticketer.py findings-down.json (five down- findings): the lab took about 8 seconds of its 10-second budget. Add a simple circuit breaker to post_all: after 3 transient failures in a row, record the remaining findings as TransientError(f"{finding['id']}: skipped, the API keeps failing") without calling the API, and reset the count after a success. The lab's version finished in about 5 seconds (the random waits move it by a second either way) with three "HTTP 503" lines, two "skipped" lines and exit status 75. Then explain why the breaker counts only transient failures: a 422 says nothing about whether the API is up. Finally, stop the test API:

deploy@web01:~/scr-py-resilience · Ubuntu 26.04 LTS
$ kill "$(cat api.pid)" && rm api.pid && echo "api stopped"
api stopped

Takeaway

Sort failures by what the caller must do, attempt every item and report the failures together, give the whole run one deadline that retries, Retry-After waits and slow answers must fit inside (with timeout or systemd outside as the backstop), retry a POST only with a key derived from what the item is, and turn SIGTERM into an exception so that ExitStack cleans up before the process dies of the signal.

Quick check
01A job's retry loop honours Retry-After by calling int(resp.headers["Retry-After"]). One day the API starts sending Retry-After: Wed, 30 Sep 2026 12:00:05 GMT. What happens, and what is the fix?
Correct — RFC 9110 allows delay-seconds or an HTTP-date; parse_retry_after handles both and returns None for anything else.
Incorrect — int() does not read part of a string; it raises ValueError. Stripping text by hand would still not turn a date into a delay.
Incorrect — urllib hands you the header text as the server sent it; converting it is the client's job, as the lab's parser does.
Incorrect — The failure happens before any wait: int() cannot parse the date. A cap is still needed, but only after a correct parse.
02ticketer posts a finding, the read times out, and the retry succeeds. Without an idempotency key, which outcome is possible, and which key design prevents it across retries and re-runs?
Incorrect — The lab showed the opposite: the server created ticket 6 and then answered too late, so the retry created ticket 7.
Incorrect — A new key per attempt tells the server each attempt is a new operation, which is exactly what creates the duplicate.
Correct — The server stores the answer per key; a stable key turns a retry or a second run (inside the server's key window) into a replay.
Incorrect — Reusing a connection does not merge two POSTs into one; the server still receives and executes both requests.
03A Python job run by systemd holds a lock and a temp directory. After systemctl stop, the temp directory is still there. The code uses with blocks for both. What is the most likely cause?
Incorrect — Context managers run their exit on any exception, BaseException included; and without a handler SIGTERM raises nothing at all.
Correct — Python's default for SIGTERM is to die; install a handler that raises, and the with blocks (or an ExitStack) unwind.
Incorrect — systemd sends SIGTERM first and SIGKILL only after TimeoutStopSec; a handler that raises lets the cleanup run in that window.
Incorrect — Where the directory lives does not matter here; the cleanup code never ran, because nothing turned SIGTERM into an exception.

Related