Failure design: exceptions, deadlines, retries and cleanup
Exception taxonomy, deadline budgets, idempotent retries, SIGTERM and ExitStack.
tar -xzf scr-py-resilience.tar.gz, which creates scr-py-resilience/. SHA-256: b6a2296242282fb86f152cf2264da477429562acff492438d299b1bd8a163ccaThis lesson designs the failure behaviour of a whole tool, not of one call. You will give a tool an exception taxonomy that every layer shares, report many failures at once with ExceptionGroup and except*, spend one time budget across every call, wait and slow answer of a run, honour Retry-After in both of its forms without letting a server park your job for an hour, make a retried POST safe with an idempotency key, and stop cleanly on SIGTERM with contextlib.ExitStack and an fcntl lock. The tool is ticketer, which opens one ticket per security finding against a local API that fails on purpose. Stdlib only; it behaves the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.
Refresher: "Logging, retries and failure design" in py-sec owns temporary versus permanent errors, raise ... from, logging to stderr, and one retry loop with full jitter and a deadline; "Calling HTTP APIs safely" owns timeouts, 429, a capped Retry-After, and read_body(), which reads an answer within a time budget. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-py-resilience. Besides the modules shown here, they hold the findings*.json batches and ticketapi.py, the test API described below.
A taxonomy the whole tool agrees on
Each kind of failure needs a different response from whoever runs the tool, so ticketer sorts every failure into one of these classes, and the class carries the exit status:
"""What can go wrong in ticketer, sorted by what the caller should do about it."""class ToolError(Exception):exit_status = 1class PermanentError(ToolError):"""Retrying cannot help: the request or the data is wrong. Exit 1."""class TransientError(ToolError):"""May work later. retry_after: the wait the server asked for, in seconds, if any."""exit_status = 75 # EX_TEMPFAIL: the scheduler should run the job again laterdef __init__(self, message: str, retry_after: float | None = None) -> None:super().__init__(message)self.retry_after = retry_afterclass DeadlineExceeded(TransientError):"""The run's time budget is spent."""class Terminated(BaseException):"""SIGTERM arrived. A BaseException, like KeyboardInterrupt, so "except Exception" inlibrary code cannot swallow a request to stop."""
PermanentError means a human must fix something, exit status 1. TransientError means "run me again later", exit status 75 (EX_TEMPFAIL from sysexits.h), and can carry the wait the server asked for. DeadlineExceeded is a kind of transient error. Terminated derives from BaseException, like KeyboardInterrupt and SystemExit, for one reason: a library's except Exception must not swallow a request to stop. Only the HTTP client knows HTTP, so it is the one place that translates status codes, socket errors and unusable answers into the taxonomy:
"""POST one finding to the ticket API; translate every failure into the taxonomy."""import http.clientimport jsonimport timeimport urllib.errorimport urllib.requestfrom errors import PermanentError, TransientErrorfrom readbody import read_bodyfrom retry_after import parse_retry_afterRETRYABLE = {429, 500, 502, 503, 504}def post_ticket(api: str, finding: dict, key: str | None, timeout: float) -> dict:"""timeout: the longest silence on the socket, and the time allowed until the body is in."""fid = finding["id"]headers = {"Content-Type": "application/json"}if key:headers["Idempotency-Key"] = keyreq = urllib.request.Request(f"{api}/tickets", data=json.dumps(finding).encode(),headers=headers, method="POST")give_up_at = time.monotonic() + timeouttry:with urllib.request.urlopen(req, timeout=timeout) as resp:kind = resp.headers.get_content_type()if kind != "application/json": # a proxy's or a captive portal's page, status 200raise PermanentError(f"{fid}: expected JSON, got {kind}")body = read_body(resp, fid, give_up_at)except urllib.error.HTTPError as err: # 4xx and 5xx answerswhat = f"{fid}: HTTP {err.code}"if err.code in RETRYABLE:raise TransientError(what, parse_retry_after(err.headers.get("Retry-After"))) from errraise PermanentError(what) from errexcept (OSError, http.client.HTTPException) as err: # timed out, refused, cut off mid-answerraise TransientError(f"{fid}: {getattr(err, 'reason', err)}") from errtry:return json.loads(body)except ValueError as err: # garbled on the way: the same key makes a retry saferaise TransientError(f"{fid}: invalid JSON: {err}") from err
The second except is py-sec's rule: OSError covers timeouts and refused connections, http.client.HTTPException an answer cut off halfway. A 200 is not proof of a ticket: a captive portal's page arrives as 200 text/html, and retrying gets the same page, so it is permanent. A body that fails to parse was garbled on the way; a retry with the same idempotency key is safe, so it is transient. read_body() belongs to the deadline section.
ticketapi.py chooses its failure from the finding id: busy- answers 429 once with Retry-After: 2, busydate- the same with an HTTP-date, slow- creates the ticket and then answers too late, bad- 422, down- always 503, later- asks for an hour, hang- never answers, html- returns a login page with status 200, and drip- sends its answer one byte every half second. Start it in the background, keeping its PID in api.pid, and run a clean batch and a mixed one:
The first batch needed three retries (two 429s and a timeout) and ended with exit status 0. In the mixed batch, f-201 was posted, the 422 and the HTML page were reported as permanent without a retry, down-203 used its four attempts, and later-204 was not retried at all, because the server's hour is longer than this tool will ever wait. The exit status is 1: when a permanent failure and transient ones happen together, the one that needs a human wins.
ticketer uses the course's exit-code contract, which "Observable jobs" states in full later in this course: 0 success, 1 a human must look (a findings file it cannot read counts as that too), 2 a command line that argparse rejects, 70 a bug, and 75 only transient failures.
Many failures at once: ExceptionGroup and except*
A batch job should attempt every item and then report every failure, not stop at the first. post_all collects the failures and raises them together as an ExceptionGroup, a built-in exception (3.11+) that carries a list of exceptions:
"""The work: one ticket per finding, every finding attempted, failures reported together."""import hashlibfrom functools import partialfrom client import post_ticketfrom deadline import Deadlinefrom errors import DeadlineExceeded, ToolErrorfrom retrypolicy import call_with_retriesdef idempotency_key(finding: dict) -> str:# From what identifies the finding, not the whole body: a field such as last_seen changes# on every scan. Same finding, same key: in a retry, and in a re-run of the whole job.return "ticketer-" + hashlib.sha256(finding["id"].encode()).hexdigest()[:32]def post_all(findings: list[dict], api: str, deadline: Deadline) -> list[dict]:posted, failed = [], []for finding in findings:if deadline.remaining() <= 0:failed.append(DeadlineExceeded(f"{finding['id']}: not attempted, no time left"))continuesend = partial(post_ticket, api, finding, idempotency_key(finding))try:posted.append(call_with_retries(send, deadline))except ToolError as err:failed.append(err) # keep going: one bad finding must not hide the othersif failed:raise ExceptionGroup(f"{len(failed)} of {len(findings)} findings failed", failed)return posted
partial(post_ticket, api, finding, key), from functools, builds a new function with those three arguments already filled in, so the retry loop only supplies the last one, the timeout. except* handles a group: each except* clause receives the sub-exceptions that match its class, as a smaller group, and several clauses can run for one group. This is the end of run() in ticketer.py (the whole file comes in the last section, once its other parts are explained):
except* PermanentError as group:for err in group.exceptions:log.error("permanent: %s", err)status = 1except* TransientError as group:for err in group.exceptions:log.error("try later: %s", err)status = status or 75return status
Two rules come with except*. A clause cannot return, break or continue, so the code sets status and returns after the try. And except and except* cannot be mixed in one try. You will meet groups even where you did not create them, because asyncio.TaskGroup wraps every failure of its tasks in one. asyncio itself is the next lesson's subject; here you only need what the program prints:
"""A TaskGroup wraps even a single failure in an ExceptionGroup."""import asyncioimport sysfrom errors import TransientErrorasync def check(name: str) -> str:await asyncio.sleep(0.1)if name == "api":raise TransientError("api: HTTP 503")return nameasync def main() -> None:async with asyncio.TaskGroup() as tg:for name in ("dns", "api", "db"):tg.create_task(check(name))def plain() -> None:try:asyncio.run(main())except TransientError as err: # never matches: the TaskGroup raises an ExceptionGroupprint("caught:", err)def star() -> None:try:asyncio.run(main())except* TransientError as group:print("caught:", [str(err) for err in group.exceptions])if __name__ == "__main__":star() if "--star" in sys.argv else plain()
The plain except TransientError did not match: what arrived was an ExceptionGroup with one TransientError inside, and the program died with exit status 1. Code moved from a loop to a TaskGroup loses its error handling this way. except* matched inside the group; it also matches a single exception outside a group, which is why run() can use it for everything.
One deadline for the whole run, and Retry-After
A per-call timeout bounds one request, not a run: three findings, four attempts each, one second per attempt plus backoff is well over ten seconds, while a cron slot or a CI step has a fixed length. Deadline holds one budget for the run, and every call and every wait draws from it:
"""One time budget shared by every call and every wait in a run."""import timefrom errors import DeadlineExceededclass Deadline:def __init__(self, seconds: float) -> None:self.end = time.monotonic() + secondsdef remaining(self) -> float:return self.end - time.monotonic()def timeout(self, cap: float) -> float:"""Timeout for the next call: at most cap seconds, never past the end of the budget."""left = self.remaining()if left <= 0:raise DeadlineExceeded("run deadline reached")return min(cap, left)
"""One retry loop that respects the run's deadline and the server's Retry-After."""import loggingimport randomimport timefrom collections.abc import Callablefrom deadline import Deadlinefrom errors import DeadlineExceeded, TransientErrorlog = logging.getLogger("retry")def call_with_retries(call: Callable[[float], dict], deadline: Deadline, attempts: int = 4,per_call: float = 1.0, max_retry_after: float = 30.0) -> dict:for attempt in range(1, attempts + 1):try:return call(deadline.timeout(per_call))except TransientError as err:if attempt == attempts:raisewait = random.uniform(0, min(8.0, 0.5 * 2 ** (attempt - 1))) # full jitterif err.retry_after is not None:if err.retry_after > max_retry_after:raise TransientError(f"{err}, Retry-After {err.retry_after:.0f} s is over"f" our {max_retry_after:.0f} s limit") from errwait = err.retry_after + random.uniform(0, 0.5) # what it asked, plus jitterif wait >= deadline.remaining():raise DeadlineExceeded(f"{err}, no time left for a {wait:.1f} s wait") from errlog.warning("%s (attempt %d), retrying in %.1f s", err, attempt, wait)time.sleep(wait)raise DeadlineExceeded("no attempts allowed") # only reached with attempts < 1
With --budget 3 the run ended after 3.1 seconds. hang-301 used all of it: in this run two attempts timed out after a second each, the third got only what was left of the budget as its timeout, and the retry loop then gave up rather than start a 0.1-second wait it had no time for. hang-302 and f-303 are reported as not attempted. The random backoff decides whether a third attempt fits (other runs of the lab stopped after two) and how the budget is split between findings; the end time does not change. Exit status 75 tells the scheduler to try later.
That worked because the server went silent. "Calling HTTP APIs safely" in py-sec showed the other case: a socket timeout bounds silence, not duration, so a server that sends one byte every half second outlasts urlopen(timeout=1), and read_body() stops it by checking the clock after every chunk. post_ticket sets give_up_at before the request and reads the body with a copy of read_body() that raises ticketer's errors. The API's drip- findings answer that way:
"""Read a response body within a time limit (py-sec's read_body, raising ticketer's errors)."""import timefrom errors import PermanentError, TransientErrorMAX_BODY = 1_000_000 # bytes; a larger answer is a bug or an attack, not a ticketdef read_body(resp, fid: str, give_up_at: float) -> bytes:body = b""while True:chunk = resp.read1(65536) # returns as soon as some bytes have arrivedif not chunk:return bodybody += chunkif len(body) > MAX_BODY:raise PermanentError(f"{fid}: answer larger than {MAX_BODY} bytes")if time.monotonic() > give_up_at: # checked per chunk: a drip cannot outlast itraise TransientError(f"{fid}: answer still incomplete at the time limit")
The first two attempts were cut off at their one-second limit although bytes kept arriving; the third had only the rest of the budget as its socket timeout and timed out, and the run ended after 3.1 seconds with a DeadlineExceeded and exit status 75. With other random waits the loop stops after one or two attempts, the last error is the read limit instead, and the run ends a little before the budget (2.3 seconds in one lab run); the exit status is 75 either way. Two limits remain. The check runs when a chunk arrives, so an attempt can overrun by one read timeout. And it does not cover the status line and headers, which urlopen() reads before it returns. A server that drips its headers (the API's driphead- mode) is stopped only from outside the process:
The 2-second budget did nothing; timeout sent SIGTERM after 6 seconds, ticketer cleaned up and died of the signal (the last section shows how), and timeout reported 124. Keep such a backstop around every scheduled job: timeout in the cron line, or TimeoutStartSec= on a systemd oneshot service ("Signals, timeouts and shutdown under systemd and Kubernetes" showed that RuntimeMaxSec= does nothing there).
The retry loop takes one more input: Retry-After. RFC 9110 allows two forms, a number of seconds or an HTTP-date, and a client that only parses the number silently ignores the date:
"""Retry-After is either delay-seconds ("120") or an HTTP-date (RFC 9110, section 10.2.3)."""import email.utilsfrom datetime import UTC, datetimedef parse_retry_after(value: str | None, now: datetime | None = None) -> float | None:"""Seconds to wait, or None when the header is missing or unreadable."""if value is None:return Nonevalue = value.strip()if value.isascii() and value.isdigit(): # isdigit() alone accepts '²', which float() refusesreturn float(value)try:when = email.utils.parsedate_to_datetime(value)except (TypeError, ValueError):return None # not a date either: use our own backoffif when.tzinfo is None:when = when.replace(tzinfo=UTC)return max(0.0, (when - (now or datetime.now(UTC))).total_seconds())
Both forms became seconds, a past date 0, and text that is neither None ("use our own backoff"). The last two values are the trap "Calling HTTP APIs safely" showed: ² passes isdigit() and makes float() raise ValueError, which would escape post_ticket as a crash, and ٣ (Arabic-Indic three) is a digit float() accepts although RFC 9110 allows only ASCII digits. isascii() rejects both. email.utils.parsedate_to_datetime reads the HTTP-date, which is the email date format. In call_with_retries, a server's value replaces the random backoff, plus up to half a second of jitter so that clients told the same second do not return together. It is capped: over 30 seconds ends the retries at once, as later-204 showed, because exiting 75 beats sleeping an hour inside a job.
Idempotency keys: retrying a POST safely
Retrying a GET is harmless. A POST creates something, and a timeout does not say whether it did. The slow- findings reproduce the dangerous case: the server creates the ticket and then answers after the client has given up.
"""Post one finding with the retry policy, with or without an Idempotency-Key."""import loggingimport sysfrom functools import partialfrom client import post_ticketfrom deadline import Deadlinefrom job import idempotency_keyfrom retrypolicy import call_with_retrieslogging.basicConfig(format="%(name)s: %(message)s")finding = {"id": sys.argv[1], "title": "ssh password spraying from 203.0.113.44"}key = idempotency_key(finding) if "--key" in sys.argv else Nonesend = partial(post_ticket, "http://127.0.0.1:18507", finding, key)print(call_with_retries(send, Deadline(10)))
Without a key, the retry created a second ticket for slow-1 (tickets 6 and 7). With --key, the POST carried an Idempotency-Key header; the server had stored the answer for that key when it created ticket 8, so the retry got ticket 8 back and created nothing. The header comes from an IETF draft and your API must support it; never retry a non-idempotent call to an API that has no equivalent.
Look at how job.py builds the key: a hash of what identifies the finding, not a random UUID per attempt and not the whole body. Scanner output carries fields that change on every scan (last_seen, counts), and hashing those would give every run a new key and a new ticket. The key is also only as good as the server's memory: the draft lets servers expire keys (24 hours is common), so a re-run after a long outage is safe only inside that window, unless the server also refuses a second ticket for the same finding id.
Stopping cleanly: one instance, SIGTERM and ExitStack
By default Python does not handle SIGTERM: the process dies on the spot, and no finally or context manager runs. ticketer installs a handler, on_sigterm, that raises Terminated. Python runs signal handlers in the main thread between bytecodes, so the exception appears wherever the program is (in a socket read, in time.sleep) and unwinds the stack. ExitStack holds several context managers and callbacks and undoes them in reverse order however the block ends. The first thing on the stack is the lock that keeps a second copy from running at the same time:
"""One run at a time: an exclusive, non-blocking flock on a file that is never deleted."""import fcntlfrom collections.abc import Iteratorfrom contextlib import contextmanagerfrom pathlib import Pathfrom errors import TransientError@contextmanagerdef single_instance(path: Path) -> Iterator[None]:with path.open("a") as f:try:fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)except BlockingIOError:raise TransientError(f"another run holds {path}") from Noneyield # the lock goes when the file is closed, or when the process dies
@contextmanager turns a generator function into a context manager: the code before yield runs when the with block starts, the code after it (here, closing the file) when the block ends, even by an exception. Now the whole tool:
"""ticketer FINDINGS.json: open one ticket per finding.Exit status: 0 all posted, 1 a permanent failure (an unreadable findings file too), 2 usage,70 a bug, 75 try again later; killed by SIGTERM after cleanup when asked to stop."""import argparseimport jsonimport loggingimport osimport signalimport tempfilefrom contextlib import ExitStackfrom pathlib import Pathfrom deadline import Deadlinefrom errors import PermanentError, Terminated, TransientErrorfrom job import post_allfrom lock import single_instancelog = logging.getLogger("ticketer")def on_sigterm(signum: int, frame: object) -> None:signal.signal(signal.SIGTERM, signal.SIG_IGN) # a second SIGTERM must not cut cleanup shortraise Terminated() # raised wherever the main thread is: in a read, in a sleepdef run(findings: list[dict], args: argparse.Namespace) -> int:status = 0try:with ExitStack() as stack: # undone in reverse order, however the block endsstack.callback(log.info, "cleanup done") # registered first, so it runs laststack.enter_context(single_instance(Path("ticketer.lock")))spool = Path(stack.enter_context(tempfile.TemporaryDirectory(prefix="ticketer-", dir="."))) # same filesystemposted = post_all(findings, args.api, Deadline(args.budget))(spool / "report.json").write_text(json.dumps(posted))os.replace(spool / "report.json", "report.json")log.info("posted %d tickets", len(posted))except* PermanentError as group:for err in group.exceptions:log.error("permanent: %s", err)status = 1except* TransientError as group:for err in group.exceptions:log.error("try later: %s", err)status = status or 75return statusdef main() -> int:p = argparse.ArgumentParser(prog="ticketer", allow_abbrev=False)p.add_argument("findings", type=Path)p.add_argument("--api", default="http://127.0.0.1:18507")p.add_argument("--budget", type=float, default=10.0, help="seconds for the whole run")args = p.parse_args() # a bad option or a missing argument exits 2 herelogging.basicConfig(level=logging.INFO, format="%(name)s: %(message)s")try:findings = json.loads(args.findings.read_text())except (OSError, ValueError) as err: # bad input, not bad usage: a human fixes the filelog.error("cannot read %s: %s", args.findings, err)return 1signal.signal(signal.SIGTERM, on_sigterm)try:return run(findings, args)except Terminated:log.warning("SIGTERM: stopped after cleanup")signal.signal(signal.SIGTERM, signal.SIG_DFL)os.kill(os.getpid(), signal.SIGTERM) # die of the signal: a clean stop to systemdraiseexcept Exception: # outside the taxonomy: a bug, so keep the tracebacklog.exception("unexpected error")return 70 # EX_SOFTWAREif __name__ == "__main__":raise SystemExit(main())
The first line of on_sigterm switches SIGTERM to SIG_IGN. Without it, a second SIGTERM (a second kill, or timeout and systemd both stopping the job) would raise Terminated inside the cleanup and skip the rest, the problem "Signals, timeouts and shutdown" solved for Bash by ignoring TERM in the handler. The Python docs warn that code cannot be made fully safe against exceptions from signal handlers; a handler that only sets a flag, checked between findings, is the sturdier design for long batch jobs. Run a job, try a second copy, and stop the first:
The second run could not take the lock and exited 75 at once. The first got SIGTERM in a hanging request: the stack released the lock and removed the spool directory, and main() logged the stop, reset SIGTERM to its default action and sent it to itself, so the process died of SIGTERM (143 in the shell) after its cleanup. systemd counts that as a clean stop and an exit 143 as a failure, as the Bash lesson showed. No spool directory was left. The last run posted all four findings again and the API answered every one from its idempotency store, so re-running the job created no duplicates.
The lock follows the rules "Temp files, locks and timeouts" in bash-ops set for flock(1): non-blocking, on a file that is never deleted, and released by the kernel when the holder dies, even by SIGKILL. What is specific to Python is inheritance. Child processes started with subprocess do not inherit the descriptor, because Python opens files as non-inheritable (close-on-exec, which "Files, descriptors and the VFS" in Advanced Linux internals explains); a child made by os.fork() without an exec does share it, and the lock with it.
Try this
When the API is down, every finding spends its four attempts. Run time python3 ticketer.py findings-down.json (five down- findings): the lab took about 8 seconds of its 10-second budget. Add a simple circuit breaker to post_all: after 3 transient failures in a row, record the remaining findings as TransientError(f"{finding['id']}: skipped, the API keeps failing") without calling the API, and reset the count after a success. The lab's version finished in about 5 seconds (the random waits move it by a second either way) with three "HTTP 503" lines, two "skipped" lines and exit status 75. Then explain why the breaker counts only transient failures: a 422 says nothing about whether the API is up. Finally, stop the test API:
Takeaway
Sort failures by what the caller must do, attempt every item and report the failures together, give the whole run one deadline that retries, Retry-After waits and slow answers must fit inside (with timeout or systemd outside as the backstop), retry a POST only with a key derived from what the item is, and turn SIGTERM into an exception so that ExitStack cleans up before the process dies of the signal.
Retry-After by calling int(resp.headers["Retry-After"]). One day the API starts sending Retry-After: Wed, 30 Sep 2026 12:00:05 GMT. What happens, and what is the fix?ticketer posts a finding, the read times out, and the retry succeeds. Without an idempotency key, which outcome is possible, and which key design prevents it across retries and re-runs?systemctl stop, the temp directory is still there. The code uses with blocks for both. What is the most likely cause?