Handling hostile input at scale
Size and time budgets, ReDoS, deserialization choices, validation and TOCTOU.
tar -xzf scr-py-secure.tar.gz, which creates scr-py-secure/. SHA-256: bb07476208639e87f5d8b64fefcf7a7f05f0b1ccf7bf8bfd1370686a6a9aaf87Automation that reads logs, uploads, API responses and files handles input someone else controls, and at scale one hostile record can take the whole job down. This lesson adds the defences that matter when the volume is large and the source is not trusted: budgets that refuse input too big or too slow to accept, regular expressions that cannot be made to hang, deserialization that cannot run code, validation that turns a raw value into a checked object, file access that a swapped-in symlink cannot redirect, and logging that a crafted field cannot forge. The tools are standard library except two matching engines, pinned in a venv; the behaviour is the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.
This lesson builds on three py-sec lessons: "JSON, CSV, YAML and TOML" owns json, yaml.safe_load and CSV injection; "Regular expressions for logs and tool output" owns the ReDoS shape and its fixes; "Files, paths and safe writes" owns pathlib, atomic writes and is_relative_to containment. Here each idea is pushed to production scale. (This lesson also absorbs the retired "Parsing at scale" lesson, whose URL now redirects here.) Unpack the Lesson files in your home directory: they create scr-py-secure/ with every script and the two NDJSON fixtures, and the commands below run inside it. The terminals show the lab's deploy account; yours shows your own.
Budgets: refuse what is too big or too slow
The first defence at scale is refusing to hold too much, or to wait too long. A decompression bomb, a file with a billion lines, one line that never ends and a sender that trickles one byte at a time all end in an outage unless a budget stops them first. limits.py defines the budgets and a read that honours two of them, total bytes and total time:
"""Budgets for untrusted input, and a read that honours two of them: total bytes and total time."""import osimport selectimport timefrom dataclasses import dataclass@dataclass(frozen=True)class Budget:max_bytes: int = 1_000_000 # never read more than this from the source at allmax_seconds: float = 5.0 # the whole input must arrive within this timemax_lines: int = 10_000 # a record count a single run should never exceedmax_line_bytes: int = 8_192 # one record is small; a huge line is an attack or a bugclass BudgetExceeded(Exception):"""The input is larger or slower than we agreed to process: reject it, do not truncate it."""def read_capped(fd: int, budget: Budget) -> bytes:# Chunk by chunk: memory stays bounded even if the source is endless, and every wait draws# from one deadline, so a sender that trickles a byte at a time cannot hold the worker.give_up_at = time.monotonic() + budget.max_secondschunks, size = [], 0while True:remaining = give_up_at - time.monotonic()if remaining <= 0 or not select.select([fd], [], [], remaining)[0]:raise BudgetExceeded(f"input still incomplete after {budget.max_seconds} s")chunk = os.read(fd, 65_536)if not chunk: # end of inputreturn b"".join(chunks)size += len(chunk)if size > budget.max_bytes:raise BudgetExceeded(f"input exceeds {budget.max_bytes} bytes")chunks.append(chunk)
read_capped reads in chunks and stops as soon as the total passes max_bytes, so an oversized input is never pulled into memory in full. The time budget works the same way: give_up_at is fixed once, and each select.select call waits only for the time that is left, so the deadline covers the whole input rather than each read. A plain stream.read() on a pipe or socket has no such limit; it waits until the bytes or the end arrive. budgets.py splits the capped bytes into records and applies the other two budgets:
"""Split capped input into records under two more budgets: line count and single-line length."""import sysfrom limits import Budget, BudgetExceeded, read_cappeddef read_records(fd: int, budget: Budget):lines = read_capped(fd, budget).split(b"\n")if lines[-1] == b"": # a final newline ends the last line, it adds nonelines.pop()if len(lines) > budget.max_lines:raise BudgetExceeded(f"more than {budget.max_lines} lines")for n, raw in enumerate(lines, 1):if len(raw) > budget.max_line_bytes:raise BudgetExceeded(f"line {n} is {len(raw)} bytes, over {budget.max_line_bytes}")if raw:yield rawif __name__ == "__main__":try:count = sum(1 for _ in read_records(sys.stdin.fileno(), Budget()))except BudgetExceeded as err:print(f"rejected: {err}", file=sys.stderr)raise SystemExit(1)print(f"accepted {count} records")
Each rejection names the budget it hit and exits 1, so the caller learns the input was refused rather than silently truncated. The line-count case shows the boundary: exactly 10,000 lines is accepted and 10,001 is not. That needs the lines.pop() line, because splitting newline-terminated text on \n leaves one empty item after the last newline. The slow sender needed ten seconds to deliver twenty bytes; the reader gave up after five, and the sender's next write went to a closed pipe. A budget is a policy decision you write down: size it to the largest legitimate input plus headroom.
ReDoS at scale: a length cap is a budget, not a fix
py-sec showed that a quantifier nested inside a repeated group, like ^(\w+\s?)+$, backtracks exponentially on an almost-matching string. At fleet scale that pattern runs on every record, so one crafted value stalls a worker. A tempting answer is to cap the field at its real maximum (a host name is at most 253 characters) and let the cap keep long attack strings away from the engine. Watch the cap fail:
"""Time the built-in re engine on a nested-quantifier pattern against an almost-matching string.The pattern validates space-separated host names. Each argument is the total length of the input."""import reimport sysimport timeHOSTS = re.compile(r"^(\w+\s?)+$") # a group that repeats, containing \w+ that also repeatsfor arg in sys.argv[1:]:text = "a" * (int(arg) - 1) + "!" # one final character the pattern cannot matchstart = time.perf_counter()HOSTS.search(text)print(f"{len(text):>3} characters: {time.perf_counter() - start:7.3f} s", flush=True)
Four more characters multiplied the time by about sixteen (two to the power four), from under a tenth of a second to over one second. At that rate 29 characters, far below any 253-character cap, needs close to twenty seconds, and timeout killed it at five with exit status 124. A cap bounds the length of the input, not the work per character. For a pattern whose cost grows with the square of the length, 253 characters is harmless; for an exponential one, 2 to the power 29 steps is already an outage. So a cap is a budget you combine with a pattern that cannot blow up. The first choice, when you own the pattern, is the fix py-sec taught: rewrite it so every repetition must start somewhere definite, or make the group atomic:
"""Fix the pattern first, then cap the field: a cap bounds the work only once the pattern is linear."""import reimport timeCAP = 253 # a host name is never longer than thisNESTED = re.compile(r"^(\w+\s?)+$") # the shape from ps-regex: exponential on a near missFLAT = re.compile(r"^\w+(?:\s\w+)*\s?$") # the same language: every repeat starts with a spaceATOMIC = re.compile(r"^(?>\w+\s?)+$") # or keep the shape and forbid backtracking into itdef valid_hosts(value: str) -> bool:if len(value) > CAP: # the budget: refuse an over-long field, never truncatereturn Falsereturn FLAT.search(value) is not Nonefor text in ("a" * 28 + "!", "a" * 40_000 + "!"):for name, pattern in (("flat", FLAT), ("atomic", ATOMIC)):start = time.perf_counter()pattern.search(text)print(f"{name:>6} on {len(text):>5} characters: {time.perf_counter() - start:.4f} s")samples = ["web01 web02", "web01 ", "web01 web02", "web01;id"]print("nested:", [bool(NESTED.search(s)) for s in samples]) # safe here: these are short and benignprint("flat :", [bool(FLAT.search(s)) for s in samples])print("capped:", valid_hosts("web01 web02"), valid_hosts("a " * 200))
Both fixed patterns finished the 29-character attack and a 40,001-character one in milliseconds, and the rewritten pattern gives the same answers as the original on the samples. With a linear pattern, the 253-character cap now does bound the work, and valid_hosts refuses the 400-character input before matching it. When you cannot rewrite the pattern (it comes from a rule file or a user), you need a time limit on the match itself. signal.alarm does interrupt re.search on CPython 3.14, because the matcher checks for pending signals while it backtracks:
"""signal.alarm does interrupt re.search on CPython 3.14, but only in the main thread."""import reimport signalimport threadingimport timeNESTED = re.compile(r"^(\w+\s?)+$")class MatchTimeout(Exception):passdef on_alarm(signum, frame):raise MatchTimeout # raised in the main thread, wherever it is: here, inside the matchersignal.signal(signal.SIGALRM, on_alarm)start = time.monotonic()signal.alarm(1) # whole seconds only; setitimer() for finer stepstry:NESTED.search("a" * 40 + "!")except MatchTimeout:print(f"main thread: re.search interrupted after {time.monotonic() - start:.1f} s")finally:signal.alarm(0)def worker() -> None:try:signal.signal(signal.SIGALRM, on_alarm)except ValueError as err:print(f"worker thread: {err}")thread = threading.Thread(target=worker)thread.start()thread.join()
The alarm stopped the match after one second. The second line shows why this does not scale to a worker pool: only the main thread may install a signal handler, and the alarm is one per process, in whole seconds, and Unix only. For matches in worker threads, use an engine that bounds itself:
"""Two engines for patterns you cannot rewrite: regex with a timeout, and google-re2 in linear time.Run under the lab venv, which pins both."""import reimport timefrom concurrent.futures import ThreadPoolExecutorimport re2import regexLONG = "a" * 40_000 + "!"def bounded_search(pattern: str, text: str) -> str:# regex checks a wall-clock timeout as it matches, in any thread, unlike signal.alarm.try:regex.search(pattern, text, timeout=0.2)return "completed"except TimeoutError:return "TimeoutError after 0.2 s, the worker is free again"with ThreadPoolExecutor(max_workers=1) as pool:print("regex in a worker thread:", pool.submit(bounded_search, r"(a|a)+$", LONG).result())start = time.perf_counter()match = re2.search(r"^(\w+\s?)+$", LONG) # the exponential pattern, on 40,001 charactersprint(f"re2 linear: {time.perf_counter() - start:.4f} s on {len(LONG)} chars, match={bool(match)}")# Before swapping re2 in, compare answers, including the inputs where the engines differ.samples = ["web01", "web01;id", "wéb01", "web01\n"]print("re :", [bool(re.search(r"^\w+$", s)) for s in samples])print("re2 :", [bool(re2.search(r"^\w+$", s)) for s in samples])
The regex module takes a timeout= and raises TimeoutError, from any thread, because it checks the clock as it works. It happens to handle (\w+\s?)+ without blowing up, so the demo uses (a|a)+$, a shape it cannot optimise. google-re2 compiles the pattern to an automaton that runs in time linear in the input and never backtracks, so the exponential pattern finished 40,001 characters in about a millisecond. Before you swap re2 in, compare answers on the inputs where the engines differ. The last two lines show two differences: in RE2 \w is ASCII only, so wéb01 fails, and $ matches only at the very end, while Python's $ also matches before a final newline, so re accepts web01\n. RE2 also lacks backreferences and lookarounds. For validation that must not let a newline through, re.fullmatch (or \Z) is the stdlib answer.
Deserialization and validation at the trust boundary
A file can be a note you read, or a note that reads itself and does what it says. pickle is the second kind: the bytes name a callable and its arguments through __reduce__, and pickle.load runs it during the load, before you inspect a single field. The payload class below only writes a marker file so the lab stays harmless, but an attacker puts anything runnable in its place:
"""A payload object. Unpickling it runs code before you read a single field. Shown to be avoided.To keep the lab harmless the payload only writes a marker file in the working directory; an attackerwould put anything runnable here. This is why pickle must never touch data you did not produce."""from pathlib import Pathdef _run_on_load(path: str) -> None: # stands in for any command an attacker could choosePath(path).write_text("arbitrary code ran during unpickle\n")class Report:def __reduce__(self):# __reduce__ tells pickle which callable to run and with what arguments when it rebuilds# the object. pickle obeys it during load, before the result is ever inspected.return (_run_on_load, ("pickle-marker",))
make_payload.py (in the Lesson files) pickles a Report into report.pkl, and the loader is the code you find in a hundred tutorials:
"""Load a pickle the way a hundred tutorials show. Unsafe on untrusted input: shown to be avoided."""import pickleimport syswith open(sys.argv[1], "rb") as f:obj = pickle.load(f) # runs the payload's __reduce__ during the loadprint("loaded object:", obj)
Loading the "cached object" created the marker: code ran during pickle.load, and the loaded value (None) is beside the point. Never call pickle.load or pickle.loads on data you did not produce and protect yourself. For data crossing a trust boundary, take JSON, which can only build strings, numbers, lists, dicts, booleans and null, never a callable, and then validate it into a typed object. parse_finding checks shape, type, range, membership and grammar, and collects every error before it fails:
"""Validate an untrusted record into a typed object, collecting every error before failing.json.loads gives you dicts and lists, not a checked object. Check shape, type, range and grammaronce, at the boundary, and reject with a clear message instead of a traceback."""import refrom dataclasses import dataclassSEVERITIES = {"low", "medium", "high", "critical"}# Letters, digits, dots and hyphens, starting and ending with a letter or digit: no spaces, no# ';', no leading '-' that a later command would read as an option, no newline, no escapes.HOSTNAME = re.compile(r"[A-Za-z0-9](?:[A-Za-z0-9.-]{0,251}[A-Za-z0-9])?")class ValidationError(Exception):def __init__(self, errors: list[str]):self.errors = errorssuper().__init__("; ".join(errors))@dataclass(frozen=True)class Finding:target: strseverity: strscore: intdef parse_finding(raw: object) -> Finding:if not isinstance(raw, dict): # valid JSON can still be a list, string or numberraise ValidationError(["record must be a JSON object"])errors: list[str] = []target = raw.get("target")if not isinstance(target, str) or len(target) > 253 or not HOSTNAME.fullmatch(target):errors.append("target must be a host name (letters, digits, '.', '-'; at most 253)")severity = raw.get("severity")if severity not in SEVERITIES:errors.append(f"severity must be one of {sorted(SEVERITIES)}")score = raw.get("score")# isinstance(True, int) is True in Python, so reject bool explicitly; a "90" string is not an intif isinstance(score, bool) or not isinstance(score, int) or not (0 <= score <= 100):errors.append("score must be an integer in 0..100")if errors:raise ValidationError(errors)return Finding(target=target, severity=severity, score=score)
The target check is more than a length check. A string of 1 to 253 characters would accept web01;id, -oProxyCommand=id or a value with a newline, and the next lesson's SSH and subprocess code would receive it as a "validated" host name. Validated has to mean "matches the grammar the next consumer expects". HOSTNAME allows letters, digits, dots and hyphens, starting and ending with a letter or digit, and fullmatch refuses a trailing newline that $ would let through. The pattern has no nested repetition and runs after the length check, so it is linear. boundary.py is the entry point: text in, a Finding or a ValidationError out.
"""The trust boundary for one record: JSON text in, a checked Finding or a ValidationError out."""import jsonimport sysfrom schema import Finding, ValidationError, parse_findingdef load_finding(text: str | bytes) -> Finding:try:raw = json.loads(text)except (ValueError, RecursionError) as err: # JSONDecodeError is only one of the ways it failsraise ValidationError([f"not usable JSON ({err.__class__.__name__})"]) from Nonereturn parse_finding(raw)if __name__ == "__main__":try:print("valid:", load_finding(sys.argv[1]))except ValidationError as err:print("rejected:", err, file=sys.stderr)raise SystemExit(1)
The bad record was rejected with both of its problems named at once. json gives "90" as a string, and the strict isinstance(score, int) check refuses it; a lax coercer such as pydantic in its default mode would turn it into 90, which is convenient inside a program and wrong at a trust boundary. The isinstance(score, bool) guard is there because isinstance(True, int) is True in Python. Both hostile host names were refused. The last terminal is the one handlers get wrong: a JSON list, 100,000 nested brackets and a 5,000-digit number are all input json.loads can be given, and none of them fails with JSONDecodeError. The list is valid JSON that is not an object; the nesting raised RecursionError; the number raised ValueError from the 4,300-digit limit on integer conversion. A handler written as except json.JSONDecodeError crashes on the last two, which is why load_finding catches (ValueError, RecursionError) and parse_finding checks the shape first.
TOCTOU-safe file access
"Files, paths and safe writes" showed how to keep an untrusted path inside a base directory with resolve() and is_relative_to. That check has a gap in a directory an attacker can write: between the check and the open, a name can be replaced with a symlink pointing somewhere else. That gap is a time-of-check-to-time-of-use (TOCTOU) race. "Trust boundaries" earlier in this course met a link planted before the run in Bash and answered it with a temp file and mv, which replaces the name without opening what it points at; that fits when you replace a whole file. The first rule in both languages is not to write as a privileged user into a directory others can write. When you must open a name inside such a tree, open it without following links. toctou.py does it three ways:
"""Write BASE/RELPATH in a tree others can write, three ways. Usage: toctou.py MODE BASE RELPATHnaive follows every symlink on the way: the write lands wherever the links pointleaf O_NOFOLLOW on the last name only: a symlinked parent directory or '..' still escapeswalk one component at a time, never following a symlink and refusing '..': stays under BASE"""import osimport sysWRITE = os.O_WRONLY | os.O_CREAT | os.O_TRUNCdef open_beneath(base: str, relpath: str, flags: int) -> int:parts = relpath.split("/")if relpath.startswith("/") or any(part in ("", ".", "..") for part in parts):raise ValueError(f"refusing {relpath!r}: absolute, empty, '.' or '..' component")fd = os.open(base, os.O_RDONLY | os.O_DIRECTORY) # BASE comes from your config, not the attackertry:for part in parts[:-1]: # each directory: no symlink, must be a dirnext_fd = os.open(part, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW, dir_fd=fd)os.close(fd)fd = next_fdreturn os.open(parts[-1], flags | os.O_NOFOLLOW, 0o644, dir_fd=fd)finally:os.close(fd)mode, base, relpath = sys.argv[1:4]try:if mode == "walk":fd = open_beneath(base, relpath, WRITE)else:base_fd = os.open(base, os.O_RDONLY | os.O_DIRECTORY)extra = os.O_NOFOLLOW if mode == "leaf" else 0try:fd = os.open(relpath, WRITE | extra, 0o644, dir_fd=base_fd)finally:os.close(base_fd)with os.fdopen(fd, "wb") as out:out.write(b"written by the ingest job\n")print(f"{mode}: wrote {relpath}")except (OSError, ValueError) as err:reason = os.strerror(err.errno) if isinstance(err, OSError) else str(err)print(f"{mode}: refused {relpath} ({reason})")raise SystemExit(1)
The naive open followed the planted link and overwrote secret.txt. With O_NOFOLLOW, the open refused the link with ELOOP ("Too many levels of symbolic links") and the secret stayed. That flag is often presented as the whole fix. It is not, because O_NOFOLLOW applies only to the last component of the path:
A symlinked parent directory (spool/sub pointing at elsewhere) and a plain .. both escaped the spool, and leaf mode wrote through them. The directory descriptor pins only the one directory it was opened on; every intermediate name in the path is still resolved and followed. open_beneath closes both holes. It refuses .., ., empty and absolute paths outright, then opens one component at a time relative to the previous directory's descriptor with O_DIRECTORY | O_NOFOLLOW, so no symlink is followed anywhere on the way:
The symlinked parent was refused with ENOTDIR ("Not a directory": with O_NOFOLLOW the kernel looks at the link itself, and a link is not a directory), .. was refused before any system call, the planted leaf link failed with ELOOP, and the legitimate in/report was written. elsewhere stayed empty and the secret was untouched. The base directory itself is opened normally, because it comes from your configuration, not from the attacker. Linux has a kernel-level version of this walk, openat2(2) with RESOLVE_BENEATH and RESOLVE_NO_SYMLINKS, but Python 3.14's os module does not expose it, so the component walk is the portable stdlib form. Add O_EXCL when the job should only create new files.
Log injection: a field is not a line
Everything parsed from untrusted input is still untrusted when it reaches your logs. If you paste a field value into a text log line, a newline inside it forges a second entry and an ANSI escape can rewrite the reader's terminal. A structured logger serialises the value, so control characters survive as data and cannot break out of the field:
"""A field from an untrusted record ends up in a log line. Two ways to write it, one safe.If you paste the value into a text line, a newline in it forges a second log entry and an ANSI escapecan rewrite your terminal. A structured logger serialises the value, so control characters survive asdata (\\n, \\u001b) and can never break out of the field."""import jsonimport loggingimport sysfrom datetime import datetime, timezoneclass JsonFormatter(logging.Formatter):def format(self, record: logging.LogRecord) -> str:payload = {"ts": datetime.fromtimestamp(record.created, timezone.utc).isoformat(),"level": record.levelname.lower(),"msg": record.getMessage(),}fields = getattr(record, "fields", None)if fields:# Nested, not merged: an untrusted "level" or "msg" key cannot replace the real one.payload["fields"] = fieldsreturn json.dumps(payload, ensure_ascii=True) # escapes newlines and control characters# A user name crafted to inject a second, fake "authentication failed" line and hide it with ANSI.hostile = "alice\n2026-01-01 00:00:00 CRITICAL auth failed for admin\x1b[2K"naive = logging.getLogger("naive")naive.addHandler(logging.StreamHandler(sys.stderr))naive.setLevel(logging.INFO)print("--- naive: value pasted into the message text", file=sys.stderr, flush=True)naive.info("login user=%s", hostile) # the newline in `hostile` becomes a real new linestructured = logging.getLogger("structured")h = logging.StreamHandler(sys.stderr)h.setFormatter(JsonFormatter())structured.addHandler(h)structured.setLevel(logging.INFO)structured.propagate = Falseprint("--- structured: value carried as a field", file=sys.stderr, flush=True)structured.info("login", extra={"fields": {"user": hostile, "level": "debug"}}) # "level" came from the record
cat -v makes the control bytes visible. The naive line broke into two: a forged CRITICAL auth failed for admin line that a log search would read as real, followed by ^[[2K (an ANSI erase). The structured line carried the value inside one JSON field, with the newline escaped to \n and the escape to \u001b. The record also carried its own level key, set to debug. Because the formatter nests the fields under "fields" instead of merging them into the top level, the real "level": "info" survived; merged, an attacker's level, ts or msg would replace the real one inside a well-formed line. "Observable jobs" later in this course builds its JSON logs on the same formatter idea.
Try this
Put the boundary defences together into ingest.py: read NDJSON from standard input with read_records under Budget(), turn each line into a Finding with load_finding, print each valid finding, count the rejects with their reasons, and exit 0 only when every record is valid. The lab's solution on findings.ndjson printed three findings and 3 valid, 0 rejected with exit status 0. On findings-mixed.ndjson (in the Lesson files; it holds a bad severity, an out-of-range score, a string score, a missing target, web01;id and a JSON list) it printed 1 valid, 6 rejected with a reason per bad record and exit status 1. On the 1.1 MB oversize.txt it printed aborted: input exceeds 1000000 bytes and exited 1 without parsing anything. Then explain why the budget abort and a validation failure both exit 1 but log different messages.
Takeaway
Bound the size, time, count and length of untrusted input before you hold it; treat a length cap as a budget that only works with a pattern that cannot backtrack exponentially, or with a match timeout; validate JSON into typed objects whose fields match the grammar the next consumer expects; and when you open a name in a tree others can write, walk it one component at a time without following links.
re.compile(r"^(\w+\s?)+$") on every incoming record and has no length limit. Traffic is fine for weeks, then one request makes a worker pin a CPU for minutes. What is the cause and the cheapest lasting fix?pickle.load. A reviewer proposes keeping pickle but adding isinstance(obj, Job) right after the load to reject anything unexpected. Why does that not make it safe?uploads/2026/report.txt under a base directory that an attacker can also write. It opens the path relative to a descriptor of the base directory, with O_NOFOLLOW. Files outside the base directory still get overwritten. What closes the hole?/base-evil as inside /base, and no string test closes the swap-after-check race.