Handling hostile input at scale

Size and time budgets, ReDoS, deserialization choices, validation and TOCTOU.

Expert45 min · lesson 11 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (16 files, 6 KB): scr-py-secure.tar.gz. Unpack it with tar -xzf scr-py-secure.tar.gz, which creates scr-py-secure/. SHA-256: bb07476208639e87f5d8b64fefcf7a7f05f0b1ccf7bf8bfd1370686a6a9aaf87

Automation that reads logs, uploads, API responses and files handles input someone else controls, and at scale one hostile record can take the whole job down. This lesson adds the defences that matter when the volume is large and the source is not trusted: budgets that refuse input too big or too slow to accept, regular expressions that cannot be made to hang, deserialization that cannot run code, validation that turns a raw value into a checked object, file access that a swapped-in symlink cannot redirect, and logging that a crafted field cannot forge. The tools are standard library except two matching engines, pinned in a venv; the behaviour is the same on Ubuntu's Python 3.14.4 and upstream 3.14.7.

This lesson builds on three py-sec lessons: "JSON, CSV, YAML and TOML" owns json, yaml.safe_load and CSV injection; "Regular expressions for logs and tool output" owns the ReDoS shape and its fixes; "Files, paths and safe writes" owns pathlib, atomic writes and is_relative_to containment. Here each idea is pushed to production scale. (This lesson also absorbs the retired "Parsing at scale" lesson, whose URL now redirects here.) Unpack the Lesson files in your home directory: they create scr-py-secure/ with every script and the two NDJSON fixtures, and the commands below run inside it. The terminals show the lab's deploy account; yours shows your own.

Budgets: refuse what is too big or too slow

The first defence at scale is refusing to hold too much, or to wait too long. A decompression bomb, a file with a billion lines, one line that never ends and a sender that trickles one byte at a time all end in an outage unless a budget stops them first. limits.py defines the budgets and a read that honours two of them, total bytes and total time:

limits.py
"""Budgets for untrusted input, and a read that honours two of them: total bytes and total time."""
import os
import select
import time
from dataclasses import dataclass
@dataclass(frozen=True)
class Budget:
max_bytes: int = 1_000_000 # never read more than this from the source at all
max_seconds: float = 5.0 # the whole input must arrive within this time
max_lines: int = 10_000 # a record count a single run should never exceed
max_line_bytes: int = 8_192 # one record is small; a huge line is an attack or a bug
class BudgetExceeded(Exception):
"""The input is larger or slower than we agreed to process: reject it, do not truncate it."""
def read_capped(fd: int, budget: Budget) -> bytes:
# Chunk by chunk: memory stays bounded even if the source is endless, and every wait draws
# from one deadline, so a sender that trickles a byte at a time cannot hold the worker.
give_up_at = time.monotonic() + budget.max_seconds
chunks, size = [], 0
while True:
remaining = give_up_at - time.monotonic()
if remaining <= 0 or not select.select([fd], [], [], remaining)[0]:
raise BudgetExceeded(f"input still incomplete after {budget.max_seconds} s")
chunk = os.read(fd, 65_536)
if not chunk: # end of input
return b"".join(chunks)
size += len(chunk)
if size > budget.max_bytes:
raise BudgetExceeded(f"input exceeds {budget.max_bytes} bytes")
chunks.append(chunk)

read_capped reads in chunks and stops as soon as the total passes max_bytes, so an oversized input is never pulled into memory in full. The time budget works the same way: give_up_at is fixed once, and each select.select call waits only for the time that is left, so the deadline covers the whole input rather than each read. A plain stream.read() on a pipe or socket has no such limit; it waits until the bytes or the end arrive. budgets.py splits the capped bytes into records and applies the other two budgets:

budgets.py
"""Split capped input into records under two more budgets: line count and single-line length."""
import sys
from limits import Budget, BudgetExceeded, read_capped
def read_records(fd: int, budget: Budget):
lines = read_capped(fd, budget).split(b"\n")
if lines[-1] == b"": # a final newline ends the last line, it adds none
lines.pop()
if len(lines) > budget.max_lines:
raise BudgetExceeded(f"more than {budget.max_lines} lines")
for n, raw in enumerate(lines, 1):
if len(raw) > budget.max_line_bytes:
raise BudgetExceeded(f"line {n} is {len(raw)} bytes, over {budget.max_line_bytes}")
if raw:
yield raw
if __name__ == "__main__":
try:
count = sum(1 for _ in read_records(sys.stdin.fileno(), Budget()))
except BudgetExceeded as err:
print(f"rejected: {err}", file=sys.stderr)
raise SystemExit(1)
print(f"accepted {count} records")
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 budgets.py < findings.ndjson
accepted 3 records
$ head -c 1100000 /dev/zero | tr "\0" a > oversize.txt python3 budgets.py < oversize.txt; echo "exit status $?"
rejected: input exceeds 1000000 bytes exit status 1
$ yes x | head -n 10000 > lines-10000.txt yes x | head -n 10001 > lines-10001.txt python3 budgets.py < lines-10000.txt python3 budgets.py < lines-10001.txt; echo "exit status $?"
accepted 10000 records rejected: more than 10000 lines exit status 1
$ python3 -c "print(\"a\" * 9000)" > longline.txt python3 budgets.py < longline.txt; echo "exit status $?"
rejected: line 1 is 9000 bytes, over 8192 exit status 1
$ start=$SECONDS for i in $(seq 20); do printf x; sleep 0.5; done | python3 budgets.py echo "exit status $? after $((SECONDS - start)) s"
rejected: input still incomplete after 5.0 s exit status 1 after 5 s

Each rejection names the budget it hit and exits 1, so the caller learns the input was refused rather than silently truncated. The line-count case shows the boundary: exactly 10,000 lines is accepted and 10,001 is not. That needs the lines.pop() line, because splitting newline-terminated text on \n leaves one empty item after the last newline. The slow sender needed ten seconds to deliver twenty bytes; the reader gave up after five, and the sender's next write went to a closed pipe. A budget is a policy decision you write down: size it to the largest legitimate input plus headroom.

ReDoS at scale: a length cap is a budget, not a fix

py-sec showed that a quantifier nested inside a repeated group, like ^(\w+\s?)+$, backtracks exponentially on an almost-matching string. At fleet scale that pattern runs on every record, so one crafted value stalls a worker. A tempting answer is to cap the field at its real maximum (a host name is at most 253 characters) and let the cap keep long attack strings away from the engine. Watch the cap fail:

redos.py
"""Time the built-in re engine on a nested-quantifier pattern against an almost-matching string.
The pattern validates space-separated host names. Each argument is the total length of the input."""
import re
import sys
import time
HOSTS = re.compile(r"^(\w+\s?)+$") # a group that repeats, containing \w+ that also repeats
for arg in sys.argv[1:]:
text = "a" * (int(arg) - 1) + "!" # one final character the pattern cannot match
start = time.perf_counter()
HOSTS.search(text)
print(f"{len(text):>3} characters: {time.perf_counter() - start:7.3f} s", flush=True)
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ timeout 5 python3 redos.py 21 25 29; echo "exit status $?"
21 characters: 0.072 s 25 characters: 1.165 s exit status 124

Four more characters multiplied the time by about sixteen (two to the power four), from under a tenth of a second to over one second. At that rate 29 characters, far below any 253-character cap, needs close to twenty seconds, and timeout killed it at five with exit status 124. A cap bounds the length of the input, not the work per character. For a pattern whose cost grows with the square of the length, 253 characters is harmless; for an exponential one, 2 to the power 29 steps is already an outage. So a cap is a budget you combine with a pattern that cannot blow up. The first choice, when you own the pattern, is the fix py-sec taught: rewrite it so every repetition must start somewhere definite, or make the group atomic:

flat.py
"""Fix the pattern first, then cap the field: a cap bounds the work only once the pattern is linear."""
import re
import time
CAP = 253 # a host name is never longer than this
NESTED = re.compile(r"^(\w+\s?)+$") # the shape from ps-regex: exponential on a near miss
FLAT = re.compile(r"^\w+(?:\s\w+)*\s?$") # the same language: every repeat starts with a space
ATOMIC = re.compile(r"^(?>\w+\s?)+$") # or keep the shape and forbid backtracking into it
def valid_hosts(value: str) -> bool:
if len(value) > CAP: # the budget: refuse an over-long field, never truncate
return False
return FLAT.search(value) is not None
for text in ("a" * 28 + "!", "a" * 40_000 + "!"):
for name, pattern in (("flat", FLAT), ("atomic", ATOMIC)):
start = time.perf_counter()
pattern.search(text)
print(f"{name:>6} on {len(text):>5} characters: {time.perf_counter() - start:.4f} s")
samples = ["web01 web02", "web01 ", "web01 web02", "web01;id"]
print("nested:", [bool(NESTED.search(s)) for s in samples]) # safe here: these are short and benign
print("flat :", [bool(FLAT.search(s)) for s in samples])
print("capped:", valid_hosts("web01 web02"), valid_hosts("a " * 200))
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 flat.py
flat on 29 characters: 0.0000 s atomic on 29 characters: 0.0000 s flat on 40001 characters: 0.0022 s atomic on 40001 characters: 0.0001 s nested: [True, True, False, False] flat : [True, True, False, False] capped: True False

Both fixed patterns finished the 29-character attack and a 40,001-character one in milliseconds, and the rewritten pattern gives the same answers as the original on the samples. With a linear pattern, the 253-character cap now does bound the work, and valid_hosts refuses the 400-character input before matching it. When you cannot rewrite the pattern (it comes from a rule file or a user), you need a time limit on the match itself. signal.alarm does interrupt re.search on CPython 3.14, because the matcher checks for pending signals while it backtracks:

alarm.py
"""signal.alarm does interrupt re.search on CPython 3.14, but only in the main thread."""
import re
import signal
import threading
import time
NESTED = re.compile(r"^(\w+\s?)+$")
class MatchTimeout(Exception):
pass
def on_alarm(signum, frame):
raise MatchTimeout # raised in the main thread, wherever it is: here, inside the matcher
signal.signal(signal.SIGALRM, on_alarm)
start = time.monotonic()
signal.alarm(1) # whole seconds only; setitimer() for finer steps
try:
NESTED.search("a" * 40 + "!")
except MatchTimeout:
print(f"main thread: re.search interrupted after {time.monotonic() - start:.1f} s")
finally:
signal.alarm(0)
def worker() -> None:
try:
signal.signal(signal.SIGALRM, on_alarm)
except ValueError as err:
print(f"worker thread: {err}")
thread = threading.Thread(target=worker)
thread.start()
thread.join()
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 alarm.py
main thread: re.search interrupted after 1.0 s worker thread: signal only works in main thread of the main interpreter

The alarm stopped the match after one second. The second line shows why this does not scale to a worker pool: only the main thread may install a signal handler, and the alarm is one per process, in whole seconds, and Unix only. For matches in worker threads, use an engine that bounds itself:

engines.py
"""Two engines for patterns you cannot rewrite: regex with a timeout, and google-re2 in linear time.
Run under the lab venv, which pins both."""
import re
import time
from concurrent.futures import ThreadPoolExecutor
import re2
import regex
LONG = "a" * 40_000 + "!"
def bounded_search(pattern: str, text: str) -> str:
# regex checks a wall-clock timeout as it matches, in any thread, unlike signal.alarm.
try:
regex.search(pattern, text, timeout=0.2)
return "completed"
except TimeoutError:
return "TimeoutError after 0.2 s, the worker is free again"
with ThreadPoolExecutor(max_workers=1) as pool:
print("regex in a worker thread:", pool.submit(bounded_search, r"(a|a)+$", LONG).result())
start = time.perf_counter()
match = re2.search(r"^(\w+\s?)+$", LONG) # the exponential pattern, on 40,001 characters
print(f"re2 linear: {time.perf_counter() - start:.4f} s on {len(LONG)} chars, match={bool(match)}")
# Before swapping re2 in, compare answers, including the inputs where the engines differ.
samples = ["web01", "web01;id", "wéb01", "web01\n"]
print("re :", [bool(re.search(r"^\w+$", s)) for s in samples])
print("re2 :", [bool(re2.search(r"^\w+$", s)) for s in samples])
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 -m venv .venv .venv/bin/python -m pip install -q --disable-pip-version-check --timeout 60 -r requirements.txt .venv/bin/python -c "import regex, re2; print(\"regex\", regex.__version__); print(\"re2 ok\")"
regex 2026.9.10 re2 ok
$ .venv/bin/python engines.py
regex in a worker thread: TimeoutError after 0.2 s, the worker is free again re2 linear: 0.0012 s on 40001 chars, match=False re : [True, False, True, True] re2 : [True, False, False, False]

The regex module takes a timeout= and raises TimeoutError, from any thread, because it checks the clock as it works. It happens to handle (\w+\s?)+ without blowing up, so the demo uses (a|a)+$, a shape it cannot optimise. google-re2 compiles the pattern to an automaton that runs in time linear in the input and never backtracks, so the exponential pattern finished 40,001 characters in about a millisecond. Before you swap re2 in, compare answers on the inputs where the engines differ. The last two lines show two differences: in RE2 \w is ASCII only, so wéb01 fails, and $ matches only at the very end, while Python's $ also matches before a final newline, so re accepts web01\n. RE2 also lacks backreferences and lookarounds. For validation that must not let a newline through, re.fullmatch (or \Z) is the stdlib answer.

Deserialization and validation at the trust boundary

A file can be a note you read, or a note that reads itself and does what it says. pickle is the second kind: the bytes name a callable and its arguments through __reduce__, and pickle.load runs it during the load, before you inspect a single field. The payload class below only writes a marker file so the lab stays harmless, but an attacker puts anything runnable in its place:

evil.py (the payload; unsafe by design)
"""A payload object. Unpickling it runs code before you read a single field. Shown to be avoided.
To keep the lab harmless the payload only writes a marker file in the working directory; an attacker
would put anything runnable here. This is why pickle must never touch data you did not produce."""
from pathlib import Path
def _run_on_load(path: str) -> None: # stands in for any command an attacker could choose
Path(path).write_text("arbitrary code ran during unpickle\n")
class Report:
def __reduce__(self):
# __reduce__ tells pickle which callable to run and with what arguments when it rebuilds
# the object. pickle obeys it during load, before the result is ever inspected.
return (_run_on_load, ("pickle-marker",))

make_payload.py (in the Lesson files) pickles a Report into report.pkl, and the loader is the code you find in a hundred tutorials:

load_pickle.py (unsafe on untrusted input)
"""Load a pickle the way a hundred tutorials show. Unsafe on untrusted input: shown to be avoided."""
import pickle
import sys
with open(sys.argv[1], "rb") as f:
obj = pickle.load(f) # runs the payload's __reduce__ during the load
print("loaded object:", obj)
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 make_payload.py
wrote report.pkl (looks like a harmless cached object)
$ rm -f pickle-marker; python3 load_pickle.py report.pkl; echo "--- marker:"; cat pickle-marker
loaded object: None --- marker: arbitrary code ran during unpickle

Loading the "cached object" created the marker: code ran during pickle.load, and the loaded value (None) is beside the point. Never call pickle.load or pickle.loads on data you did not produce and protect yourself. For data crossing a trust boundary, take JSON, which can only build strings, numbers, lists, dicts, booleans and null, never a callable, and then validate it into a typed object. parse_finding checks shape, type, range, membership and grammar, and collects every error before it fails:

schema.py
"""Validate an untrusted record into a typed object, collecting every error before failing.
json.loads gives you dicts and lists, not a checked object. Check shape, type, range and grammar
once, at the boundary, and reject with a clear message instead of a traceback."""
import re
from dataclasses import dataclass
SEVERITIES = {"low", "medium", "high", "critical"}
# Letters, digits, dots and hyphens, starting and ending with a letter or digit: no spaces, no
# ';', no leading '-' that a later command would read as an option, no newline, no escapes.
HOSTNAME = re.compile(r"[A-Za-z0-9](?:[A-Za-z0-9.-]{0,251}[A-Za-z0-9])?")
class ValidationError(Exception):
def __init__(self, errors: list[str]):
self.errors = errors
super().__init__("; ".join(errors))
@dataclass(frozen=True)
class Finding:
target: str
severity: str
score: int
def parse_finding(raw: object) -> Finding:
if not isinstance(raw, dict): # valid JSON can still be a list, string or number
raise ValidationError(["record must be a JSON object"])
errors: list[str] = []
target = raw.get("target")
if not isinstance(target, str) or len(target) > 253 or not HOSTNAME.fullmatch(target):
errors.append("target must be a host name (letters, digits, '.', '-'; at most 253)")
severity = raw.get("severity")
if severity not in SEVERITIES:
errors.append(f"severity must be one of {sorted(SEVERITIES)}")
score = raw.get("score")
# isinstance(True, int) is True in Python, so reject bool explicitly; a "90" string is not an int
if isinstance(score, bool) or not isinstance(score, int) or not (0 <= score <= 100):
errors.append("score must be an integer in 0..100")
if errors:
raise ValidationError(errors)
return Finding(target=target, severity=severity, score=score)

The target check is more than a length check. A string of 1 to 253 characters would accept web01;id, -oProxyCommand=id or a value with a newline, and the next lesson's SSH and subprocess code would receive it as a "validated" host name. Validated has to mean "matches the grammar the next consumer expects". HOSTNAME allows letters, digits, dots and hyphens, starting and ending with a letter or digit, and fullmatch refuses a trailing newline that $ would let through. The pattern has no nested repetition and runs after the length check, so it is linear. boundary.py is the entry point: text in, a Finding or a ValidationError out.

boundary.py
"""The trust boundary for one record: JSON text in, a checked Finding or a ValidationError out."""
import json
import sys
from schema import Finding, ValidationError, parse_finding
def load_finding(text: str | bytes) -> Finding:
try:
raw = json.loads(text)
except (ValueError, RecursionError) as err: # JSONDecodeError is only one of the ways it fails
raise ValidationError([f"not usable JSON ({err.__class__.__name__})"]) from None
return parse_finding(raw)
if __name__ == "__main__":
try:
print("valid:", load_finding(sys.argv[1]))
except ValidationError as err:
print("rejected:", err, file=sys.stderr)
raise SystemExit(1)
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 boundary.py '{"target": "web01", "severity": "high", "score": 90}'
valid: Finding(target='web01', severity='high', score=90)
$ python3 boundary.py '{"target": "web01", "severity": "urgent", "score": 900}'; echo "exit status $?"
rejected: severity must be one of ['critical', 'high', 'low', 'medium']; score must be an integer in 0..100 exit status 1
$ python3 boundary.py '{"target": "web01", "severity": "low", "score": "90"}'; echo "exit status $?"
rejected: score must be an integer in 0..100 exit status 1
$ python3 boundary.py '{"target": "web01;id", "severity": "low", "score": 5}' python3 boundary.py '{"target": "-oProxyCommand=id", "severity": "low", "score": 5}'
rejected: target must be a host name (letters, digits, '.', '-'; at most 253) rejected: target must be a host name (letters, digits, '.', '-'; at most 253)
$ python3 boundary.py '[1]' python3 boundary.py "$(head -c 100000 /dev/zero | tr '\0' '[')" python3 boundary.py "$(head -c 5000 /dev/zero | tr '\0' 9)"; echo "exit status $?"
rejected: record must be a JSON object rejected: not usable JSON (RecursionError) rejected: not usable JSON (ValueError) exit status 1

The bad record was rejected with both of its problems named at once. json gives "90" as a string, and the strict isinstance(score, int) check refuses it; a lax coercer such as pydantic in its default mode would turn it into 90, which is convenient inside a program and wrong at a trust boundary. The isinstance(score, bool) guard is there because isinstance(True, int) is True in Python. Both hostile host names were refused. The last terminal is the one handlers get wrong: a JSON list, 100,000 nested brackets and a 5,000-digit number are all input json.loads can be given, and none of them fails with JSONDecodeError. The list is valid JSON that is not an object; the nesting raised RecursionError; the number raised ValueError from the 4,300-digit limit on integer conversion. A handler written as except json.JSONDecodeError crashes on the last two, which is why load_finding catches (ValueError, RecursionError) and parse_finding checks the shape first.

TOCTOU-safe file access

"Files, paths and safe writes" showed how to keep an untrusted path inside a base directory with resolve() and is_relative_to. That check has a gap in a directory an attacker can write: between the check and the open, a name can be replaced with a symlink pointing somewhere else. That gap is a time-of-check-to-time-of-use (TOCTOU) race. "Trust boundaries" earlier in this course met a link planted before the run in Bash and answered it with a temp file and mv, which replaces the name without opening what it points at; that fits when you replace a whole file. The first rule in both languages is not to write as a privileged user into a directory others can write. When you must open a name inside such a tree, open it without following links. toctou.py does it three ways:

toctou.py
"""Write BASE/RELPATH in a tree others can write, three ways. Usage: toctou.py MODE BASE RELPATH
naive follows every symlink on the way: the write lands wherever the links point
leaf O_NOFOLLOW on the last name only: a symlinked parent directory or '..' still escapes
walk one component at a time, never following a symlink and refusing '..': stays under BASE"""
import os
import sys
WRITE = os.O_WRONLY | os.O_CREAT | os.O_TRUNC
def open_beneath(base: str, relpath: str, flags: int) -> int:
parts = relpath.split("/")
if relpath.startswith("/") or any(part in ("", ".", "..") for part in parts):
raise ValueError(f"refusing {relpath!r}: absolute, empty, '.' or '..' component")
fd = os.open(base, os.O_RDONLY | os.O_DIRECTORY) # BASE comes from your config, not the attacker
try:
for part in parts[:-1]: # each directory: no symlink, must be a dir
next_fd = os.open(part, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW, dir_fd=fd)
os.close(fd)
fd = next_fd
return os.open(parts[-1], flags | os.O_NOFOLLOW, 0o644, dir_fd=fd)
finally:
os.close(fd)
mode, base, relpath = sys.argv[1:4]
try:
if mode == "walk":
fd = open_beneath(base, relpath, WRITE)
else:
base_fd = os.open(base, os.O_RDONLY | os.O_DIRECTORY)
extra = os.O_NOFOLLOW if mode == "leaf" else 0
try:
fd = os.open(relpath, WRITE | extra, 0o644, dir_fd=base_fd)
finally:
os.close(base_fd)
with os.fdopen(fd, "wb") as out:
out.write(b"written by the ingest job\n")
print(f"{mode}: wrote {relpath}")
except (OSError, ValueError) as err:
reason = os.strerror(err.errno) if isinstance(err, OSError) else str(err)
print(f"{mode}: refused {relpath} ({reason})")
raise SystemExit(1)
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ printf SECRET > secret.txt; mkdir -p spool/in elsewhere ln -sfn "$PWD/secret.txt" spool/report; ls -l spool/report
lrwxrwxrwx 1 deploy deploy 37 Sep 28 18:41 spool/report -> /home/deploy/scr-py-secure/secret.txt
$ python3 toctou.py naive spool report; cat secret.txt printf SECRET > secret.txt python3 toctou.py leaf spool report; cat secret.txt; echo
naive: wrote report written by the ingest job leaf: refused report (Too many levels of symbolic links) SECRET

The naive open followed the planted link and overwrote secret.txt. With O_NOFOLLOW, the open refused the link with ELOOP ("Too many levels of symbolic links") and the secret stayed. That flag is often presented as the whole fix. It is not, because O_NOFOLLOW applies only to the last component of the path:

deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ ln -sfn "$PWD/elsewhere" spool/sub python3 toctou.py leaf spool sub/report python3 toctou.py leaf spool ../secret.txt ls elsewhere; cat secret.txt
leaf: wrote sub/report leaf: wrote ../secret.txt report written by the ingest job

A symlinked parent directory (spool/sub pointing at elsewhere) and a plain .. both escaped the spool, and leaf mode wrote through them. The directory descriptor pins only the one directory it was opened on; every intermediate name in the path is still resolved and followed. open_beneath closes both holes. It refuses .., ., empty and absolute paths outright, then opens one component at a time relative to the previous directory's descriptor with O_DIRECTORY | O_NOFOLLOW, so no symlink is followed anywhere on the way:

deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ printf SECRET > secret.txt; rm -f elsewhere/report python3 toctou.py walk spool sub/report python3 toctou.py walk spool ../secret.txt python3 toctou.py walk spool report python3 toctou.py walk spool in/report ls elsewhere spool/in; cat secret.txt; echo
walk: refused sub/report (Not a directory) walk: refused ../secret.txt (refusing '../secret.txt': absolute, empty, '.' or '..' component) walk: refused report (Too many levels of symbolic links) walk: wrote in/report elsewhere: spool/in: report SECRET

The symlinked parent was refused with ENOTDIR ("Not a directory": with O_NOFOLLOW the kernel looks at the link itself, and a link is not a directory), .. was refused before any system call, the planted leaf link failed with ELOOP, and the legitimate in/report was written. elsewhere stayed empty and the secret was untouched. The base directory itself is opened normally, because it comes from your configuration, not from the attacker. Linux has a kernel-level version of this walk, openat2(2) with RESOLVE_BENEATH and RESOLVE_NO_SYMLINKS, but Python 3.14's os module does not expose it, so the component walk is the portable stdlib form. Add O_EXCL when the job should only create new files.

Log injection: a field is not a line

Everything parsed from untrusted input is still untrusted when it reaches your logs. If you paste a field value into a text log line, a newline inside it forges a second entry and an ANSI escape can rewrite the reader's terminal. A structured logger serialises the value, so control characters survive as data and cannot break out of the field:

logsafe.py
"""A field from an untrusted record ends up in a log line. Two ways to write it, one safe.
If you paste the value into a text line, a newline in it forges a second log entry and an ANSI escape
can rewrite your terminal. A structured logger serialises the value, so control characters survive as
data (\\n, \\u001b) and can never break out of the field."""
import json
import logging
import sys
from datetime import datetime, timezone
class JsonFormatter(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
payload = {
"ts": datetime.fromtimestamp(record.created, timezone.utc).isoformat(),
"level": record.levelname.lower(),
"msg": record.getMessage(),
}
fields = getattr(record, "fields", None)
if fields:
# Nested, not merged: an untrusted "level" or "msg" key cannot replace the real one.
payload["fields"] = fields
return json.dumps(payload, ensure_ascii=True) # escapes newlines and control characters
# A user name crafted to inject a second, fake "authentication failed" line and hide it with ANSI.
hostile = "alice\n2026-01-01 00:00:00 CRITICAL auth failed for admin\x1b[2K"
naive = logging.getLogger("naive")
naive.addHandler(logging.StreamHandler(sys.stderr))
naive.setLevel(logging.INFO)
print("--- naive: value pasted into the message text", file=sys.stderr, flush=True)
naive.info("login user=%s", hostile) # the newline in `hostile` becomes a real new line
structured = logging.getLogger("structured")
h = logging.StreamHandler(sys.stderr)
h.setFormatter(JsonFormatter())
structured.addHandler(h)
structured.setLevel(logging.INFO)
structured.propagate = False
print("--- structured: value carried as a field", file=sys.stderr, flush=True)
structured.info("login", extra={"fields": {"user": hostile, "level": "debug"}}) # "level" came from the record
deploy@web01:~/scr-py-secure · Ubuntu 26.04 LTS
$ python3 logsafe.py 2>&1 | cat -v
--- naive: value pasted into the message text login user=alice 2026-01-01 00:00:00 CRITICAL auth failed for admin^[[2K --- structured: value carried as a field {"ts": "2026-09-28T18:41:23.908522+00:00", "level": "info", "msg": "login", "fields": {"user": "alice\n2026-01-01 00:00:00 CRITICAL auth failed for admin\u001b[2K", "level": "debug"}}

cat -v makes the control bytes visible. The naive line broke into two: a forged CRITICAL auth failed for admin line that a log search would read as real, followed by ^[[2K (an ANSI erase). The structured line carried the value inside one JSON field, with the newline escaped to \n and the escape to \u001b. The record also carried its own level key, set to debug. Because the formatter nests the fields under "fields" instead of merging them into the top level, the real "level": "info" survived; merged, an attacker's level, ts or msg would replace the real one inside a well-formed line. "Observable jobs" later in this course builds its JSON logs on the same formatter idea.

Where one hostile record becomes an outage
Volume
size / time / line budgets
refuse, never truncate
one deadline for the read
a trickle cannot hold the worker
Matching
rewrite or atomic group
then the length cap bounds the work
regex timeout= / re2
for patterns you cannot rewrite
Meaning
never pickle.load untrusted
it runs code on load
json + shape + grammar
catch ValueError, RecursionError
Files and logs
walk with O_NOFOLLOW
every component, no ..
nested structured fields
a field cannot forge a line
Each column is a way one hostile record turns into an outage, and the defence that stops it.

Try this

Put the boundary defences together into ingest.py: read NDJSON from standard input with read_records under Budget(), turn each line into a Finding with load_finding, print each valid finding, count the rejects with their reasons, and exit 0 only when every record is valid. The lab's solution on findings.ndjson printed three findings and 3 valid, 0 rejected with exit status 0. On findings-mixed.ndjson (in the Lesson files; it holds a bad severity, an out-of-range score, a string score, a missing target, web01;id and a JSON list) it printed 1 valid, 6 rejected with a reason per bad record and exit status 1. On the 1.1 MB oversize.txt it printed aborted: input exceeds 1000000 bytes and exited 1 without parsing anything. Then explain why the budget abort and a validation failure both exit 1 but log different messages.

Takeaway

Bound the size, time, count and length of untrusted input before you hold it; treat a length cap as a budget that only works with a pattern that cannot backtrack exponentially, or with a match timeout; validate JSON into typed objects whose fields match the grammar the next consumer expects; and when you open a name in a tree others can write, walk it one component at a time without following links.

Quick check
01A parser validates a host-list field with re.compile(r"^(\w+\s?)+$") on every incoming record and has no length limit. Traffic is fine for weeks, then one request makes a worker pin a CPU for minutes. What is the cause and the cheapest lasting fix?
Incorrect — A cap alone does not help here: the lab's 29-character near-match already ran past five seconds, because the cost is exponential in the length, not proportional to it.
Incorrect — Unicode matching is linear; the blow-up is the nested quantifier, and re.ASCII leaves it in place.
Correct — The rewrite removes the exponential backtracking, and once the pattern is linear the length cap bounds the work per record.
Incorrect — A read timeout bounds waiting on the network, not the CPU the regex burns after the full request has arrived.
02A service loads job definitions from files with pickle.load. A reviewer proposes keeping pickle but adding isinstance(obj, Job) right after the load to reject anything unexpected. Why does that not make it safe?
Correct — Unpickling attacker bytes executes code as it rebuilds the object, so no post-load check can help; take JSON and validate instead.
Incorrect — isinstance reports the real type fine; the problem is that code already ran during the load, not that the check is unreliable.
Incorrect — pickle does no such conversion; the object's type is not the issue, the code executed during load is.
Incorrect — No threads or races are involved; the payload runs synchronously inside pickle.load before the next line runs.
03A job writes reports to paths like uploads/2026/report.txt under a base directory that an attacker can also write. It opens the path relative to a descriptor of the base directory, with O_NOFOLLOW. Files outside the base directory still get overwritten. What closes the hole?
Incorrect — The check and the open are still two moments: a directory in the path can be swapped for a symlink after the check passes.
Incorrect — startswith is weaker, not stricter: it treats /base-evil as inside /base, and no string test closes the swap-after-check race.
Incorrect — O_DIRECTORY only requires the final name to be a directory; it does nothing about links or .. earlier in the path.
Correct — O_NOFOLLOW covers only the last component, so a symlinked parent or .. still escapes; walking component by component leaves no name that is followed.

Related