Capstone: an IOC sweep tool
Config, secrets, hashing, an HTTPS API, subprocess, output, tests and install.
tar -xzf ps-capstone.tar.gz, which creates ps-capstone/. SHA-256: 685f5013ab38e5b2f91c23476b1b1a7652b590a98c3a2adacba5a31d8c6363bdThis capstone builds one tool you could schedule on a real server and trust with evidence. ioc-sweep hashes every regular file under a directory such as an upload share, looks the hashes up in a threat-intelligence API, runs a scanner over each file, reports findings and can quarantine them. Every part comes from an earlier lesson. You will read it module by module with its tests, run it against a local HTTPS API that fails in every way the tool claims to handle, break a test, install the tool from a wheel and schedule it; the Try this section is where you write code. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/ps-capstone: they hold every file, including one this page describes but does not list, lab/mock_api.py.
What the tool promises
The command line is ioc-sweep scan DIR --config sweep.toml [--dry-run] [--quarantine] [--format ndjson|csv] [-v]. Write down what one run promises before writing code, because each promise becomes a check in the code and a test in the suite:
The contract separates "fix it" from "try later". Exit 2 keeps the meaning from the argparse lesson, "the command could not run as asked": a usage error, or a setting, API path, certificate or token that fails the same way on every run and must not be retried (scripting-adv's "Observable jobs" later gives the setup case its own code, 78). An API outage is 4, which a wrapper may retry in an hour. A partial run beats findings: a run that skipped files must never look complete. A crash is 70, as in hashcheck, never Python's own 1, which would read as "findings, every file checked". The project uses the src layout from the packaging lesson; lab/ holds what stands in for the outside world.
[build-system]requires = ["hatchling==1.32.4"]build-backend = "hatchling.build"[project]name = "ioc-sweep"version = "0.1.0"description = "Sweep a directory for known-bad files: hashes, threat intel, a scanner, quarantine."requires-python = ">=3.14"dependencies = [] # standard library only[project.scripts]ioc-sweep = "ioc_sweep.cli:main"[tool.hatch.build.targets.sdist]include = ["/src"]
dependencies is empty: urllib, tomllib, hashlib, subprocess and csv cover everything, so production has nothing to lock except the tool itself. pytest lives only in the development venv, pinned as in the testing lesson. The settings file:
# ioc-sweep settings. Every key is required and an unknown key is an error.# Paths are absolute: cron and systemd do not start in this directory.[api]url = "https://localhost:18160/v1"ca_file = "/home/deploy/ps-capstone/lab/ca.pem"token_file = "/home/deploy/ps-capstone/api-token"timeout = 5 # seconds for the connect, and for each readbatch_size = 3 # hashes per requestmax_pages = 5 # per batch; more means the API is brokenattempts = 3 # per request, the first one includeddeadline = 15 # seconds for one batch: all its pages and retriesmax_retry_after = 5 # a 429 that asks for longer ends the batch[scan]max_size = 1048576 # bytes; a bigger file is not checked and the run is partialscanner = "/home/deploy/ps-capstone/lab/fakescan"scanner_timeout = 10 # seconds per file[output]summary_csv = "/home/deploy/ps-capstone/out/summary.csv"quarantine_dir = "/home/deploy/ps-capstone/quarantine"
The paths name the lab machine's home directory, /home/deploy. Point them at yours (on the lab machine the command changes nothing, and grep shows the lines it would change):
pip install -e . made an editable install with an ioc-sweep launcher in .venv/bin. The help text and the exit-status epilog come from cli_args.py, shown with main() later.
Settings, the token and the file walk
"""Load sweep.toml and check every value once, so the rest of the tool can trust them."""import tomllibfrom pathlib import Path# Every table, every key in it, and the type its value must have. Nothing else is accepted.SCHEMA = {"api": {"url": str, "ca_file": Path, "token_file": Path, "timeout": float,"batch_size": int, "max_pages": int, "attempts": int, "deadline": float,"max_retry_after": float},"scan": {"max_size": int, "scanner": Path, "scanner_timeout": float},"output": {"summary_csv": Path, "quarantine_dir": Path},}class ConfigError(Exception):"""sweep.toml is missing, is not TOML, or has a wrong, missing or unknown setting."""def load_config(path: Path) -> dict:"""Return {key: value} for every key in SCHEMA, or raise ConfigError with one line."""try:with open(path, "rb") as f: # tomllib reads bytes and decodes UTF-8 itselfdata = tomllib.load(f)except OSError as err:raise ConfigError(f"{path}: {err.strerror}") from errexcept tomllib.TOMLDecodeError as err:raise ConfigError(f"{path}: {err}") from errunknown = sorted(set(data) - set(SCHEMA)) # a typo is reported, never ignoredif unknown:raise ConfigError(f"{path}: unknown table [{unknown[0]}]")settings = {}for table, keys in SCHEMA.items():values = data.get(table)if not isinstance(values, dict):raise ConfigError(f"{path}: no [{table}] table")unknown = sorted(set(values) - set(keys))if unknown:raise ConfigError(f"{path}: unknown key {table}.{unknown[0]}")for key, kind in keys.items():settings[key] = check(f"{path}: {table}.{key}", values.get(key), kind)if not settings["url"].startswith("https://"):raise ConfigError(f"{path}: api.url must start with https://")return settingsdef check(name: str, value, kind: type):if value is None:raise ConfigError(f"{name} is missing")if kind is Path: # a path is written as a string and must be absoluteif type(value) is not str or not Path(value).is_absolute():raise ConfigError(f"{name} must be an absolute path, not {value!r}")return Path(value)if kind is float and type(value) is int:value = float(value) # timeout = 5 means 5.0# type(), not isinstance(): True is an instance of int, and "attempts = true" is a mistake.if type(value) is not kind:raise ConfigError(f"{name} must be {kind.__name__}, not {type(value).__name__}")if kind in (int, float) and value <= 0:raise ConfigError(f"{name} must be greater than 0, not {value}")return value
SCHEMA is the single list of what the file may contain. Unknown tables and keys are checked first, with set difference (set(values) - set(keys) is the names in the file that the schema lacks), so batch_sise is reported as a typo, not ignored while a default takes its place. check() compares with type() rather than isinstance() for the reason in its comment: True is an int to isinstance(). Paths must be absolute because cron starts jobs in the home directory. Every problem becomes one ConfigError line.
"""The API token: read from a file only its owner can read, and kept out of every log line."""import loggingimport osimport statfrom pathlib import Pathclass SecretError(Exception):"""The token file is missing, empty, or readable by other users."""def read_token(path: Path) -> str:"""The token: systemd's credential api-token if the service has one, else the file at path."""cred_dir = os.environ.get("CREDENTIALS_DIRECTORY") # set by LoadCredential=api-token:...systemd = bool(cred_dir) and Path(cred_dir, "api-token").is_file()if systemd:path = Path(cred_dir, "api-token")try:with open(path, encoding="utf-8") as f:info = os.fstat(f.fileno()) # the file we opened, not whatever the name points to later# systemd's copy is root-owned and readable by this service only (an ACL): the owner# and mode checks are for a file we manage ourselvesif not systemd and info.st_uid != os.getuid():raise SecretError(f"{path}: owned by uid {info.st_uid}, not by uid {os.getuid()}")if not systemd and info.st_mode & 0o077:raise SecretError(f"{path}: mode {stat.filemode(info.st_mode)} lets other users read it")token = f.read().strip()except OSError as err:raise SecretError(f"{path}: {err.strerror}") from errif not token:raise SecretError(f"{path}: empty")return tokenclass RedactSecrets(logging.Filter):"""Replace known secret values in every record before the handler writes it."""def __init__(self, secrets: list[str]):super().__init__()self.secrets = [s for s in secrets if s]def scrub(self, text: str) -> str:for secret in self.secrets:text = text.replace(secret, "[REDACTED]")return textdef filter(self, record: logging.LogRecord) -> bool:record.msg = self.scrub(record.getMessage())record.args = Noneif record.exc_info: # a traceback can hold the secret toorecord.exc_text = self.scrub(logging.Formatter().formatException(record.exc_info))return True
Both parts are the secrets lesson's code. read_token() follows that lesson's ranking: a systemd credential api-token first (from LoadCredential=api-token:/etc/credstore/ioc-sweep-token in the unit), without the owner check that systemd's root-owned copy would fail; otherwise read_secret_file(), which refuses a file other accounts could read. RedactSecrets is the logging filter. The token goes into one place only, the Authorization header; never into arguments, and never into the scanner's environment.
"""List the regular files under a directory with their SHA-256, never following a symbolic link."""import hashlibimport loggingimport statfrom pathlib import Pathlog = logging.getLogger(__name__)def inventory(root: Path, max_size: int) -> list[dict]:"""One record per regular file: its path and SHA-256, or the problems that stopped the hash."""records = []def cannot_list(err: OSError) -> None: # walk() calls this for a directory it cannot readrecords.append({"path": Path(err.filename), "sha256": "", "problems": [err.strerror]})for dirpath, dirnames, filenames in root.walk(on_error=cannot_list):dirnames.sort() # walk() visits the subdirectories in this list's orderfor name in sorted(filenames):path = dirpath / namerecord = {"path": path, "sha256": "", "problems": []}try:info = path.lstat() # the entry itself: a link is seen as a linkif not stat.S_ISREG(info.st_mode):log.info("skipped %s: not a regular file", path)continueif info.st_size > max_size:record["problems"].append(f"{info.st_size} bytes, over scan.max_size")else:with open(path, "rb") as f:record["sha256"] = hashlib.file_digest(f, "sha256").hexdigest()except OSError as err: # removed or made unreadable while we walkedrecord["problems"].append(err.strerror)records.append(record)return records
Path.walk() (Python 3.12 and later) works like find: for each directory, top down, it yields the directory and the names of its subdirectories and files, follows no symbolic links, and descends in the order of dirnames, sorted in place here. By default it skips a directory it cannot read without a word, a false all-clear in a sweep, so on_error= records it. lstat() describes the entry itself, and stat.S_ISREG() is true only for a regular file: not a link, and not a FIFO, whose read would block forever. The gap between lstat() and open() is the one the files lesson named; scripting-adv closes it.
The tests share fixtures in tests/conftest.py:
"""Fixtures: a lab CA, the mock API on a free port, a token file, a small tree and a config."""import subprocessimport threadingfrom pathlib import Pathimport mock_api # lab/mock_api.py, on sys.path through pytest.tomlimport pytestLAB = Path(__file__).parent.parent / "lab"TOKEN = "lab-token-not-secret"SAMPLE_A = b"ioc-sweep lab sample A: stands in for a malicious file\n" # in lab/known-bad.txtKNOWN_BAD = sorted(line.split()[0] for line in (LAB / "known-bad.txt").read_text().splitlines())TEMPLATE = Path(__file__).parent / "sweep.template.toml" # sweep.toml with {placeholders}@pytest.fixture(scope="session")def certs(tmp_path_factory):"""A throwaway CA and server certificate, made once per test run with the lab's own script."""directory = tmp_path_factory.mktemp("ca")subprocess.run(["bash", LAB / "make_lab_ca.sh"], cwd=directory, check=True,capture_output=True, timeout=30)return directory@pytest.fixturedef api(certs):"""The mock API on 127.0.0.1, on a port the OS picks. A test sets api.mode to make it misbehave."""server = mock_api.make_server(certs, TOKEN)threading.Thread(target=server.serve_forever, daemon=True).start()server.url = f"https://localhost:{server.server_address[1]}/v1"yield serverserver.shutdown()server.server_close()@pytest.fixturedef tree(tmp_path):"""A known-bad sample, two ordinary files, and a symbolic link that points out of the tree."""root = tmp_path / "tree"(root / "docs").mkdir(parents=True)(root / "invoice.pdf").write_bytes(SAMPLE_A) # SHA-256 1f72...(root / "docs" / "notes.txt").write_text("meeting notes\n") # 2f96...(root / "docs" / "backup.log").write_text("backup done\n") # 83b7..., the 8-f shard(root / "passwd").symlink_to("/etc/passwd")return root@pytest.fixturedef make_config(tmp_path, certs, api):"""Return a function that writes sweep.toml for this test's API and files; keywords change values."""token_file = tmp_path / "api-token"token_file.write_text(TOKEN + "\n")token_file.chmod(0o600)def make(**changes) -> Path:values = {"url": api.url, "certs": certs, "tmp": tmp_path, "deadline": 5,"max_retry_after": 2, "scanner": LAB / "fakescan"}values.update(changes)path = tmp_path / "sweep.toml"path.write_text(TEMPLATE.read_text().format(**values))return pathreturn make
certs runs the lab's CA script once per test session (scope="session"); api is the next section's subject. make_config returns a function; **changes collects its keyword arguments into a dict, so make_config(deadline=2) writes a sweep.toml with one value changed, filling the {placeholders} of sweep.template.toml with str.format(). pytest.toml puts lab/ on the import path, so conftest.py can import the mock API:
[api]url = "{url}"ca_file = "{certs}/ca.pem"token_file = "{tmp}/api-token"timeout = 1batch_size = 2max_pages = 3attempts = 3deadline = {deadline}max_retry_after = {max_retry_after}[scan]max_size = 1000scanner = "{scanner}"scanner_timeout = 5[output]summary_csv = "{tmp}/summary.csv"quarantine_dir = "{tmp}/quarantine"
[pytest]testpaths = ["tests"]pythonpath = ["lab"] # the mock API is a lab script, not part of the package
"""load_config() and read_token(): what they accept, and the one line they say about the rest."""import pytestfrom ioc_sweep.config import ConfigError, load_configfrom ioc_sweep.secret import SecretError, read_tokendef test_values_are_typed(make_config):settings = load_config(make_config())assert settings["timeout"] == 1.0 and type(settings["timeout"]) is floatassert settings["batch_size"] == 2assert settings["scanner"].is_absolute()@pytest.mark.parametrize("old, new, message",[("batch_size = 2", "batch_sise = 2", "unknown key api.batch_sise"),("max_pages = 3", 'max_pages = "3"', "api.max_pages must be int, not str"),("attempts = 3", "attempts = true", "api.attempts must be int, not bool"),("attempts = 3", "attempts = 0", "api.attempts must be greater than 0"),('url = "https', 'url = "http', "api.url must start with https://"),('summary_csv = "/', 'summary_csv = "', "output.summary_csv must be an absolute path"),("[output]", "[extra]\nx = 1\n[output]", "unknown table [extra]"),("timeout = 1", "timeout = 1s", "(at line 5, column 12)"),],ids=["typo", "string", "bool", "zero", "http", "relative", "table", "not-toml"],)def test_bad_setting_is_one_clear_error(make_config, old, new, message):path = make_config()path.write_text(path.read_text().replace(old, new, 1))with pytest.raises(ConfigError) as err:load_config(path)assert message in str(err.value)def test_token_file_must_be_private(tmp_path):token_file = tmp_path / "api-token"token_file.write_text("lab-token-not-secret\n")token_file.chmod(0o644)with pytest.raises(SecretError, match="-rw-r--r-- lets other users read it"):read_token(token_file)token_file.chmod(0o600)assert read_token(token_file) == "lab-token-not-secret"def test_systemd_credential_comes_first(tmp_path, monkeypatch):creds = tmp_path / "credentials"creds.mkdir()(creds / "api-token").write_text("lab-token-from-systemd\n")(creds / "api-token").chmod(0o440) # group-readable, as systemd's copy is through its ACLmonkeypatch.setenv("CREDENTIALS_DIRECTORY", str(creds))assert read_token(tmp_path / "no-such-file") == "lab-token-from-systemd"
"""inventory(): regular files only, links skipped, a size cap, and unreadable places recorded."""import hashlibfrom conftest import SAMPLE_Afrom ioc_sweep.walk import inventorydef test_regular_files_hashed_and_link_skipped(tree):records = inventory(tree, max_size=1000)assert [r["path"].relative_to(tree).as_posix() for r in records] == ["invoice.pdf", "docs/backup.log", "docs/notes.txt"]assert records[0]["sha256"] == hashlib.sha256(SAMPLE_A).hexdigest()assert all(not r["problems"] for r in records)def test_problems_are_recorded_not_skipped(tree):(tree / "big.iso").write_bytes(b"x" * 2000)(tree / "locked").mkdir()(tree / "locked").chmod(0o000) # a directory walk() cannot listtry:problems = {r["path"].name: r["problems"] for r in inventory(tree, max_size=1000) if r["problems"]}finally:(tree / "locked").chmod(0o755) # so pytest can delete it afterwardsassert problems == {"big.iso": ["2000 bytes, over scan.max_size"], "locked": ["Permission denied"]}
Each parametrized case is one way a person gets a setting wrong, with the one line they will read. The walk test shows the link skipped, and a 2000-byte file and a directory with mode 000 recorded as problems instead of disappearing.
Talking to the threat-intel API
The API takes a batch of digests, GET /v1/hashes?sha256=H1,H2,H3, and answers {"matches": [...], "next_cursor": ...} with one match per page. lab/mock_api.py plays it over TLS, knows the two hashes in lab/known-bad.txt, and misbehaves on request: in --mode partial any batch holding a hash that starts with 8 to f gets 503; down always answers 503, throttle answers the first request with 429 and Retry-After: 1, slow stalls, drip sends one byte at a time and loop repeats its cursor. Another path gets 404, and a wrong token gets 401 with a reason phrase that repeats the token, as some real APIs do.
"""One GET request to the threat-intel API: verified TLS, the token, time limits, checked JSON."""import http.clientimport jsonimport sslimport timeimport urllib.errorimport urllib.requestMAX_BODY = 1_000_000 # bytes; a larger answer is a bug or an attack, not dataclass ApiError(Exception):"""This request will not work however often it is sent: a wrong status or a wrong answer."""class SetupError(ApiError):"""Every request would fail the same way: a wrong api.url, a certificate that does not verify."""class AuthError(SetupError):"""401 or 403: the token is wrong or not allowed. Retrying cannot help and may lock it."""class Unavailable(ApiError):"""A timeout, a refused or dropped connection, 5xx or 429: the same request may work later."""def __init__(self, message: str, retry_after: float | None = None):super().__init__(message)self.retry_after = retry_after # seconds the server asked us to wait, if it saiddef get_json(url: str, token: str, context: ssl.SSLContext, timeout: float, deadline: float) -> dict:"""timeout limits each step; deadline (a time.monotonic() value) also limits reading the body."""request = urllib.request.Request(url, headers={"Accept": "application/json"})request.add_unredirected_header("Authorization", f"Bearer {token}") # not resent on a redirecttry:with urllib.request.urlopen(request, timeout=timeout, context=context) as response:content_type = response.headers.get_content_type()if content_type != "application/json":raise ApiError(f"expected JSON, got {content_type}")body = read_body(response, deadline)except urllib.error.HTTPError as err: # the server answered with 4xx or 5xxmessage = f"HTTP {err.code} {err.reason}"if err.code in (401, 403):raise AuthError(message) from errif err.code == 404: # the path in api.url is wrong, for every request alikeraise SetupError(message) from errif err.code == 429 or err.code >= 500:raise Unavailable(message, retry_after_seconds(err.headers.get("Retry-After"))) from errraise ApiError(message) from errexcept (OSError, http.client.HTTPException) as err: # no usable answer (see the HTTP lesson)reason = getattr(err, "reason", err)if isinstance(reason, ssl.SSLCertVerificationError): # wrong name, expired, unknown CAraise SetupError(f"TLS: {reason}") from errraise Unavailable(str(reason)) from errtry:data = json.loads(body)except ValueError as err: # not JSON, or not UTF-8raise ApiError(f"invalid JSON: {err}") from errif not isinstance(data, dict):raise ApiError(f"expected a JSON object, got {type(data).__name__}")return datadef read_body(response, deadline: float) -> bytes:"""Chunk by chunk, so that a server sending one byte at a time cannot outlast the deadline."""body = b""while chunk := response.read1(65536):body += chunkif len(body) > MAX_BODY:raise ApiError(f"answer larger than {MAX_BODY} bytes")if time.monotonic() > deadline:raise Unavailable("answer still incomplete at api.deadline")return bodydef retry_after_seconds(value: str | None) -> float | None:"""Retry-After in seconds; an HTTP date or anything else counts as 'too long'."""if value is None:return Nonereturn float(value) if value.isascii() and value.isdigit() else float("inf")
This is get_json() from the HTTP lesson, with its handling of dropped and garbled answers and of a server that sends one byte at a time (read_body() checks the batch's deadline between chunks). The changes are in the classes. SetupError means "every request would fail the same way": a 404 on the API path, or a certificate that does not verify (a wrong name, an expired certificate, an unknown CA, or an interception) needs a person, not a retry. AuthError, for 401 and 403, is a kind of SetupError. Unavailable carries the server's Retry-After.
"""One request, retried while the API is unavailable: bounded, jittered, inside a deadline."""import loggingimport randomimport timefrom .api import Unavailable, get_jsonlog = logging.getLogger(__name__)def get_with_retries(url: str, token: str, context, settings: dict, deadline: float) -> dict:for attempt in range(1, settings["attempts"] + 1):left = deadline - time.monotonic()if left <= 0:raise Unavailable(f"api.deadline of {settings['deadline']:g}s reached")try:return get_json(url, token, context, timeout=min(settings["timeout"], left), deadline=deadline)except Unavailable as err:if attempt == settings["attempts"]: # the last attempt: give up now, no sleepraise Unavailable(f"{err} ({attempt} attempts)") from errdelay = random.uniform(0, min(4.0, 0.5 * 2 ** (attempt - 1))) # full jitterif err.retry_after is not None: # a 429 that says how long to waitif err.retry_after > settings["max_retry_after"]:raise Unavailable(f"{err}, Retry-After {err.retry_after:g}s is too long") from errdelay = err.retry_afterif delay >= deadline - time.monotonic():raise Unavailable(f"{err}, no time left before api.deadline") from errlog.warning("%s; retry %d in %.1fs", err, attempt, delay)time.sleep(delay)
The loop is retry() from the errors lesson: bounded attempts, full jitter, no sleep after the last attempt, and every attempt's timeout cut to the time left before the deadline. A 429 replaces the jitter with the server's wait, but only up to api.max_retry_after; a server asking for longer ends the batch.
"""Look digests up in batches: every page of a batch, inside limits the API cannot break."""import loggingimport timefrom urllib.parse import urlencodefrom .api import ApiError, SetupError, Unavailablefrom .retry import get_with_retrieslog = logging.getLogger(__name__)VERDICTS = {"malicious", "suspicious"}def lookup_all(digests: list[str], settings: dict, token: str, context) -> tuple[dict, set]:"""Return ({sha256: match}, {digests whose batch failed}). SetupError stops everything."""size = settings["batch_size"]batches = [digests[i:i + size] for i in range(0, len(digests), size)]matches, failed, temporary = {}, set(), Falsefor number, batch in enumerate(batches, 1):try:matches.update(lookup_batch(batch, settings, token, context))except SetupError:raise # the same URL, certificate or token would fail every batch: stop sendingexcept ApiError as err: # Unavailable included: this batch gave up, try the nextlog.warning("batch %d of %d failed: %s", number, len(batches), err)failed.update(batch)temporary = temporary or isinstance(err, Unavailable)if digests and len(failed) == len(digests) and not temporary:# every batch got a wrong answer, none of them temporary: waiting will not fix thatraise SetupError(f"all {len(batches)} batches failed, none of them temporary")return matches, faileddef lookup_batch(batch: list[str], settings: dict, token: str, context) -> dict:"""Every page of one batch; one deadline covers all its pages and retries."""deadline = time.monotonic() + settings["deadline"]matches, seen, cursor = {}, set(), Nonefor page in range(1, settings["max_pages"] + 1):query = {"sha256": ",".join(batch)}if cursor is not None:query["cursor"] = cursorurl = f"{settings['url']}/hashes?{urlencode(query)}"data = get_with_retries(url, token, context, settings, deadline)items, cursor = data.get("matches"), data.get("next_cursor")if not isinstance(items, list) or not (cursor is None or isinstance(cursor, str)):raise ApiError(f"page {page}: not a {{matches, next_cursor}} answer")for item in items:if not isinstance(item, dict) or item.get("sha256") not in batch \or item.get("verdict") not in VERDICTS or not isinstance(item.get("name"), str):raise ApiError(f"page {page}: unexpected match {item!r}")matches[item["sha256"]] = itemif cursor is None:return matchesif cursor in seen:raise ApiError(f"page {page} repeats cursor {cursor!r}: the API is looping")seen.add(cursor)raise ApiError(f"more than api.max_pages ({settings['max_pages']}) pages")
lookup_batch() is the pagination loop from the HTTP lesson, with both guards, a deadline per batch, and a check of every match before it is kept. urlencode() builds the query string and percent-encodes each value. lookup_all() decides what a failure costs: a failed batch is logged and its digests go into failed, so the next batch still runs, but a bare raise passes a SetupError on, because the same URL, certificate or token would fail every batch. So does "every batch failed, none of them temporarily" (an API that loops or sends the wrong data): waiting an hour will not fix it.
"""lookup_all() against the mock API over real HTTPS, once for each way the API behaves."""import sslimport timeimport pytestfrom conftest import KNOWN_BAD, TOKENfrom ioc_sweep.api import AuthError, SetupErrorfrom ioc_sweep.config import load_configfrom ioc_sweep.intel import lookup_allA, B = KNOWN_BAD # both start with 0-7LOW, HIGH = "0" * 64, "e" * 64 # unknown hashes, one on each shard@pytest.fixturedef lookup(make_config):"""Call lookup_all() with this test's config; keyword arguments change config values."""def run(digests, token=TOKEN, **changes):settings = load_config(make_config(**changes))context = ssl.create_default_context(cafile=settings["ca_file"])return lookup_all(digests, settings, token, context)return rundef test_two_matches_two_pages(api, lookup):assert lookup([A, B]) == ({A: {"sha256": A, "verdict": "malicious", "name": "lab-dropper-a"},B: {"sha256": B, "verdict": "suspicious", "name": "lab-loader-b"}}, set())assert api.requests == 2 # one batch of two, one match per pagedef test_wrong_token_is_not_retried(api, lookup):with pytest.raises(AuthError, match="401"):lookup([A], token="lab-token-revoked")assert api.requests == 1def test_429_waits_for_retry_after(api, lookup):api.mode = "throttle"start = time.monotonic()assert list(lookup([A])[0]) == [A]assert api.requests == 2 and time.monotonic() - start >= 1@pytest.mark.parametrize("mode, failed, requests", [("down", {LOW, A, HIGH}, 6), # two batches, three attempts each("partial", {HIGH}, 4), # batch 2 holds a hash from the broken shard], ids=["down", "partial"])def test_failed_batches_are_reported(api, lookup, mode, failed, requests):api.mode = modeassert lookup([LOW, A, HIGH])[1] == failedassert api.requests == requestsdef test_every_batch_wrong_is_a_setup_error(api, lookup):api.mode = "loop" # in each batch, page 2 repeats the cursor of page 1: nothing temporarywith pytest.raises(SetupError, match="none of them temporary"):lookup([LOW, A, HIGH])assert api.requests == 4def test_dripping_api_is_bounded_by_the_deadline(api, lookup):api.mode, api.delay = "drip", 0.2 # every read gets a byte in time; the whole answer never doesstart = time.monotonic()assert lookup([A], deadline=2)[1] == {A}assert time.monotonic() - start < 3def test_stalled_api_is_bounded_by_the_deadline(api, lookup):api.mode, api.delay = "slow", 3start = time.monotonic()assert lookup([A], deadline=2)[1] == {A}assert time.monotonic() - start < 2.5
The request counts are the behaviour: two pages for two matches, one request for the refused token, six for down (three attempts for each of two batches), four for partial, and four for loop, stopped on page 2 of each batch and reported as a setup error. The dripping and the stalled API each stopped at about two seconds, the deadline; the drip needed over half a minute.
The scanner, the output and the quarantine
#!/bin/sh# fakescan: a stand-in for a malware scanner, for the ioc-sweep lab.# Usage: fakescan --json -- FILE# Prints a JSON report; a finding when FILE holds the lab's test signature.# Exit status: 0 report printed, 2 usage error or unreadable file.if [ "$#" -ne 3 ] || [ "$1" != --json ] || [ "$2" != -- ]; thenecho "usage: fakescan --json -- FILE" >&2exit 2figrep -q -F -e LAB-TEST-SIGNATURE -- "$3"case $? in0) echo '{"findings": [{"rule": "lab-test-signature", "severity": "high"}]}' ;;1) echo '{"findings": []}' ;;*) exit 2 ;;esac
"""Run the external scanner on one file: a list, --, a timeout and a scrubbed environment."""import jsonimport subprocessfrom pathlib import PathCHILD_ENV = {"PATH": "/usr/bin:/bin", "LANG": "C.UTF-8"} # no token, nothing else inheritedclass ScanError(Exception):"""No usable answer from the scanner: missing, failed, hung, or not its JSON report."""def scan_file(scanner: Path, path: Path, timeout: float) -> list[dict]:try:result = subprocess.run([scanner, "--json", "--", path], # "--": a file named -x is not an optioncapture_output=True, timeout=timeout, env=CHILD_ENV,encoding="utf-8", errors="replace", # a file name in its output may not be UTF-8)except OSError as err: # FileNotFoundError, PermissionError: it never startedraise ScanError(f"cannot run {scanner}: {err.strerror}") from errexcept subprocess.TimeoutExpired as err:raise ScanError(f"no answer within {timeout:g}s") from errif result.returncode != 0:raise ScanError(f"exited {result.returncode}: {result.stderr.strip()}")try:findings = json.loads(result.stdout)["findings"]return [{"rule": str(f["rule"]), "severity": str(f["severity"])} for f in findings]except (json.JSONDecodeError, KeyError, TypeError) as err:raise ScanError("its output is not the JSON report") from err
fakescan stands in for a malware scanner and flags any file holding LAB-TEST-SIGNATURE. scan_file() is the testing lesson's scan_file() with four changes: the scanner's absolute path comes from the config; env=CHILD_ENV gives it a scrubbed environment; except OSError also covers a file that is not executable; and errors="replace" decodes output that repeats a file name which is not UTF-8. A timeout kills the scanner process only; the subprocess lesson showed the process-group kill for scanners that start children.
"""scan_file() with the lab's fakescan and with stubs that fail, hang or print their environment."""from pathlib import Pathimport pytestfrom conftest import LABfrom ioc_sweep.scanner import ScanError, scan_file@pytest.fixturedef stub(tmp_path):"""Return a function that writes an executable stub scanner running the given sh code."""def write(sh_code: str) -> Path:path = tmp_path / "stubscan"path.write_text(f"#!/bin/sh\n{sh_code}\n")path.chmod(0o755)return pathreturn writedef test_signature_found_and_option_like_name_is_a_file(tmp_path, monkeypatch):monkeypatch.chdir(tmp_path)Path("--help").write_text("text LAB-TEST-SIGNATURE text\n") # a name that looks like an optionassert scan_file(LAB / "fakescan", Path("--help"), 5) == [{"rule": "lab-test-signature", "severity": "high"}]@pytest.mark.parametrize("sh_code, timeout, message", [("echo 'signature database missing' >&2; exit 2", 5, "exited 2: signature database missing"),("echo 'Scanning... done'", 5, "not the JSON report"),("exec sleep 30", 0.5, "no answer within 0.5s"), # exec: the kill on timeout reaches sleep], ids=["fails", "garbage", "hangs"])def test_scanner_failures(stub, sh_code, timeout, message):with pytest.raises(ScanError, match=message):scan_file(stub(sh_code), Path("x"), timeout)def test_missing_scanner(tmp_path):with pytest.raises(ScanError, match="No such file or directory"):scan_file(tmp_path / "nowhere", Path("x"), 5)def test_scanner_gets_no_token(stub, tmp_path, monkeypatch):monkeypatch.setenv("IOC_SWEEP_TOKEN", "lab-token-not-secret") # as if a wrapper exported itscan_file(stub(f"""env > {tmp_path}/env; echo '{{"findings": []}}'"""), Path("x"), 5)env = (tmp_path / "env").read_text()assert "lab-token-not-secret" not in envassert "PATH=/usr/bin:/bin\n" in env
"""Findings on stdout as NDJSON or CSV, and the per-file summary CSV, replaced atomically."""import csvimport jsonimport osimport sysimport tempfilefrom pathlib import PathFINDING_FIELDS = ["path", "sha256", "source", "detail"]SUMMARY_FIELDS = ["path", "sha256", "status", "detail"]def printable(path) -> str:"""A path as text. Linux file names are bytes; bytes that are not UTF-8 are written as \\xNN."""return os.fsencode(path).decode("utf-8", "backslashreplace")def neutralise(value) -> str:"""Stop a spreadsheet from running a cell as a formula: file names come from outside."""text = str(value)return "'" + text if text.startswith(("=", "+", "-", "@", "\t", "\r", "\n")) else textdef write_findings(findings: list[dict], fmt: str) -> None:if fmt == "ndjson":for finding in findings:print(json.dumps(finding)) # one object per line; json.dumps escapes any file namereturnwriter = csv.DictWriter(sys.stdout, FINDING_FIELDS)writer.writeheader()for finding in findings:writer.writerow({key: neutralise(finding[key]) for key in FINDING_FIELDS})def write_summary(rows: list[dict], target: Path) -> None:"""Readers see the old summary or the new one, never half of one."""fd, tmp = tempfile.mkstemp(dir=target.parent, prefix=f".{target.name}.")try:with open(fd, "w", encoding="utf-8", newline="") as f:writer = csv.DictWriter(f, SUMMARY_FIELDS)writer.writeheader()for row in rows:writer.writerow({key: neutralise(row[key]) for key in SUMMARY_FIELDS})f.flush()os.fsync(f.fileno())os.replace(tmp, target)except BaseException:os.unlink(tmp)raise
Findings are NDJSON by default, one json.dumps() object per line as in the logs lesson, or CSV. Every CSV cell goes through neutralise() from the data lesson, because file names are chosen by whoever uploaded them. For the same reason every path goes through printable(). A Linux file name is bytes; Python represents bytes that are not valid UTF-8 as placeholder characters (lone surrogates), and writing one to a UTF-8 file raises UnicodeEncodeError: one oddly named upload would crash the run before the summary and the quarantine. os.fsencode() gives back the bytes, and backslashreplace writes each invalid one as \xNN. The summary uses the files lesson's temporary file, fsync() and os.replace(); a SIGTERM from timeout ends Python without running except and can leave a .summary.csv.* behind (scripting-adv handles signals). quarantine.py is the argparse lesson's module, unchanged.
"""Everything that changes files: the summary and the quarantine. --dry-run stops here, only here."""import loggingfrom .output import printable, write_summaryfrom .quarantine import QuarantineError, destination, prepare, quarantinelog = logging.getLogger(__name__)def apply(records: list[dict], findings: list[dict], settings: dict, move: bool, dry_run: bool) -> bool:"""Write the summary and, with move, quarantine each file with a finding. False: not all done."""found = {} # path -> the details of its findingsfor f in findings:found.setdefault(f["path"], []).append(f"{f['source']} {f['detail']}")rows = [summary_row(r, found.get(printable(r["path"]))) for r in records]done = Trueif dry_run:log.info("dry run: would write %d rows to %s", len(rows), settings["summary_csv"])else:try:write_summary(rows, settings["summary_csv"])except OSError as err:log.error("summary not written: %s: %s", settings["summary_csv"], err.strerror)done = Falseif not move:return doneinto = settings["quarantine_dir"]try:prepare(into, create=not dry_run)except (OSError, QuarantineError) as err:log.error("nothing quarantined: %s", err)return Falsefor r in records:if printable(r["path"]) not in found:continueif dry_run: # the plan comes from the same code path as the real moveslog.info("dry run: would move %s -> %s", r["path"], destination(into, r["path"], r["sha256"]))continuetry:log.info("quarantined %s -> %s", r["path"], quarantine(r["path"], r["sha256"], into))except (OSError, QuarantineError) as err: # FileExistsError included: never overwritelog.warning("%s: not moved: %s", r["path"], err)done = Falsereturn donedef summary_row(record: dict, details: list[str] | None) -> dict:if details:status = "finding"else:status, details = ("incomplete" if record["problems"] else "clean"), record["problems"]return {"path": printable(record["path"]), "sha256": record["sha256"], "status": status,"detail": "; ".join(details)}
Everything that changes a file is in apply(), so an honest --dry-run is one if in each place, on the same code path as the real run. found.setdefault(path, []) returns the list stored for a path, creating it the first time. A summary not written, an unsafe quarantine directory or a file not moved makes the run partial.
main(), and the sweep end to end
"""The checks: hash the tree, look the digests up in threat intel, run the scanner on each file."""import loggingfrom .intel import lookup_allfrom .output import printablefrom .scanner import ScanError, scan_filefrom .walk import inventorylog = logging.getLogger(__name__)class ApiUnavailable(Exception):"""Every lookup failed. Without threat intel, "no findings" would be a false all-clear."""def check_tree(root, settings: dict, token: str, context) -> tuple[list[dict], list[dict]]:"""Return (records, findings). A record's problems say why its file was not fully checked."""records = inventory(root, settings["max_size"])if not records: # an unmounted share looks exactly like thislog.warning("no files at all under %s: is it the right directory, and mounted?", root)digests = sorted({r["sha256"] for r in records if r["sha256"]}) # each content oncelog.info("%d files under %s, %d distinct digests to look up", len(records), root, len(digests))matches, failed = lookup_all(digests, settings, token, context)if digests and len(failed) == len(digests):raise ApiUnavailable(f"all {len(digests)} lookups failed")findings = []for record in records:if not record["sha256"]:continue # not hashed: the reason is already in its problemsif record["sha256"] in failed:record["problems"].append("lookup failed")match = matches.get(record["sha256"])if match:findings.append(finding(record, "intel", f"{match['verdict']} {match['name']}"))try:hits = scan_file(settings["scanner"], record["path"], settings["scanner_timeout"])except ScanError as err:log.warning("%s: scanner: %s", record["path"], err)record["problems"].append(f"scanner: {err}")continuefor hit in hits:findings.append(finding(record, "scanner", f"{hit['severity']} {hit['rule']}"))return records, findingsdef finding(record: dict, source: str, detail: str) -> dict:return {"path": printable(record["path"]), "sha256": record["sha256"], "source": source, "detail": detail}
Identical files are looked up once (sorted() of a set). If every lookup failed, check_tree() raises instead of returning: with no threat intel, "no findings" would be a false all-clear. A file whose lookup failed is still scanned and recorded as incomplete. A tree with no files at all is legal but suspicious, since an unmounted share looks exactly like that, so it is logged as a warning; the exit status stays 0, and a monitor should watch for that line.
"""The command line of ioc-sweep."""import argparsefrom pathlib import Pathdef build_parser() -> argparse.ArgumentParser:parser = argparse.ArgumentParser(prog="ioc-sweep", description="Sweep a directory for known-bad files.", allow_abbrev=False,epilog="exit status: 0 clean, 1 findings, 2 fix the setup, 3 partial, 4 API unavailable, ""70 internal error")commands = parser.add_subparsers(dest="command", required=True, metavar="COMMAND")scan = commands.add_parser("scan", allow_abbrev=False, help="sweep one directory tree")scan.add_argument("dir", type=Path, metavar="DIR", help="the directory to sweep")scan.add_argument("--config", required=True, type=Path, metavar="FILE", help="settings (TOML)")scan.add_argument("--dry-run", action="store_true", help="look everything up, change nothing")scan.add_argument("--quarantine", action="store_true", help="move files with findings away")scan.add_argument("--format", choices=["ndjson", "csv"], default="ndjson",help="how findings are written to stdout (default: %(default)s)")scan.add_argument("-v", "--verbose", action="store_true", help="log debug messages too")return parser
"""ioc-sweep: sweep a directory for known-bad files.Exit status: 0 clean; 1 findings; 2 usage, config, token or API setup problem (fix it, then runagain); 3 partial: some files were not fully checked (wins over 1); 4 threat-intel APIunavailable (try later); 70 internal error: a bug or an input the tool did not expect."""import loggingimport sslfrom .actions import applyfrom .api import AuthError, SetupErrorfrom .cli_args import build_parserfrom .config import ConfigError, load_configfrom .output import write_findingsfrom .secret import RedactSecrets, SecretError, read_tokenfrom .sweep import ApiUnavailable, check_treeCLEAN, FINDINGS, USAGE, PARTIAL, UNAVAILABLE, INTERNAL = 0, 1, 2, 3, 4, 70log = logging.getLogger("ioc_sweep") # the parent of every module's logger in the packagedef main(argv: list[str] | None = None) -> int:args = build_parser().parse_args(argv) # a usage error exits 2 herehandler = logging.StreamHandler() # stderrhandler.setFormatter(logging.Formatter("ioc-sweep: %(levelname)s %(message)s"))log.handlers = [handler] # replace, not add: tests call main() many times in one processlog.setLevel(logging.DEBUG if args.verbose else logging.INFO)try:return sweep(args, handler)except Exception: # never let a crash exit 1, which would read as "findings, all checked"log.exception("internal error; the sweep is incomplete")return INTERNALdef sweep(args, handler: logging.Handler) -> int:try:settings = load_config(args.config)token = read_token(settings["token_file"])context = ssl.create_default_context(cafile=settings["ca_file"])except (ConfigError, SecretError) as err:log.error("%s", err)return USAGEexcept OSError as err: # the CA file: missing, or not a certificate (ssl.SSLError)log.error("api.ca_file %s: %s", settings["ca_file"], err)return USAGEhandler.addFilter(RedactSecrets([token])) # from here on, no log line can show the tokenif not args.dir.is_dir() or settings["quarantine_dir"].resolve().is_relative_to(args.dir.resolve()):log.error("%s: not a directory, or output.quarantine_dir is inside it", args.dir)return USAGEtry:records, findings = check_tree(args.dir, settings, token, context)except AuthError as err: # a subclass of SetupError, so it comes firstlog.error("the API refused the token: %s (not retried)", err)return USAGEexcept SetupError as err:log.error("the API cannot work with these settings: %s (not retried)", err)return USAGEexcept ApiUnavailable as err:log.error("threat-intel API unavailable: %s; nothing reported", err)return UNAVAILABLEwrite_findings(findings, args.format)if not apply(records, findings, settings, move=args.quarantine, dry_run=args.dry_run):return PARTIALif any(r["problems"] for r in records):return PARTIALreturn FINDINGS if findings else CLEAN
Logging is set up on the package's logger, ioc_sweep, the parent of every getLogger(__name__) in the package. The handler list is replaced rather than added to, because tests call main() many times in one process; basicConfig(), the errors lesson's one-liner, does nothing once the root logger has a handler, and pytest gives it one. main() wraps the whole sweep in the internal-error handler; inside sweep() the redaction filter goes on the handler once the token is known, so even a traceback is scrubbed, and the checks run in contract order.
"""ioc-sweep through main(argv): exit statuses, stdout and stderr, and the files it changes."""import jsonimport osimport pytestfrom ioc_sweep import clifrom ioc_sweep.cli import maindef files_under(root):return sorted(str(p.relative_to(root)) for p in root.rglob("*"))def sweep(tree, config, *options):return main(["scan", str(tree), "--config", str(config), *options])@pytest.mark.parametrize("mode, status", [("ok", 1), ("partial", 3), ("down", 4)])def test_exit_status_follows_the_api(api, tree, make_config, capsys, mode, status):api.mode = modeassert sweep(tree, make_config()) == statusout = capsys.readouterr().outif status == 4:assert out == "" # without threat intel nothing is reported, not even "clean"else:assert json.loads(out)["path"] == str(tree / "invoice.pdf")def test_clean_tree(api, tree, make_config):(tree / "invoice.pdf").unlink()assert sweep(tree, make_config()) == 0def test_config_error_and_usage_error(api, tree, make_config, capsys):config = make_config()config.write_text(config.read_text().replace("batch_size", "batch_sise"))assert sweep(tree, config) == 2assert "unknown key api.batch_sise" in capsys.readouterr().errwith pytest.raises(SystemExit) as exc:sweep(tree, config, "--format", "yaml")assert exc.value.code == 2@pytest.mark.parametrize("old, new, requests", [("/v1", "/v2", 1), # a wrong path: 404, the same for every request("localhost", "127.0.0.1", 0), # a name the certificate does not list: no request at all], ids=["404", "tls-name"])def test_setup_error_is_2_not_4(api, tree, make_config, old, new, requests):assert sweep(tree, make_config(url=api.url.replace(old, new))) == 2assert api.requests == requests # stopped at once, never retrieddef test_refused_token_is_redacted(api, tree, make_config, capsys):config = make_config()(config.parent / "api-token").write_text("lab-token-revoked\n")assert sweep(tree, config) == 2err = capsys.readouterr().errassert "Unknown token [REDACTED]" in err and "lab-token-revoked" not in errdef test_dry_run_changes_nothing(api, tree, make_config, tmp_path):config = make_config()before = files_under(tmp_path)assert sweep(tree, config, "--dry-run", "--quarantine") == 1assert files_under(tmp_path) == beforedef test_quarantine_then_rerun_is_a_no_op(api, tree, make_config, tmp_path):config = make_config()assert sweep(tree, config, "--quarantine") == 1manifest = (tmp_path / "quarantine" / "manifest.ndjson").read_text()assert not (tree / "invoice.pdf").exists() and manifest.count("\n") == 1assert (tmp_path / "quarantine").stat().st_mode & 0o777 == 0o700assert sweep(tree, config, "--quarantine") == 0assert (tmp_path / "quarantine" / "manifest.ndjson").read_text() == manifestdef test_name_that_is_not_utf8(api, tree, make_config, tmp_path, capsys):odd = tree / os.fsdecode(b"caf\xe9.pdf") # a Latin-1 name, as an upload can haveodd.write_bytes((tree / "invoice.pdf").read_bytes())assert sweep(tree, make_config()) == 1paths = [json.loads(line)["path"] for line in capsys.readouterr().out.splitlines()]assert f"{tree}/caf\\xe9.pdf" in pathsassert "caf\\xe9.pdf,1f72" in (tmp_path / "summary.csv").read_text()def test_internal_error_is_70_not_findings(api, tree, make_config, monkeypatch, capsys):def broken(*args):raise RuntimeError("a bug")monkeypatch.setattr(cli, "check_tree", broken) # the name cli.sweep() looks upassert sweep(tree, make_config()) == 70assert "internal error" in capsys.readouterr().err
The setup errors stop after one request (the 404) or none (the TLS handshake), the Latin-1 name comes out as caf\xe9.pdf, and the forced bug ends with 70. Now the real thing. make_tree.sh builds a clean tree and an upload tree with two different samples called invoice.pdf, a file carrying the scanner's signature, a file whose name is not valid UTF-8, and a link to /etc/passwd:
#!/usr/bin/env bash# Builds the trees the sweep runs on, from scratch each time:# srv/clean three ordinary files# srv/uploads two known-bad samples that share a name, a file with the scanner's test# signature, ordinary files, a file whose name is not valid UTF-8, and a# symbolic link that points out of the treeset -euo pipefailrm -rf srvmkdir -p srv/clean/docs srv/uploads/alice srv/uploads/bob srv/uploads/sharedprintf 'quarterly report, nothing unusual\n' > srv/clean/docs/report.txtprintf 'meeting notes\n' > srv/clean/docs/notes.txtprintf 'backup done\n' > srv/clean/backup.logprintf 'ioc-sweep lab sample A: stands in for a malicious file\n' > srv/uploads/alice/invoice.pdfprintf 'ioc-sweep lab sample B: a different file with the same name\n' > srv/uploads/bob/invoice.pdfprintf 'macro loader text LAB-TEST-SIGNATURE end\n' > srv/uploads/bob/macro.docmprintf 'quarterly report, nothing unusual\n' > srv/uploads/shared/report.txtprintf 'meeting notes\n' > srv/uploads/shared/notes.txtprintf 'staff rota for October\n' > srv/uploads/shared/rota.txt# A name with byte 0xE9 (a Latin-1 "é"), which is not valid UTF-8: whoever uploads chooses the name.printf 'meeting notes\n' > "srv/uploads/shared/caf$(printf '\351').txt"ln -s /etc/passwd srv/uploads/shared/passwd# %q shows each name as Bash would quote it, so the byte 0xE9 appears as \351find srv -type f -print0 | sort -z | while IFS= read -r -d '' name; do printf '%q\n' "$name"; donefind srv -type l -printf '%p -> %l\n'
The lab CA comes from the HTTP lesson's script, unchanged, run in a subshell so that your shell stays in ~/ps-capstone. The token file is a 0600 copy of the token the mock API accepts, and the API starts in the background:
The clean tree exits 0 with one log line. The upload tree exits 1: two intelligence matches and one scanner finding on stdout, and the skipped link on stderr. The API log shows the first batch of three taking two pages (cursor c1) and the second one page. The other output formats:
The first block is the findings as CSV on stdout; the second is out/summary.csv, one row per file, with clean rows that prove those files were checked, including caf\xe9.txt, the name make_tree.sh printed as $'...\351...'. Next, the dry run, checked on the disk rather than trusted:
The plan named the summary and three moves, the listing of srv/ and out/ is unchanged, and no quarantine directory exists, yet the status is 1, as the real run returns. The real run stored all three files under hash-prefixed names with mode 0600 in a 0700 directory. The manifest is 0664 because the argparse lesson's open(..., "a") follows the umask; the 0700 directory is what keeps it private. The second run found only the clean files, exited 0 and added no manifest line. Now the failures:
With the API stopped, three attempts with jittered waits were refused, and the run ended with 4 and no findings. In partial mode the first batch answered and the second got 503 three times: the matches still reached stdout, the status is 3, and the summary marks two files incomplete with the detail lookup failed. The setup failures:
A typo, an abbreviated option (allow_abbrev=False on both parsers), a wrong API path, a URL whose name the certificate does not list and a token file readable by others each exit 2 with one line, after at most one request: a wrapper that follows the contract calls a person instead of retrying every hour. The revoked token got one 401, not retried; the API's reason phrase repeated the token, and the redaction filter printed [REDACTED] in its place.
The whole suite, a broken test, and the wheel
39 tests in about 25 seconds, mostly real waits on the mock API. Now a plausible edit: someone "simplifies" the type check to isinstance():
DID NOT RAISE means the block inside pytest.raises finished without the exception: attempts = true was accepted as the integer 1, and a sweep would silently try each request once. Restoring the line makes all eleven pass. Last, ship it. The wheel is built in the development venv, and production gets a new venv with only the wheel, installed with --no-index because nothing needs downloading:
env -i ran the launcher with an empty environment, less than cron provides, and it worked: the launcher names the venv's Python and the config uses absolute paths. A crontab line such as ioc-sweep scan ... > /var/tmp/ioc-sweep.ndjson would undo the Bash capstone: another account can create that fixed name first, the file follows cron's umask, > empties it so a reader sees half a result, and slow nights can overlap. nightly-sweep.sh applies the Bash capstone's answers:
#!/usr/bin/env bash# One scheduled sweep, run the way cron should run it: one run at a time, a time limit, and the# findings and the log kept in a directory that only this user can read.# Exit status: ioc-sweep's own (0, 1, 2, 3, 4, 70); 75 another sweep still holds the lock;# 124 the time limit was reached (137 if the sweep then had to be killed).set -euo pipefailumask 077dir=${1:?usage: nightly-sweep.sh DIR}root=$(cd -- "$(dirname -- "$0")" && pwd -P)results=$root/resultsmkdir -p -- "$results"# The script opens its own log only after mkdir: a crontab redirection into results/ would be# opened by the shell before this script runs, and fail while results/ does not exist yet.exec 2>> "$results/sweep.log"printf 'nightly-sweep: %s sweep of %s\n' "$(date -u +%FT%TZ)" "$dir" >&2exec 9> "$results/.lock"if ! flock -n 9; thenecho "nightly-sweep: another sweep still holds the lock" >&2exit 75fitmp=$(mktemp -- "$results/.findings.XXXXXX")trap 'rm -f -- "$tmp"' EXITstatus=0timeout -k 1m 6h "$root/prod/bin/ioc-sweep" scan "$dir" --config "$root/sweep.toml" > "$tmp" || status=$?case $status in0 | 1 | 3) mv -- "$tmp" "$results/findings.ndjson" ;; # a result, complete or partial: publish it*) echo "nightly-sweep: findings.ndjson not replaced" >&2 ;;esacecho "nightly-sweep: ioc-sweep exit status $status" >&2exit "$status"
umask 077 and a results/ directory in the project keep everything private. Once mkdir has made it, exec 2>> sends all later stderr, ioc-sweep's included, to results/sweep.log. flock -n allows one sweep at a time and exits 75 (EX_TEMPFAIL) otherwise. timeout -k 1m 6h bounds the whole run, including what no deadline inside the tool covers. The findings replace findings.ndjson in one mv, and only when the sweep produced a result (0, 1 or 3). Run it with an almost empty environment, then again while another process holds the lock:
The terminal shows only the exit status; the log has the rest. Every file is -rw-------.
Why the script opens its own log: a crontab line ending in >> "$HOME/ps-capstone/results/sweep.log" 2>&1 makes /bin/sh open that file before starting the script, and a fresh server has no results/ yet:
sh could not create the log, so the script never started: exit 2, still no results/. cron would mail that message, and a server without a mail program loses it. So the crontab line has no redirection. sweep.cron runs the job every minute for the lab; on a server use 15 2 * * * and the real share. Add it the Bash capstone's way, which keeps your existing lines, wait for cron (the mock API is still running; the wait gives up after 150 seconds), and remove it:
cron started the job with results/ missing: the script made it 0700, logged the run, published three findings and ended with 1. The diff proves your crontab is back. As a systemd service (a Type=oneshot unit started by a timer, see the Linux course's scheduling lesson), LoadCredential= delivers the token and TimeoutStartSec= replaces timeout: a oneshot unit counts as starting until its command exits, so RuntimeMaxSec= would have no effect on it (scripting-adv's lesson on signals shows both); alerting on a failed or missing run is in scripting-adv. Stop the mock API:
Try this
Write tests/test_throttle.py with two tests that use the api, tree and make_config fixtures with api.mode = "throttle". With the test settings, main() waits out the one-second Retry-After, exits 1 and the API counts three requests. With make_config(max_retry_after=0.5), the first batch is given up, the run exits 3, and stderr contains Retry-After 1s is too long. Both pass. Then break retry.py (keep a copy): delete the two lines that compare err.retry_after with max_retry_after. The second test fails with assert 1 == 3; the captured log shows the tool sleeping the full second. Restore the file and both pass.
Takeaway
A tool that runs unattended must say exactly what it proved: check every setting before the first request, count a failed lookup as unknown rather than clean, keep every change behind one dry-run check, and give each kind of failure its own exit status, so the scheduler knows whether to wait, retry or call a person.
Next: Advanced scripting for DevSecOps.
check_tree() raises ApiUnavailable, and the lab run ended with 4 and no findings.api.attempts and api.deadline whatever the result; the lab's down run gave up after three attempts.ioc-sweep stop at once with exit 2 instead of retrying and reporting 4?AuthError is the tool's own class, raised in get_json(). Any exception can be caught and retried; the question is whether it should be.walk.py passes on_error=cannot_list to Path.walk(). What would happen without it when one subdirectory of the tree has mode 000?Path.walk() ignores such errors by default. Nothing is raised, which is the danger.000 directory is not a link, and walk() follows no links by default. Its contents stay unread.on_error, the test showed the directory recorded as a problem, which makes the run partial.Permission denied.