Capstone: a certificate-expiry and rotation runner
A Bash entry point, a packaged Python core, a fleet, a flaky CA API and release gates.
tar -xzf scr-capstone.tar.gz, which creates scr-capstone/. SHA-256: d82da6e9216f939f5c8e5f6b05bcccaa7694182d418a0ec532cb3cbb994b65a4This capstone builds certrun, a job that finds TLS certificates about to expire and replaces them, and runs it the way production would: installed root-owned, started by systemd as an unprivileged service account, with nobody watching. You will read the parts that are new here, test them, build, sign and install a release, and run it end to end: a plan, a run stopped by SIGTERM in a CA call, a run stopped by systemctl stop in a deploy, the run that finishes both and rolls back a bad certificate, a run with nothing left to do, and a tampered release that is refused.
Refresher: the entry point is from "When a script has outgrown Bash", deadlines and signals from "Signals, timeouts and shutdown", typed settings from "Designing a CLI tool", retries and idempotency keys from "Failure design", pinned SSH from "Fleet automation", logs, metrics and audit lines from "Observable jobs", and the lock file, SBOM and signature from the last two lessons.
What one run does, and what it promises
certrun.sh plan|apply DESTINATION is the only command systemd knows; a destination is a set of hosts and endpoints with its own settings file. The service account is the unprivileged account the job runs as. In this lab your own account plays it; in production it is a system account with no login shell.
The exit contract is the interface to everything that watches the job. 0: every endpoint is fine or was rotated. 1: some endpoints need a person (or the endpoint metrics could not be written), everything else was done. 75 (EX_TEMPFAIL): only temporary failures, such as a host that did not answer; run again later. 78 (EX_CONFIG): fix the settings, retrying is pointless. 70 (EX_SOFTWARE): a bug outside any one endpoint's work. The entry point adds 2 for usage, 75 for a held lock or a spent deadline, and 77 (EX_NOPERM) when it refuses the core. 70, 75, 77 and 78 come from sysexits.h; 0, 1 and 2 are the shell's conventions.
The lab, and the entry point
What you need: Ubuntu 26.04 with sudo; openssh-server, jq, curl, and bats with bats-support and bats-assert (as in "Testing shell automation"); network access to GitHub releases and PyPI; free ports 18941-18946, 18951-18952 and 18961; membership in adm or systemd-journal to read the journal. Unpack the lesson files in your home directory (cd ~ && tar -xzf scr-capstone.tar.gz && cd scr-capstone); every terminal runs in ~/scr-capstone. No file names an account: the scripts that need one take $USER.
The lab stands in for the outside world; its scripts are in the lesson files and are not walked through. lab/get-tools.sh downloads uv 0.12.19, cosign v3.1.3 and Syft 1.52.0 into bin/, each checked against a pinned SHA-256. lab/hosts.sh (as root) plays the hosts web1 and web2: two throwaway sshd units on 127.0.0.1 (never port 22) with their keys and pinned known_hosts in root's /run/scr-capstone-sshd, and /srv/scr-capstone (yours) for the hosts' certificate files. lab/make-pki.sh makes a throwaway CA, a certificate per endpoint with chosen dates, the CA API's certificate, a test token, two CAs for failure cases and, for the tests, a valid certificate under the expired CA: 8 certificates in all. lab/ca_api.py is the CA's API, which fails on purpose. lab/tlsfarm.py is the hosts' web server: like nginx, it loads certificate files at start and on SIGHUP, never per connection, and a reload takes effect 2 seconds after the signal.
Six endpoints, one path each: api has 200 days left, billing 12, legacy expired three days ago, shop has 20, admin serves a certificate for the wrong name (www), and search lives on web2. The CA's faults give legacy a 500 then a slow answer, billing a 429 then a slow answer, and shop a certificate from an issuer nobody trusts. In the last step, ssh-keygen -lf prints key fingerprints and ssh-keyscan fetches a server's host key without logging in: lines 1-2 are the pins, lines 3-4 what the hosts offer now. web2 offers a key other than its pin, the key it had before it was rebuilt; certrun cannot tell a rebuild from an attacker in the middle, so it must refuse. sudo lab/hosts.sh stop undoes the hosts (see Clean up).
certrun.sh owns what Bash does well. It checks its arguments (the destination is a name, never a path), takes an flock on /var/lib/scr-capstone/state/DESTINATION.lock so two runs for one destination never overlap, and makes a private run directory. The EXIT trap removes that directory on every path and, for an apply that got the lock, records when the run ended and how. A run refused at the lock did nothing, so it records nothing: the metric's age shows that no run gets through. The TERM and INT traps record the signal and pass a stop request to the core; once the core has finished, reraise sends the same signal to the script itself, so a supervisor sees the death it asked for.
Then the gate (explained with the release), and the job file, written with jq --arg, which escapes each value. The core runs under timeout with what is left of the budget, with only HOME, LANG and PATH in its environment and python -I -B (the -B is explained with the release). The last case maps its status: the core's contract passes through, 124 is a spent deadline (75), and 137 is a deadline only if the deadline has passed; a core SIGKILLed earlier, say by the out-of-memory killer, is 70.
#!/usr/bin/env bash# certrun.sh plan|apply DESTINATION: the entry point the systemd unit runs as the service account.# It owns the process: one run per destination, a deadline for the whole run, signals, cleanup,# the run-level metrics, and a gate that refuses a core that is not the signed release.# Exit status: the core's contract passed through (0 ok, 1 a person must look at an endpoint or# at the metrics, 70 internal, 75 temporary, 78 config), plus 2 usage, 75 lock held or deadline# reached, 77 core refused, 78 no settings or a bad CERTRUN_BUDGET, 70 core killed. Stopped by# SIGTERM, it cleans up and dies of SIGTERM (143 in a shell).# CERTRUN_BUDGET: seconds for the whole run (default 60). CERTRUN_VAR: state root (for tests).set -Eeuo pipefailshopt -s inherit_errexithere=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd -P) # /opt/scr-capstone: root'srelease=$here/releasewheel=$release/certrun-0.1.0-py3-none-any.whlvar=${CERTRUN_VAR:-/var/lib/scr-capstone} # the service account's: state/ and textfile/budget=${CERTRUN_BUDGET:-60}mode='' dest='' run='' child='' stop='' locked=''log() { # log LEVEL MESSAGE: one JSON line on stderr, like the core'sjq -ncM --arg level "$1" --arg msg "$2" --arg dest "$dest" \'{ts: (now | todate), level: $level, logger: "certrun.sh", msg: $msg, destination: $dest}' >&2}# shellcheck disable=SC2329 # called from on_exit, the EXIT traprun_metrics() { # run_metrics STATUS: when the last apply ended and how, rewritten in one renamelocal file=$var/textfile/certrun-$dest-run.prom label="{destination=\"$dest\"}" tmptmp=$(mktemp "$file.XXXXXX") || returnprintf '%s\n' "# TYPE certrun_last_run_seconds gauge" "certrun_last_run_seconds$label $(date +%s)" \"# TYPE certrun_last_run_status gauge" "certrun_last_run_status$label $1" >"$tmp"chmod 0644 "$tmp" && mv -f -- "$tmp" "$file"}# shellcheck disable=SC2329 # called from the EXIT trapon_exit() { # every way out: the run metrics (an apply that got the lock), then the run directorylocal status=$1trap '' TERM INT # a second signal must not cut this short[[ -z $stop ]] || status=$((128 + $(kill -l "$stop")))[[ $mode != apply || -z $locked ]] || run_metrics "$status" || log warning "run metrics not written"[[ -z $run ]] || rm -rf -- "$run"}# shellcheck disable=SC2329 # called from the TERM and INT trapson_signal() { # remember the signal and pass a stop request on to the core, if it runsstop=$1[[ -z $child ]] || kill -TERM "$child" 2>/dev/null || true}reraise() { # die of the signal we were sent; the EXIT trap still runs firstlog warning "stopped by SIG$stop"trap - "$stop"kill -s "$stop" "$$"}usage() {echo "usage: certrun.sh plan|apply DESTINATION" >&2exit 2}trap 'on_exit $?' EXITtrap 'on_signal TERM' TERMtrap 'on_signal INT' INT(($# == 2)) || usage[[ $1 == plan || $1 == apply ]] || usage[[ $2 =~ ^[a-z0-9][a-z0-9-]{0,62}$ ]] || usage # a name, never a pathmode=$1 dest=$2config=$here/etc/$dest.toml[[ -f $config ]] || { log error "no settings file $config"; exit 78; }[[ $budget =~ ^[1-9][0-9]*$ ]] || { log error "CERTRUN_BUDGET must be whole seconds"; exit 78; }deadline=$((SECONDS + budget))[[ -w $var/state ]] || { log error "$var/state is missing or not writable: see install.sh"; exit 78; }# One run per destination. The core inherits this descriptor on purpose: if this script is# killed, the lock stays held until the core has gone too. A run refused here did nothing, so it# leaves the run metrics alone (their age shows that no run gets through).exec {lock}>>"$var/state/$dest.lock"flock -n "$lock" || { log warning "another run holds the lock"; exit 75; }locked=yesrun=$(mktemp -d -t "certrun-$dest.XXXXXX")# The gate: verify the signature on a private copy of the wheel with the pinned key, then check# the installed core against that copy, with the system python (-I -S: no site, so nothing in the# venv runs before the check). set -e is off inside a function called in a condition: check each.verify_core() {cp -- "$wheel" "$run/core.whl" || return 1cp -- "$wheel.sigstore.json" "$run/core.bundle" || return 1"$here/bin/cosign" verify-blob --key "$here/etc/cosign.pub" --insecure-ignore-tlog=true \--bundle "$run/core.bundle" "$run/core.whl" || return 1/usr/bin/python3 -I -S "$here/verify_core.py" "$run/core.whl" "$release/venv"}if ! verify_core >"$run/gate.out" 2>&1; thenlog error "refusing to run: the core is not the signed release"cat -- "$run/gate.out" >&2exit 77fi[[ -z $stop ]] || reraisejq -n --arg mode "$mode" --arg destination "$dest" --arg config "$config" --arg var "$var" \--arg run_id "$(date -u +%Y%m%dT%H%M%SZ)-$$" \'{mode: $mode, destination: $destination, config: $config, run_id: $run_id, var: $var}' >"$run/job.json"left=$((deadline - SECONDS))((left > 0)) || { log error "deadline reached before the core started"; exit 75; }# Only HOME, LANG and PATH reach the core; -I: no PYTHONPATH, no user site, no current# directory on sys.path. The bytecode is root's, compiled by install.sh next to the source; -B# writes none at run time (PYTHONDONTWRITEBYTECODE would not do: -I ignores PYTHON* variables).timeout -k 10 "$left" env -i HOME="$HOME" LANG=C.UTF-8 PATH=/usr/bin:/bin \"$release/venv/bin/python" -I -B -m certrun "$run/job.json" &child=$![[ -z $stop ]] || kill -TERM "$child" # a signal that came before $! was knownstatus=0wait "$child" || status=$?while kill -0 "$child" 2>/dev/null; do # a trapped signal ends wait early: wait for the corestatus=0wait "$child" || status=$?donechild=[[ -z $stop ]] || reraise # the core has finished its own cleanup by nowcase $status in0 | 1 | 70 | 75 | 78) ;; # the core's own contract124)log error "deadline of ${budget}s reached: the core was stopped"status=75;;137)if ((SECONDS >= deadline)); thenlog error "deadline of ${budget}s reached: the core ignored SIGTERM and was killed"status=75elselog error "the core was killed with SIGKILL before the deadline (out of memory?)"status=70fi;;*)log error "the core ended with status $status, which is not in its contract"status=70;;esacexit "$status"
Two details matter under stress. The trap does not exit: the while kill -0 loop waits for the core's cleanup, because a trapped signal ends wait early. And the script records the signal itself, because timeout's status cannot say who stopped the core: Ubuntu 26.04's uutils timeout -k exits 124 when it is itself sent SIGTERM (143 without -k), while GNU timeout passes on the child's status; the lab checked each case. tests/certrun.bats (in the lesson files) tests the entry point with a fake release: tests/fake/cosign passes unless FAKE_COSIGN_FAIL is set, and tests/fake/python plays the core as each test tells it.
Captured without a terminal (--tap). Test 6: a core SIGKILLed right away is 70, not a deadline. Test 7: SIGTERM mid-run gives 143 (128 + 15), the core ran its handler, the run directory is gone, the lock is free, and the metric says 143. Tests 1, 4 and 8: every apply that gets the lock records its status, 77 included; an unknown destination, a run refused at the lock and a plan record nothing.
The core: what it sees, and how it deploys
The core is a package, certrun, with two dependencies: paramiko for SSH and cryptography for X.509. pyproject.toml and the lab's uv.lock are in the lesson files, so your dependency tree matches the one below. The settings for the destination lab (44 of 79 lines; the other five endpoints look like api):
Relative paths are relative to this file, which is installed root-owned in /opt/scr-capstone/etc next to the keys it names. Each host's reload command runs over SSH after a deploy. Where state and metrics live belongs to the installation, so certrun.sh passes it in the job.
These modules are in the lesson files. config.py and loader.py: the dataclasses and loader of "Designing a CLI tool", now resolving relative paths and requiring an absolute pem_path. retry_after.py: the parser of "Failure design", unchanged (isascii() before isdigit(), because '²'.isdigit() is true). ssh.py: the pinned-key connect of "Fleet automation", with a host that is down or slow as TransientError and a key mismatch, refused login or failing command as TargetError. obs.py: the JSON logs, textfile writer and fsynced audit append of "Observable jobs". errors.py: those errors, ConfigError, Stopped and the exit codes. tlscheck.py: one verified handshake per endpoint, bounded by a semaphore and asyncio.timeout, and on failure a non-verifying one that only reads the certificate. cli.py: the job file and the exit statuses. It prints the results table before it writes the endpoint metrics, so a failed write (a full disk, a missing textfile/) cannot hide what the run did: it logs "endpoint metrics not written" and the run exits 1.
remote.py works on a host, each step over its own SSH connection. The install is a short script in the remote shell with the new PEM on standard input: it keeps a backup of the old file (once per rotation, so a resumed rotation never backs up its own new file), starts the new file as a copy of the old one so owner and mode stay what the server expects, and replaces the old one in one rename. The server finds the old pair or the new one, never half of each, which is why key and certificate share a file. A symlink is refused: the rename would replace the link, not the file the server reads.
"""What certrun does on a host, each over its own SSH connection: read the certificate in everyPEM file, replace one file with a backup, put the backup back, and have the server reload.Blocking (paramiko), so the async code calls these through asyncio.to_thread."""import shlexfrom .config import Config, Endpoint, Hostfrom .errors import TargetError, TransientErrorfrom .ssh import connect, run# Runs in the remote login shell with the new PEM on stdin. The old file is kept as BACKUP (once# per rotation: a resumed rotation must not back up its own new file). The new file starts as a# copy of the old one, so it keeps the owner and mode the server expects, and replaces it in one# rename(2): the server finds the old pair or the new one, never half of each. A symlink is# refused, because the rename would replace the link instead of the file the server reads.INSTALL = """set -ef={path} b={backup}if [ -L "$f" ]; then echo "$f is a symlink: not replacing it" >&2; exit 1; fi[ -e "$b" ] || cp -p -- "$f" "$b"cp -p -- "$f" "$f.certrun-new"cat >"$f.certrun-new"mv -f -- "$f.certrun-new" "$f""""def deployed(host: Host, cfg: Config) -> dict[str, str | Exception]:"""The SHA-256 of the certificate in each of this host's PEM files, or why it is unknown."""found: dict[str, str | Exception] = {}with connect(host, cfg) as client:for ep in (e for e in cfg.endpoints if e.host == host.name):try:out = run(client, f"openssl x509 -noout -fingerprint -sha256 -in {shlex.quote(ep.pem_path)}", cfg)found[ep.name] = out.strip().split("=", 1)[1].replace(":", "").lower()except (TargetError, TransientError) as err:found[ep.name] = errexcept (IndexError, ValueError):found[ep.name] = TargetError(f"{ep.pem_path}: unexpected openssl output")return founddef install(host: Host, ep: Endpoint, pem: bytes, backup: str, cfg: Config) -> None:command = INSTALL.format(path=shlex.quote(ep.pem_path), backup=shlex.quote(backup))with connect(host, cfg) as client:run(client, command, cfg, data=pem)def rollback(host: Host, ep: Endpoint, backup: str, cfg: Config) -> None:"""Put back the file install() replaced, again in one rename."""with connect(host, cfg) as client:run(client, f"mv -f -- {shlex.quote(backup)} {shlex.quote(ep.pem_path)}", cfg)def forget(host: Host, backup: str, cfg: Config) -> None:"""The backup holds the old private key: remove it once the new certificate is live."""with connect(host, cfg) as client:run(client, f"rm -f -- {shlex.quote(backup)}", cfg)def reload(host: Host, cfg: Config) -> None:with connect(host, cfg) as client:run(client, host.reload, cfg)
decide.py turns what the survey saw into an action, with no network and no side effects. The host comes first: never issue a certificate you cannot deliver. Then any record a stopped run left (next section). Then the file on disk must be what the endpoint serves; if not, a reload is pending or failed, and another certificate would not help. Expiry comes from the leaf's own not_after, not from the verify code: code 10, "certificate has expired", is also what an expired CA above a valid leaf gives.
"""decide(): what to do with one endpoint, from what it serves, what its host has on disk and therecord of any rotation a stopped run left behind. No network and no side effects."""from dataclasses import dataclassfrom datetime import datetimefrom .config import Config, Endpointfrom .errors import TransientErrorfrom .pending import Pendingfrom .tlscheck import EXPIRED, Seen@dataclassclass Result:endpoint: straction: str # ok, rotate, finish, rollback (plan); rotated, failed (apply)detail: strtemporary: bool = False # failed, and the next run may succeedseen: Seen | None = Nonerecord: Pending | None = Nonenot_after: datetime | None = Nonedef failed(ep: Endpoint, err: Exception, seen: Seen | Exception) -> Result:not_after = seen.not_after if isinstance(seen, Seen) else None # still worth a metricreturn Result(ep.name, "failed", str(err), temporary=isinstance(err, TransientError), not_after=not_after)def decide(ep: Endpoint, seen: Seen | Exception, disk: str | Exception, rec: Pending | None,cfg: Config) -> Result:for outcome in (disk, seen): # the host first: no certificate for a host we cannot trustif isinstance(outcome, Exception):return failed(ep, outcome, seen)assert isinstance(seen, Seen) and isinstance(disk, str)def act(action: str, detail: str) -> Result:return Result(ep.name, action, detail, seen=seen, record=rec, not_after=seen.not_after)days = seen.days_left()if rec and rec.rolled_back:return act("failed", f"a deploy was rolled back ({rec.rolled_back}); fix the cause,"f" then delete pending/{ep.name}.json")if rec and rec.deployed and seen.key == rec.key: # deployed, never checked: check it nowif not seen.problem:return act("finish", "a stopped run deployed the new certificate; the audit line is missing")return act("rollback", f"a stopped run deployed a certificate that fails its check: {seen.problem}")if rec and rec.replacing in (seen.fingerprint, disk):return act("rotate", "resume the rotation a stopped run began")if rec:return act("failed", f"pending/{ep.name}.json replaces {rec.replacing[:12]}, which is gone")if disk != seen.fingerprint:return act("failed", f"{ep.host} has {disk[:12]} on disk but serves {seen.fingerprint[:12]}:"" a reload is pending or failed")if seen.problem and seen.code != EXPIRED: # a wrong name, an unknown issuer: not ours to fixreturn act("failed", f"serves a certificate that does not verify: {seen.problem}")if seen.problem and not seen.expired: # code 10, but the leaf itself is validreturn act("failed", f"the chain does not verify ({seen.problem}), but the certificate itself"f" has {days} days left: an issuer's certificate expired")if seen.expired:return act("rotate", f"EXPIRED {-days} days ago")return act("rotate" if days <= cfg.policy.renew_within_days else "ok", f"{days} days left")
14 packages: certrun, its seven runtime dependencies, and pytest with its own (the development group). The 24 tests run the CA API and TLS endpoints in-process on free ports (tests/conftest.py); test_an_expired_ca_does_not_make_a_valid_leaf_expired is the code-10 rule, and test_cli.py holds the metrics rules of the end-to-end section. SSH is not faked: the end-to-end runs use real sshd.
Rotation that survives a stop
A rotation spans several systems, and a stop can come between any two steps. The rule that keeps it safe: write down what you are about to do before you do it. pending.py saves a record before the CA is asked: the new private key, the exact request bytes and a random idempotency key, 0600 in the 0700 state directory, fsynced with its directory. Just before the deploy, the record is marked deployed: the new file may be live, but nobody has checked it. The next run takes one of four paths:
"""A rotation in progress, written to disk before the CA is asked: the new private key, the exactrequest and its idempotency key, and, from just before the deploy, that the new file may be livebut is not yet checked. A run stopped anywhere after that leaves the record, and the next runresends the same bytes (the CA answers from its own record) or checks the deployed certificate:the missing audit line if it verifies, a rollback if it does not.One file per endpoint, 0600 in a 0700 directory."""import dataclassesimport hashlibimport jsonimport osimport secretsfrom datetime import UTC, datetimefrom pathlib import Pathfrom cryptography import x509from cryptography.hazmat.primitives import hashes, serializationfrom cryptography.hazmat.primitives.asymmetric import ecfrom cryptography.x509.oid import NameOIDfrom .config import Endpointfrom .tlscheck import DER, SPKI, SeenPEM = serialization.Encoding.PEM@dataclasses.dataclass(frozen=True)class Pending:endpoint: strreplacing: str # fingerprint of the certificate being replacedreplacing_key: str # and of its public key: the new key must differkey: str # SHA-256 of the new public keykey_pem: str # the new private key (unencrypted, like the PEM the server loads)request: str # the request body, resent byte for byteidempotency_key: strcreated: strdeployed: bool = False # the new file may be on the host: deployed, not yet checkedrolled_back: str | None = None # why a deploy of this certificate was undonedef start(state: Path, ep: Endpoint, seen: Seen) -> Pending:"""A new key and request for ep, saved before anything is sent."""key = ec.generate_private_key(ec.SECP256R1()) # a new key for every rotation, never reusedsubject = x509.Name([x509.NameAttribute(NameOID.COMMON_NAME, ep.server_name)])csr = (x509.CertificateSigningRequestBuilder().subject_name(subject).add_extension(x509.SubjectAlternativeName([x509.DNSName(ep.server_name)]), critical=False).sign(key, hashes.SHA256())) # ECDSA signatures are random: a new CSR never repeatsrec = Pending(endpoint=ep.name, replacing=seen.fingerprint, replacing_key=seen.key,key=hashlib.sha256(key.public_key().public_bytes(DER, SPKI)).hexdigest(),key_pem=key.private_bytes(PEM, serialization.PrivateFormat.PKCS8,serialization.NoEncryption()).decode(),request=json.dumps({"name": ep.name, "csr": csr.public_bytes(PEM).decode()}),idempotency_key="certrun-" + secrets.token_hex(16),created=datetime.now(UTC).isoformat(timespec="seconds"))save(state, rec)return recdef load(state: Path, name: str) -> Pending | None:try:return Pending(**json.loads((state / "pending" / f"{name}.json").read_text()))except FileNotFoundError:return Nonedef save(state: Path, rec: Pending) -> None:folder = state / "pending"folder.mkdir(mode=0o700, parents=True, exist_ok=True) # parents: 0700 state from install.shtmp = folder / f".{rec.endpoint}.tmp"with os.fdopen(os.open(tmp, os.O_WRONLY | os.O_CREAT | os.O_TRUNC, 0o600), "w") as f:json.dump(dataclasses.asdict(rec), f)f.flush()os.fsync(f.fileno())os.replace(tmp, folder / f"{rec.endpoint}.json")sync_dir(folder) # the new name is on disk too, not only the bytesdef drop(state: Path, name: str) -> None:(state / "pending" / f"{name}.json").unlink()sync_dir(state / "pending")def sync_dir(folder: Path) -> None:fd = os.open(folder, os.O_RDONLY | os.O_DIRECTORY)try:os.fsync(fd)finally:os.close(fd)
Why the exact bytes: a certificate request (CSR) is signed, and ECDSA signatures are random, so rebuilding it gives new bytes. A CA that follows the IETF Idempotency-Key draft fingerprints the whole request and answers 422 to a known key with another body; the lab CA does, and a test proves that a rebuilt CSR is refused. The key sits unencrypted, like the PEM the server loads, only until the rotation finishes; generating it on the target host, so only the CSR travels, is the production refinement.
ca.py retries a 429, 500, 502, 503 or 504, network errors and timeouts: after exactly the Retry-After the server named, unless it exceeds max_retry_after, which gives up and leaves the endpoint to the next run; otherwise after a jittered backoff. A CA certificate that does not verify, and any other status, is not retried. Redirects are refused: urllib would follow a 302 as a GET and copy the Authorization header to wherever Location points, plain HTTP included (a test proves both), so NoRedirects returns None and the 3xx becomes an error.
"""The CA API client. Every attempt of one rotation sends the same Idempotency-Key with the samerequest bytes, so a request whose answer was lost is answered from the CA's record instead ofissuing a second certificate. Redirects are refused: the bearer token only ever goes to ca.url."""import asyncioimport http.clientimport jsonimport loggingimport randomimport sslimport urllib.errorimport urllib.requestfrom .config import CAfrom .errors import TargetError, TransientErrorfrom .retry_after import parse_retry_afterlog = logging.getLogger("certrun.ca")RETRYABLE = {429, 500, 502, 503, 504}class NoRedirects(urllib.request.HTTPRedirectHandler):def redirect_request(self, req, fp, code, msg, headers, newurl): # type: ignore[no-untyped-def]return None # urllib then raises HTTPError for the 3xx, and the token stays heredef post(ca: CA, token: str, key: str, body: str) -> str:"""One attempt; blocking. The certificate PEM, or an error that says what should happen next."""https = urllib.request.HTTPSHandler(context=ssl.create_default_context(cafile=ca.ca_file))opener = urllib.request.build_opener(NoRedirects, https)req = urllib.request.Request(f"{ca.url}/certificates", data=body.encode(), method="POST",headers={"Content-Type": "application/json", "Authorization": f"Bearer {token}","Idempotency-Key": key})try:with opener.open(req, timeout=ca.timeout) as resp:answer = json.load(resp)except urllib.error.HTTPError as err:if err.code in RETRYABLE:raise TransientError(f"HTTP {err.code}", parse_retry_after(err.headers["Retry-After"])) from Nonemoved = f" to {err.headers['Location']}, not followed" if 300 <= err.code < 400 else ""raise TargetError(f"the CA answered HTTP {err.code}{moved}") from Noneexcept urllib.error.URLError as err:if isinstance(err.reason, ssl.SSLError): # the CA's certificate does not verify: not transientraise TargetError(f"CA TLS: {err.reason}") from Noneraise TransientError(repr(err.reason)) from Noneexcept (OSError, http.client.HTTPException, ValueError) as err: # timeouts, resets, a garbled bodyraise TransientError(repr(err)) from Noneif not isinstance(answer, dict) or not isinstance(answer.get("pem"), str):raise TargetError("the CA's answer holds no certificate")return answer["pem"]async def issue(ca: CA, token: str, key: str, body: str, name: str) -> str:for attempt in range(1, ca.attempts + 1):try:return await asyncio.to_thread(post, ca, token, key, body)except TransientError as err:if attempt == ca.attempts:raisewait = random.uniform(0, 0.5 * 2**attempt) # full jitter: up to 1 s, 2 s, 4 sif err.retry_after is not None:if err.retry_after > ca.max_retry_after: # give up now; the next run tries againraise TransientError(f"{err}, Retry-After {err.retry_after:.0f} s is too long") from Nonewait = err.retry_after # the server said when: not soonerlog.warning("CA call failed, retrying", extra={"fields": {"endpoint": name, "error": str(err), "attempt": attempt, "wait_s": round(wait, 1)}})await asyncio.sleep(wait)raise AssertionError("not reached: attempts is at least 1")
rotate.py resumes the record or starts one, asks the CA, and checks that the certificate carries the recorded key. After the install and the reload it asks the endpoint what it serves, every half second for up to reload_wait seconds. If the new certificate is served and verifies, finish() removes the backup (it holds the old private key), writes the audit line, then deletes the record; a stop in between leaves the record, so the next run finishes the job. If not, roll_back() marks the record rolled back first (so no later run asks the CA again), renames the backup into place, reloads and checks. A stop after the reload but before the check leaves a deployed record, and the next run checks what is served: the audit line if it verifies, roll_back() if not.
"""Rotate one endpoint: never issue twice, never deploy blind. The request is on disk before it issent (pending.py); the certificate must match the recorded key; the record says "deployed, not yetchecked" before the deploy, which keeps a backup; the server is reloaded and checked; a failedcheck is rolled back; only then is the change audited."""import asyncioimport dataclassesimport hashlibimport loggingfrom typing import NoReturnfrom cryptography import x509from . import ca, obs, pending, remote, tlscheckfrom .config import Config, Endpointfrom .errors import TargetError, TransientErrorfrom .pending import Pendingfrom .tlscheck import DER, Seenlog = logging.getLogger("certrun")async def rotate(ep: Endpoint, rec: Pending | None, seen: Seen, cfg: Config, token: str) -> Seen:"""rec: the record a stopped run left, resumed as it is; None starts a new rotation."""rec = rec or pending.start(cfg.paths.state, ep, seen)if rec.key == rec.replacing_key:raise TargetError("the new key is the key being replaced; not requesting")pem = await ca.issue(cfg.ca, token, rec.idempotency_key, rec.request, ep.name)cert = x509.load_pem_x509_certificate(pem.encode())if tlscheck.key_id(cert) != rec.key:raise TargetError("the CA returned a certificate for another key; not deployed")new = hashlib.sha256(cert.public_bytes(DER)).hexdigest()host = cfg.hosts[ep.host]rec = dataclasses.replace(rec, deployed=True) # from here the new file may be live, uncheckedpending.save(cfg.paths.state, rec)await asyncio.to_thread(remote.install, host, ep, (pem + rec.key_pem).encode(), backup_of(ep, rec), cfg)await asyncio.to_thread(remote.reload, host, cfg)try:now = await served(ep, cfg, new)why = now.problem or (None if now.fingerprint == new else "the reload did not take effect")except TransientError as err: # no answer at all after the reloadwhy = str(err)if why is None:await finish(ep, rec, now, cfg)return nowawait roll_back(ep, rec, why, cfg)async def roll_back(ep: Endpoint, rec: Pending, why: str, cfg: Config) -> NoReturn:"""Put the backup back, reload, check. Marked first, so no later run asks the CA again forthis rotation: a person decides. Also what the next run does with a failing deployed record."""pending.save(cfg.paths.state, dataclasses.replace(rec, rolled_back=why))host = cfg.hosts[ep.host]await asyncio.to_thread(remote.rollback, host, ep, backup_of(ep, rec), cfg)await asyncio.to_thread(remote.reload, host, cfg)back = await served(ep, cfg, rec.replacing)state = "rolled back" if back.fingerprint == rec.replacing else "ROLLBACK FAILED"raise TargetError(f"the new certificate failed its check ({why}); {state}, serves {back.fingerprint[:12]}")def backup_of(ep: Endpoint, rec: Pending) -> str:return f"{ep.pem_path}.certrun-{rec.idempotency_key[-8:]}" # one backup per rotationasync def served(ep: Endpoint, cfg: Config, want: str) -> Seen:"""What ep serves once a reload has taken effect: checked every half second until it serves`want`, for at most policy.reload_wait seconds (a graceful reload takes a moment)."""loop = asyncio.get_running_loop()end = loop.time() + cfg.policy.reload_waitwhile True:seen = await tlscheck.check(ep, cfg, asyncio.Semaphore(1))if seen.fingerprint == want or loop.time() >= end:return seenawait asyncio.sleep(0.5)async def finish(ep: Endpoint, rec: Pending, now: Seen, cfg: Config, recovered: bool = False) -> None:"""The new certificate is live and verified: remove the backup (it holds the old private key),record the change, then delete the record. Anything that stops this half way leaves the record,and the next run finishes the job: the backup is never left without a record naming it."""try:await asyncio.to_thread(remote.forget, cfg.hosts[ep.host], backup_of(ep, rec), cfg)except (TargetError, TransientError) as err:raise type(err)(f"the new certificate is live; the old key file stays until the next run: {err}") from Noneobs.audit(cfg.paths.audit, {"action": "rotate", "endpoint": ep.name, "host": ep.host, "old": rec.replacing,"new": now.fingerprint, "old_key": rec.replacing_key, "new_key": rec.key,"not_after": now.not_after.isoformat(), "idempotency_key": rec.idempotency_key,**({"recovered": True} if recovered else {})})pending.drop(cfg.paths.state, ep.name)log.info("rotated, audit line written", extra={"fields": {"endpoint": ep.name, "recovered": recovered}})
Two more modules are in the lesson files. runner.py surveys every endpoint and host at once in a TaskGroup, then acts one endpoint at a time, most urgent first; any exception in one endpoint's work, even an unexpected one, becomes that endpoint's failure, so 70 is left for bugs outside endpoint handling. stop.py handles SIGTERM the asyncio way. "Failure design" raised from a signal.signal handler, but in an event loop that exception would surface inside the loop's own code; so the handler is registered with loop.add_signal_handler and cancels the main task, and CancelledError arrives at the current await. A to_thread call cannot be cancelled, so asyncio.run waits for it, bounded by that call's timeouts (a CA read 2 s; one SSH call about 25 s: connect, banner and login 5 s each, then 5 s per idle read of its output and of its errors) and finally by TimeoutStopSec.
Release, install and the unit
release.sh builds the wheel, installs it and the hash-checked locked dependencies into a relocatable release/venv as copies, signs the wheel with cosign and writes an SBOM. The signing key is made for this release in a temporary directory and deleted when the script ends; only cosign.pub survives, next to release/, not in it. The offline signing of "Release gates" carries the same caveat (a transparency log is the production form). In production the key lives in a KMS or HSM, or signing is keyless: never on the runner.
#!/usr/bin/env bash# Build one release of the certrun core into release/: the wheel and its signature bundle, a venv# with the locked dependencies and the wheel, and an SBOM of that venv. The signing key is made# for this one release in a private temporary directory and deleted when the script ends; only# its public key survives, as cosign.pub next to release/ (not in it), for install.sh to pin.# In production the key lives in a KMS or HSM, or signing is keyless: never on the runner.# Needs bin/uv, bin/cosign and bin/syft (lab/get-tools.sh).set -euo pipefailcd "$(dirname "$0")"whl=certrun-0.1.0-py3-none-any.whl[[ ! -e release ]] || { echo "release.sh: release/ exists; a release is never overwritten" >&2; exit 1; }keys=$(mktemp -d)trap 'rm -rf -- "$keys"' EXIT./bin/uv lock --check -q # the lock still matches pyproject.tomlrm -rf dist./bin/uv build -q --wheel # dist/$whl./bin/uv export -q --locked --no-dev --no-emit-project -o dist/requirements.txt # with hashes# Every file is copied (no links into uv's cache), so the release owns its bytes; the venv is# relocatable because install.sh moves it to /opt../bin/uv venv -q --relocatable --python "$(command -v python3)" release/venv./bin/uv pip install -q --link-mode=copy --python release/venv/bin/python \--require-hashes -r dist/requirements.txt./bin/uv pip install -q --link-mode=copy --python release/venv/bin/python --no-deps "dist/$whl"cp "dist/$whl" release/# The password protects the key file for the few seconds it exists; only cosign sees it.pass=$(openssl rand -hex 16)COSIGN_PASSWORD=$pass ./bin/cosign generate-key-pair --output-key-prefix "$keys/cosign" 2>/dev/nullCOSIGN_PASSWORD=$pass ./bin/cosign sign-blob --yes --key "$keys/cosign.key" \--signing-config offline-signing.json --bundle "release/$whl.sigstore.json" "release/$whl"cp "$keys/cosign.pub" cosign.pubchmod -R u=rwX,go=rX release # readable by the service account; install.sh makes it root's./bin/syft scan -q dir:release/venv -o cyclonedx-json=release/sbom.cdx.jsonprintf 'release: %s, %s components in the SBOM; signing key deleted, public key in cosign.pub\n' \"$whl" "$(jq '.components | length' release/sbom.cdx.json)"
A signature on the wheel says nothing about the installed copies Python will import. verify_core.py compares them with the signed wheel's RECORD and refuses extra files. certrun.sh runs it with the system python3 -I -S, so nothing in the venv (a .pth file in site-packages runs at startup) runs before the check. The gate checks the core's own files only: the interpreter, the dependencies, pyvenv.cfg and site hooks such as a .pth file rely on root ownership alone. The bytecode is root's too: install.sh compiles it next to the source in checked-hash mode (a .pyc is used only while it matches the source beside it, which the gate checked), and the core runs with -B, so it writes none. Under -I, PYTHONDONTWRITEBYTECODE=1 would be ignored; the lab checked both.
"""Is the installed core the signed wheel? Every file the wheel's RECORD lists for the certrunpackage must be installed in VENV with that hash, and nothing else may sit in the packagedirectory. Run by certrun.sh after cosign has verified the wheel, with the system python -I -S,so nothing in the venv (a .pth file in site-packages runs at startup) runs before the check.Standard library only. The dependencies are not checked: root ownership of the release is whatkeeps them as installed. Usage: python3 -I -S verify_core.py WHEEL VENV exit 0 same files, 1 not"""import base64import csvimport hashlibimport ioimport sysimport zipfilefrom pathlib import PathPACKAGE = "certrun/"def main(wheel: str, venv: str) -> int:site = next(Path(venv).glob("lib/python3*/site-packages"), None)if site is None:print(f"verify_core: no site-packages in {venv}", file=sys.stderr)return 1with zipfile.ZipFile(wheel) as zf:record = next(n for n in zf.namelist() if n.endswith(".dist-info/RECORD"))rows = csv.reader(io.StringIO(zf.read(record).decode()))signed = {path: digest for path, digest, _size in rows if path.startswith(PACKAGE)}problems = []for path, digest in sorted(signed.items()):algorithm, _, expected = digest.partition("=")try:data = (site / path).read_bytes()except OSError:problems.append(f"{path}: missing")continueactual = base64.urlsafe_b64encode(hashlib.new(algorithm, data).digest()).rstrip(b"=")if actual.decode() != expected:problems.append(f"{path}: differs from the signed wheel")for file in sorted((site / PACKAGE).rglob("*")):name = file.relative_to(site).as_posix()if file.is_file() and "__pycache__" not in file.parts and name not in signed:problems.append(f"{name}: not in the signed wheel")for problem in problems:print(f"verify_core: {problem}", file=sys.stderr)return 1 if problems else 0if __name__ == "__main__":sys.exit(main(sys.argv[1], sys.argv[2]))
ls finds only cosign.pub: the private key is gone. Syft listed 8 packages (library), certrun and its seven locked dependencies, plus 23 file entries it hashed.
The gate is code, and a gate the service account can edit is no gate: owning certrun.sh, the verifier, the public key or the directory holding the release, it could replace any of them. install.sh puts everything that decides what runs under /opt/scr-capstone, root's all the way up, and lets the account write only to /var/lib/scr-capstone. The CA token and its SSH key are root's, readable by its group only. Metrics live outside any home directory, because Ubuntu 26.04 homes are 0750 and node_exporter could not reach inside.
#!/usr/bin/env bash# Install the certrun runner for the service ACCOUNT, as root, from the project directory:# /opt/scr-capstone root:root, 0755: certrun.sh, the release, the gate and its trust# anchors, the settings. The account can run all of it, change none.# /var/lib/scr-capstone state/ (locks, pending rotations, audit; 0700) and textfile/ (0755,# for node_exporter): the only places the account can write.# /etc/systemd/system the unit and timer, with User=ACCOUNT. The timer is not enabled.# Undo: sudo rm -rf /opt/scr-capstone /var/lib/scr-capstone \# /etc/systemd/system/scr-capstone-certrun@.{service,timer} && sudo systemctl daemon-reloadset -euo pipefail((EUID == 0)) || { echo "install.sh: run as root" >&2; exit 1; }account=${1:?usage: install.sh ACCOUNT}group=$(id -gn -- "$account")opt=/opt/scr-capstone var=/var/lib/scr-capstone sshd=/run/scr-capstone-sshdcd -- "$(dirname -- "${BASH_SOURCE[0]}")"[[ ! -e $opt ]] || { echo "install.sh: $opt exists; remove it first" >&2; exit 1; }install -d -m 0755 "$opt" "$opt/bin" "$opt/etc"install -m 0755 certrun.sh "$opt/"install -m 0644 verify_core.py "$opt/"install -m 0755 bin/cosign "$opt/bin/"cp -R release "$opt/release" && chown -R root:root "$opt/release"# The bytecode, compiled here by root next to the source, so the account cannot write it (and# certrun.sh runs python -B: nothing is written at run time). checked-hash: a .pyc is used only# while it matches the source file beside it, and the gate checks that source."$opt/release/venv/bin/python" -I -m compileall -q -j 0 --invalidation-mode checked-hash "$opt/release/venv/lib"install -m 0644 cosign.pub etc/lab.toml "$opt/etc/" # the pinned signing key, the settingsinstall -m 0644 lab/pki/ca.pem "$opt/etc/lab-ca.pem"install -m 0644 "$sshd/known_hosts" "$opt/etc/known_hosts"# Secrets the account must read and must not change: root's, readable by its group only.install -m 0640 -g "$group" lab/pki/ca-token "$opt/etc/ca-token"install -m 0640 -g "$group" "$sshd/id_certrun" "$opt/etc/id_certrun"install -d -m 0755 "$var"install -d -m 0700 -o "$account" -g "$group" "$var/state"install -d -m 0755 -o "$account" -g "$group" "$var/textfile"sed "s/^User=.*/User=$account/" systemd/scr-capstone-certrun@.service \>/etc/systemd/system/scr-capstone-certrun@.serviceinstall -m 0644 systemd/scr-capstone-certrun@.timer /etc/systemd/system/systemctl daemon-reloadfind "$opt" "$var" -maxdepth 2 -not -path "$opt/release/*" -printf '%M %u:%g %p\n'
All three fail, the mv too: renaming a directory needs write access to its parent. The 242 .pyc files install.sh compiled are root's, 0644; the lab also checked that no later run wrote one. The threat model: this layout stops the service account, or code running as it, from changing what runs; the gate stops a core nobody signed (an edit in place, a corrupted copy, a development build) from running. Neither protects against root or whoever holds the signing key. The undo is in install.sh's header.
The unit is a template: in scr-capstone-certrun@lab.service, the part after @ (%i) is the destination. Type=oneshot makes systemctl start wait for the run. For a oneshot unit, death by SIGTERM counts as a failure (systemd.service(5), SuccessExitStatus=), so every shutdown would look like a broken job; listing SIGTERM makes it a clean stop. KillMode=mixed sends SIGTERM to certrun.sh alone and SIGKILL to whatever is left once it has exited, which is why the script waits for the core. Restart= wires the contract to the scheduler: 75 is retried in 15 minutes, every other failure waits for a person.
# One apply run of certrun for the destination %i, e.g. scr-capstone-certrun@lab.service.# Started by scr-capstone-certrun@%i.timer, or by hand with systemctl start.[Unit]Description=certrun apply %i: renew expiring TLS certificates[Service]Type=oneshot# install.sh puts the service account here (in production, a system account of its own)User=certrunExecStart=/opt/scr-capstone/certrun.sh apply %iEnvironment=CERTRUN_BUDGET=120# A stop (shutdown, deploy, an operator) ends the run by SIGTERM. For Type=oneshot that counts# as a failure unless it is listed here, and a failure is what OnFailure= alerts react to.SuccessExitStatus=SIGTERM# SIGTERM to certrun.sh only, which forwards it to the core; SIGKILL to whatever is left once# certrun.sh has exited, or after TimeoutStopSec. A call running in a thread cannot be cancelled,# so a stop waits for it: 30 s is above the longest, one SSH call of about 25 s (connect, banner# and auth 5 s each, then 5 s per idle read of its output and of its errors).KillMode=mixedTimeoutStopSec=30# 75 means "temporary, run again later": retry it in 15 minutes. The other failures need a# person, not a loop.Restart=on-failureRestartSec=15minRestartPreventExitStatus=1 2 70 77 78NoNewPrivileges=yesPrivateTmp=yesProtectSystem=strictReadWritePaths=/var/lib/scr-capstone
A timer is a second unit that starts the service on a schedule (systemd.timer(5)): every day at 00:00 and 12:00, catching up a run missed while the host was down, spread over half an hour across runners.
# Twice a day for the destination %i. In production:# systemctl enable --now scr-capstone-certrun@lab.timer[Unit]Description=certrun apply %i, twice a day[Timer]OnCalendar=*-*-* 00,12:00:00# A run missed while the host was down happens at the next boot; many runners spread their# requests to the CA over half an hour instead of all arriving at 00:00:00.Persistent=trueRandomizedDelaySec=30min[Install]WantedBy=timers.target
systemd-analyze verify loads the units as systemd would (the grep keeps lines about ours: none), and calendar shows when the timer would fire. It stays disabled and inactive (is-active exits 3): an armed timer could fire mid-lesson and race your manual runs, so the lesson starts the service directly. In production, systemctl enable --now scr-capstone-certrun@lab.timer arms it.
With the release installed, the first failures a person meets:
A mistyped mode is 2. An unknown destination is 78 from the script. A settings file with renew_within_days = "30" passes the gate, and the core's loader names the setting, the type and the value, then exits 78, which Restart= does not retry.
End to end
Plan first. It looks at everything and changes nothing, not even the metrics:
Three to rotate, one fine, two for a person, so the plan exits 1, as the job would. The search line gives both fingerprints, which is what a person needs to decide. On stderr, one JSON object per event, with the run ID in each.
Now an apply stopped in a CA call: the step waits until the CA has issued legacy's certificate and is still holding the answer, then sends SIGTERM. At a prompt, Bash may also print a job notice.
legacy went first, being expired. Its first request got a 500; the retry was in flight when SIGTERM arrived, the core stopped, and certrun.sh re-raised: 143. No run directory, the lock is free, no core process, the run metric says 143, and one file remains on purpose: pending/legacy.json, holding the 473 request bytes the CA has already answered.
Next, the unit, started without waiting; the step waits for the web server's second reload request (billing's deploy) and runs systemctl stop there, as a shutdown would.
inactive and success, "Deactivated successfully": that is SuccessExitStatus=SIGTERM; without it the same stop leaves the unit failed with result signal (the lab checked both). legacy resumed: the recorded bytes went out again, the CA replayed its answer, and deploy, reload, check and audit followed. billing got a 429, a slow answer (TimeoutError) and a replay, and was stopped after its deploy. Once that reload landed, billing serves a certificate with the key hash in pending/billing.json, and has no audit line. Now let the unit run to the end:
The run exited 1, so the unit entered the failed state (status=1/FAILURE), which is what OnFailure= reacts to, and Restart= did not retry. billing needed no CA request: its live certificate carries the record's key, so the run only wrote the audit line. shop shows the rollback: a certificate from an untrusted issuer failed the post-deploy check, and shop serves its old one again. The CA's log: seven requests, three certificates, two replays; each (4 s late) line is the CA answering a client that had stopped waiting.
Nothing to do, and the CA's counters did not move: shop's rolled-back record blocks new requests until a person looks. One audit line per change, billing's marked recovered, each new key different from the old. certrun_last_run_seconds and certrun_last_run_status, written by certrun.sh on every apply that got the lock, say whether the job runs and how it last ended. certrun_last_success_seconds is when a run last handled every endpoint (exit 0 or 1); a 75 run leaves it alone, so a CA or host that stays down shows as its age. In "Observable jobs" success was exit 0, but here 1 is normal while any endpoint waits for a person, and a success stuck at 0 could never show a stopped timer; its age is that alert. certrun_needs_attention (3) pages a person, and the per-endpoint expiry catches a certificate near its end whatever the job does.
Last, a core that is not the release: a wheel with one byte appended, then a "hotfix" in an installed file (each needs sudo, which is the point):
Both 77, recorded in the run metric, and the core never started. cosign refused the changed wheel; with the wheel intact (Verified OK), verify_core.py still refused, naming the edited file. The WARNING is cosign's reminder that the transparency log was skipped. After the repair, the run is back to its honest 1.
Try this
Fix search and shop the right way, before you clean up. Someone with console access to web2 reads its host key fingerprint there (here: ssh-keygen -lf /run/scr-capstone-sshd/host-web2.pub) and finds what web2 offered in the plan. Pinning is a root-owned change: in /opt/scr-capstone/etc/known_hosts, replace the [127.0.0.1]:18952 line with [127.0.0.1]:18952 and the first two fields of host-web2.pub (ssh-ed25519 AAAA...), then run /opt/scr-capstone/certrun.sh apply lab. Expected: search rotated, exit 1, 8 CA requests with "search": 1. The CA's fault for shop was a one-off, so delete /var/lib/scr-capstone/state/lab/pending/shop.json and apply again. Expected: shop rotated, "shop": 2, and certrun_needs_attention 1. Then explain what must change about admin (its certificate names www) before certrun may touch it, and why it does not decide that itself.
Clean up
Stop the lab servers and the hosts, reset the failed unit, remove the installation, and delete the project (everything left in it is yours):
0 units left. To start over, run this clean-up without its last line and begin at the first terminal: a restarted CA fails the same way again.
Takeaway
Give an unattended job an entry point that owns the process, install whatever decides what runs where the job's account cannot change it, and write each change down before making it, so a retry, crash or stop at any step is finished by the next run instead of repeated or lost.
This is the end of the Scripting track. The CI/CD courses take the release half further. Start with "Software supply chain security" (/courses/ci-supply/) for signing, attestations and verification across a pipeline, then go deeper into SLSA, in-toto and Sigstore in "Software supply chain in depth" (/courses/supplychain-deep/). "Secure CI/CD with GitLab" (/courses/ci-gitlab/) puts gates like these into a GitLab pipeline.
legacy's certificate but before certrun heard the answer. The next run starts. What keeps the CA from issuing a second certificate?certrun does not rely on one; a CA that had it would also block real renewals.Type=oneshot with KillMode=mixed, and systemctl stop sends SIGTERM to certrun.sh while the core is running. Why does the script forward the signal, wait for the core and re-raise, instead of calling exit 143 in its TERM trap?exit works in a trap; the problem is what happens to the core once the script has gone.KillMode=mixed, whatever is left after the main process exits gets SIGKILL. And status 143 would be a plain failure, while SuccessExitStatus=SIGTERM makes a death by the signal a clean stop.systemctl stop never triggers Restart=.exit too; the bats test showed cleanup and the 143 metric on the re-raise path as well.api's certificate still has 200 days left. The TLS check now fails with OpenSSL verify code 10, "certificate has expired". What does certrun do with api?certrun keeps temporary for hosts and CAs that do not answer.not_after; with 200 days left, code 10 must come from the chain, which only a person can fix.