Capstone: a certificate-expiry and rotation runner

A Bash entry point, a packaged Python core, a fleet, a flaky CA API and release gates.

Expert90 min · lesson 15 of 15
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (42 files, 53 KB): scr-capstone.tar.gz. Unpack it with tar -xzf scr-capstone.tar.gz, which creates scr-capstone/. SHA-256: d82da6e9216f939f5c8e5f6b05bcccaa7694182d418a0ec532cb3cbb994b65a4

This capstone builds certrun, a job that finds TLS certificates about to expire and replaces them, and runs it the way production would: installed root-owned, started by systemd as an unprivileged service account, with nobody watching. You will read the parts that are new here, test them, build, sign and install a release, and run it end to end: a plan, a run stopped by SIGTERM in a CA call, a run stopped by systemctl stop in a deploy, the run that finishes both and rolls back a bad certificate, a run with nothing left to do, and a tampered release that is refused.

Refresher: the entry point is from "When a script has outgrown Bash", deadlines and signals from "Signals, timeouts and shutdown", typed settings from "Designing a CLI tool", retries and idempotency keys from "Failure design", pinned SSH from "Fleet automation", logs, metrics and audit lines from "Observable jobs", and the lock file, SBOM and signature from the last two lessons.

What one run does, and what it promises

certrun.sh plan|apply DESTINATION is the only command systemd knows; a destination is a set of hosts and endpoints with its own settings file. The service account is the unprivileged account the job runs as. In this lab your own account plays it; in production it is a system account with no login shell.

One run of certrun.sh apply lab
1certrun.sh: arguments, settings file, lock, traps
bad arguments 2; no settings file 78; another run holds the lock 75
2certrun.sh: the verify gate
wheel signature, then the installed files; any doubt 77, the core never starts
3certrun.sh: the core under a deadline
a JSON job file, an environment of three variables; deadline reached 75
4core: load the settings, then look at everything at once
a bad setting 78; TLS handshakes and SSH logins, bounded
5core: decide per endpoint, then act, most urgent first
record, CA request, deploy, reload, check; audit or roll back
6core: report
table on stdout, JSON logs on stderr, metrics; 0 ok, 1 a person must look, 75 retry later
SIGTERM at any point: the core cancels and dies of SIGTERM, certrun.sh does the same (a shell sees 143), and SuccessExitStatus=SIGTERM in the unit makes that a clean stop for systemd. Every apply that gets the lock, however it ends, records when and how in a metric.

The exit contract is the interface to everything that watches the job. 0: every endpoint is fine or was rotated. 1: some endpoints need a person (or the endpoint metrics could not be written), everything else was done. 75 (EX_TEMPFAIL): only temporary failures, such as a host that did not answer; run again later. 78 (EX_CONFIG): fix the settings, retrying is pointless. 70 (EX_SOFTWARE): a bug outside any one endpoint's work. The entry point adds 2 for usage, 75 for a held lock or a spent deadline, and 77 (EX_NOPERM) when it refuses the core. 70, 75, 77 and 78 come from sysexits.h; 0, 1 and 2 are the shell's conventions.

The lab, and the entry point

What you need: Ubuntu 26.04 with sudo; openssh-server, jq, curl, and bats with bats-support and bats-assert (as in "Testing shell automation"); network access to GitHub releases and PyPI; free ports 18941-18946, 18951-18952 and 18961; membership in adm or systemd-journal to read the journal. Unpack the lesson files in your home directory (cd ~ && tar -xzf scr-capstone.tar.gz && cd scr-capstone); every terminal runs in ~/scr-capstone. No file names an account: the scripts that need one take $USER.

The lab stands in for the outside world; its scripts are in the lesson files and are not walked through. lab/get-tools.sh downloads uv 0.12.19, cosign v3.1.3 and Syft 1.52.0 into bin/, each checked against a pinned SHA-256. lab/hosts.sh (as root) plays the hosts web1 and web2: two throwaway sshd units on 127.0.0.1 (never port 22) with their keys and pinned known_hosts in root's /run/scr-capstone-sshd, and /srv/scr-capstone (yours) for the hosts' certificate files. lab/make-pki.sh makes a throwaway CA, a certificate per endpoint with chosen dates, the CA API's certificate, a test token, two CAs for failure cases and, for the tests, a valid certificate under the expired CA: 8 certificates in all. lab/ca_api.py is the CA's API, which fails on purpose. lab/tlsfarm.py is the hosts' web server: like nginx, it loads certificate files at start and on SIGHUP, never per connection, and a reload takes effect 2 seconds after the signal.

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ lab/get-tools.sh
uv 0.12.19 (aarch64-unknown-linux-gnu) GitVersion: v3.1.3 syft 1.52.0
$ sudo lab/hosts.sh start "$USER"
hosts: web1 on 127.0.0.1:18951, web2 on 127.0.0.1:18952, files in /srv/scr-capstone
$ lab/make-pki.sh for f in /srv/scr-capstone/*/*/tls.pem; do printf "%-39s %-23s until %s\n" "$f" "$(openssl x509 -in "$f" -noout -ext subjectAltName | tail -n 1 | tr -d " ")" \ "$(openssl x509 -in "$f" -noout -enddate | cut -d= -f2)" done
make-pki: 3 CAs (lab, untrusted, expired) and 8 certificates /srv/scr-capstone/web1/admin/tls.pem DNS:www.lab.internal until Feb 26 03:13:56 2027 GMT /srv/scr-capstone/web1/api/tls.pem DNS:api.lab.internal until Apr 17 03:13:55 2027 GMT /srv/scr-capstone/web1/billing/tls.pem DNS:billing.lab.internal until Oct 11 03:13:55 2026 GMT /srv/scr-capstone/web1/legacy/tls.pem DNS:legacy.lab.internal until Sep 26 03:13:55 2026 GMT /srv/scr-capstone/web1/shop/tls.pem DNS:shop.lab.internal until Oct 19 03:13:56 2026 GMT /srv/scr-capstone/web2/search/tls.pem DNS:search.lab.internal until Oct 8 03:13:56 2026 GMT
$ python3 lab/ca_api.py lab/pki lab/ca-faults.json 18961 >lab/ca.log 2>&1 & echo $! >lab/ca.pid python3 lab/tlsfarm.py etc/lab.toml /srv/scr-capstone/tlsfarm.pid >lab/tlsfarm.log 2>&1 & timeout 10 bash -c "until curl -fs --cacert lab/pki/ca.pem https://127.0.0.1:18961/v1/stats; do sleep 0.2; done"; echo timeout 10 bash -c "until [ -s /srv/scr-capstone/tlsfarm.pid ]; do sleep 0.2; done" cat lab/ca-faults.json
{"requests": 0, "issued": {}, "replayed": 0} {"legacy": ["500", "slow"], "billing": ["429", "slow"], "shop": ["untrusted"]}
$ ssh-keygen -lf /run/scr-capstone-sshd/known_hosts ssh-keyscan -q -t ed25519 -p 18951 127.0.0.1 | ssh-keygen -lf - ssh-keyscan -q -t ed25519 -p 18952 127.0.0.1 | ssh-keygen -lf -
256 SHA256:ekuaplecONFYqRig9asmJVGA/9V7zKp+tvawFWdAT74 [127.0.0.1]:18951 (ED25519) 256 SHA256:+E80JU0VjFDBurfzC9sw6N4L2bIMeaIjk3//CSd7iDQ [127.0.0.1]:18952 (ED25519) 256 SHA256:ekuaplecONFYqRig9asmJVGA/9V7zKp+tvawFWdAT74 [127.0.0.1]:18951 (ED25519) 256 SHA256:gCq4H8S0hfCbGZn7XZPzjPIqN/N/mra0Cl7Hd5FBOmM [127.0.0.1]:18952 (ED25519)

Six endpoints, one path each: api has 200 days left, billing 12, legacy expired three days ago, shop has 20, admin serves a certificate for the wrong name (www), and search lives on web2. The CA's faults give legacy a 500 then a slow answer, billing a 429 then a slow answer, and shop a certificate from an issuer nobody trusts. In the last step, ssh-keygen -lf prints key fingerprints and ssh-keyscan fetches a server's host key without logging in: lines 1-2 are the pins, lines 3-4 what the hosts offer now. web2 offers a key other than its pin, the key it had before it was rebuilt; certrun cannot tell a rebuild from an attacker in the middle, so it must refuse. sudo lab/hosts.sh stop undoes the hosts (see Clean up).

certrun.sh owns what Bash does well. It checks its arguments (the destination is a name, never a path), takes an flock on /var/lib/scr-capstone/state/DESTINATION.lock so two runs for one destination never overlap, and makes a private run directory. The EXIT trap removes that directory on every path and, for an apply that got the lock, records when the run ended and how. A run refused at the lock did nothing, so it records nothing: the metric's age shows that no run gets through. The TERM and INT traps record the signal and pass a stop request to the core; once the core has finished, reraise sends the same signal to the script itself, so a supervisor sees the death it asked for.

Then the gate (explained with the release), and the job file, written with jq --arg, which escapes each value. The core runs under timeout with what is left of the budget, with only HOME, LANG and PATH in its environment and python -I -B (the -B is explained with the release). The last case maps its status: the core's contract passes through, 124 is a spent deadline (75), and 137 is a deadline only if the deadline has passed; a core SIGKILLed earlier, say by the out-of-memory killer, is 70.

certrun.sh
#!/usr/bin/env bash
# certrun.sh plan|apply DESTINATION: the entry point the systemd unit runs as the service account.
# It owns the process: one run per destination, a deadline for the whole run, signals, cleanup,
# the run-level metrics, and a gate that refuses a core that is not the signed release.
# Exit status: the core's contract passed through (0 ok, 1 a person must look at an endpoint or
# at the metrics, 70 internal, 75 temporary, 78 config), plus 2 usage, 75 lock held or deadline
# reached, 77 core refused, 78 no settings or a bad CERTRUN_BUDGET, 70 core killed. Stopped by
# SIGTERM, it cleans up and dies of SIGTERM (143 in a shell).
# CERTRUN_BUDGET: seconds for the whole run (default 60). CERTRUN_VAR: state root (for tests).
set -Eeuo pipefail
shopt -s inherit_errexit
here=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd -P) # /opt/scr-capstone: root's
release=$here/release
wheel=$release/certrun-0.1.0-py3-none-any.whl
var=${CERTRUN_VAR:-/var/lib/scr-capstone} # the service account's: state/ and textfile/
budget=${CERTRUN_BUDGET:-60}
mode='' dest='' run='' child='' stop='' locked=''
log() { # log LEVEL MESSAGE: one JSON line on stderr, like the core's
jq -ncM --arg level "$1" --arg msg "$2" --arg dest "$dest" \
'{ts: (now | todate), level: $level, logger: "certrun.sh", msg: $msg, destination: $dest}' >&2
}
# shellcheck disable=SC2329 # called from on_exit, the EXIT trap
run_metrics() { # run_metrics STATUS: when the last apply ended and how, rewritten in one rename
local file=$var/textfile/certrun-$dest-run.prom label="{destination=\"$dest\"}" tmp
tmp=$(mktemp "$file.XXXXXX") || return
printf '%s\n' "# TYPE certrun_last_run_seconds gauge" "certrun_last_run_seconds$label $(date +%s)" \
"# TYPE certrun_last_run_status gauge" "certrun_last_run_status$label $1" >"$tmp"
chmod 0644 "$tmp" && mv -f -- "$tmp" "$file"
}
# shellcheck disable=SC2329 # called from the EXIT trap
on_exit() { # every way out: the run metrics (an apply that got the lock), then the run directory
local status=$1
trap '' TERM INT # a second signal must not cut this short
[[ -z $stop ]] || status=$((128 + $(kill -l "$stop")))
[[ $mode != apply || -z $locked ]] || run_metrics "$status" || log warning "run metrics not written"
[[ -z $run ]] || rm -rf -- "$run"
}
# shellcheck disable=SC2329 # called from the TERM and INT traps
on_signal() { # remember the signal and pass a stop request on to the core, if it runs
stop=$1
[[ -z $child ]] || kill -TERM "$child" 2>/dev/null || true
}
reraise() { # die of the signal we were sent; the EXIT trap still runs first
log warning "stopped by SIG$stop"
trap - "$stop"
kill -s "$stop" "$$"
}
usage() {
echo "usage: certrun.sh plan|apply DESTINATION" >&2
exit 2
}
trap 'on_exit $?' EXIT
trap 'on_signal TERM' TERM
trap 'on_signal INT' INT
(($# == 2)) || usage
[[ $1 == plan || $1 == apply ]] || usage
[[ $2 =~ ^[a-z0-9][a-z0-9-]{0,62}$ ]] || usage # a name, never a path
mode=$1 dest=$2
config=$here/etc/$dest.toml
[[ -f $config ]] || { log error "no settings file $config"; exit 78; }
[[ $budget =~ ^[1-9][0-9]*$ ]] || { log error "CERTRUN_BUDGET must be whole seconds"; exit 78; }
deadline=$((SECONDS + budget))
[[ -w $var/state ]] || { log error "$var/state is missing or not writable: see install.sh"; exit 78; }
# One run per destination. The core inherits this descriptor on purpose: if this script is
# killed, the lock stays held until the core has gone too. A run refused here did nothing, so it
# leaves the run metrics alone (their age shows that no run gets through).
exec {lock}>>"$var/state/$dest.lock"
flock -n "$lock" || { log warning "another run holds the lock"; exit 75; }
locked=yes
run=$(mktemp -d -t "certrun-$dest.XXXXXX")
# The gate: verify the signature on a private copy of the wheel with the pinned key, then check
# the installed core against that copy, with the system python (-I -S: no site, so nothing in the
# venv runs before the check). set -e is off inside a function called in a condition: check each.
verify_core() {
cp -- "$wheel" "$run/core.whl" || return 1
cp -- "$wheel.sigstore.json" "$run/core.bundle" || return 1
"$here/bin/cosign" verify-blob --key "$here/etc/cosign.pub" --insecure-ignore-tlog=true \
--bundle "$run/core.bundle" "$run/core.whl" || return 1
/usr/bin/python3 -I -S "$here/verify_core.py" "$run/core.whl" "$release/venv"
}
if ! verify_core >"$run/gate.out" 2>&1; then
log error "refusing to run: the core is not the signed release"
cat -- "$run/gate.out" >&2
exit 77
fi
[[ -z $stop ]] || reraise
jq -n --arg mode "$mode" --arg destination "$dest" --arg config "$config" --arg var "$var" \
--arg run_id "$(date -u +%Y%m%dT%H%M%SZ)-$$" \
'{mode: $mode, destination: $destination, config: $config, run_id: $run_id, var: $var}' >"$run/job.json"
left=$((deadline - SECONDS))
((left > 0)) || { log error "deadline reached before the core started"; exit 75; }
# Only HOME, LANG and PATH reach the core; -I: no PYTHONPATH, no user site, no current
# directory on sys.path. The bytecode is root's, compiled by install.sh next to the source; -B
# writes none at run time (PYTHONDONTWRITEBYTECODE would not do: -I ignores PYTHON* variables).
timeout -k 10 "$left" env -i HOME="$HOME" LANG=C.UTF-8 PATH=/usr/bin:/bin \
"$release/venv/bin/python" -I -B -m certrun "$run/job.json" &
child=$!
[[ -z $stop ]] || kill -TERM "$child" # a signal that came before $! was known
status=0
wait "$child" || status=$?
while kill -0 "$child" 2>/dev/null; do # a trapped signal ends wait early: wait for the core
status=0
wait "$child" || status=$?
done
child=
[[ -z $stop ]] || reraise # the core has finished its own cleanup by now
case $status in
0 | 1 | 70 | 75 | 78) ;; # the core's own contract
124)
log error "deadline of ${budget}s reached: the core was stopped"
status=75
;;
137)
if ((SECONDS >= deadline)); then
log error "deadline of ${budget}s reached: the core ignored SIGTERM and was killed"
status=75
else
log error "the core was killed with SIGKILL before the deadline (out of memory?)"
status=70
fi
;;
*)
log error "the core ended with status $status, which is not in its contract"
status=70
;;
esac
exit "$status"

Two details matter under stress. The trap does not exit: the while kill -0 loop waits for the core's cleanup, because a trapped signal ends wait early. And the script records the signal itself, because timeout's status cannot say who stopped the core: Ubuntu 26.04's uutils timeout -k exits 124 when it is itself sent SIGTERM (143 without -k), while GNU timeout passes on the child's status; the lab checked each case. tests/certrun.bats (in the lesson files) tests the entry point with a fake release: tests/fake/cosign passes unless FAKE_COSIGN_FAIL is set, and tests/fake/python plays the core as each test tells it.

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ bats --tap tests/certrun.bats
1..8 ok 1 usage errors exit 2; an unknown destination or a bad budget exits 78 ok 2 the core's exit contract passes through; anything else becomes 70 ok 3 a core that fails verification never starts: 77 ok 4 a second run while the lock is held exits 75, starts nothing, records nothing ok 5 the whole-run deadline stops a hung core: 75, nothing left running ok 6 a core killed by SIGKILL before the deadline is not called a deadline: 70 ok 7 SIGTERM mid-run: core stopped, run directory gone, lock free, status 143 ok 8 every apply that gets the lock records its status, a plan records nothing

Captured without a terminal (--tap). Test 6: a core SIGKILLed right away is 70, not a deadline. Test 7: SIGTERM mid-run gives 143 (128 + 15), the core ran its handler, the run directory is gone, the lock is free, and the metric says 143. Tests 1, 4 and 8: every apply that gets the lock records its status, 77 included; an unknown destination, a run refused at the lock and a plan record nothing.

The core: what it sees, and how it deploys

The core is a package, certrun, with two dependencies: paramiko for SSH and cryptography for X.509. pyproject.toml and the lab's uv.lock are in the lesson files, so your dependency tree matches the one below. The settings for the destination lab (44 of 79 lines; the other five endpoints look like api):

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ cat etc/lab.toml
# certrun settings for the destination "lab": what to check, and how. Reviewed like code and # installed root-owned with it (/opt/scr-capstone/etc). Relative paths are relative to this file. # No secrets here, only the files that hold them. [ca] url = "https://127.0.0.1:18961/v1" ca_file = "lab-ca.pem" # trust anchor for the CA API's own certificate token_file = "ca-token" # others must have no access, or the run refuses timeout = 2.0 # seconds for the connect, and for each read attempts = 4 # per certificate, the first one included max_retry_after = 10.0 # a longer Retry-After ends the attempts until the next run [policy] trust_anchor = "lab-ca.pem" # what the endpoints' certificates must chain to renew_within_days = 30 check_concurrency = 3 # TLS handshakes in flight at once check_timeout = 5.0 # seconds per handshake reload_wait = 5.0 # seconds a reloaded server gets to serve the new certificate [ssh] user = "" # empty: the account certrun runs as identity = "id_certrun" known_hosts = "known_hosts" # the pinned host keys timeout = 5.0 # In the lab, one TLS server process plays the web servers of both hosts; SIGHUP reloads it. [[hosts]] name = "web1" address = "127.0.0.1" port = 18951 reload = 'kill -HUP "$(cat /srv/scr-capstone/tlsfarm.pid)"' [[hosts]] name = "web2" address = "127.0.0.1" port = 18952 reload = 'kill -HUP "$(cat /srv/scr-capstone/tlsfarm.pid)"' [[endpoints]] name = "api" host = "web1" port = 18941 server_name = "api.lab.internal" pem_path = "/srv/scr-capstone/web1/api/tls.pem" # certificate and key in one file …

Relative paths are relative to this file, which is installed root-owned in /opt/scr-capstone/etc next to the keys it names. Each host's reload command runs over SSH after a deploy. Where state and metrics live belongs to the installation, so certrun.sh passes it in the job.

These modules are in the lesson files. config.py and loader.py: the dataclasses and loader of "Designing a CLI tool", now resolving relative paths and requiring an absolute pem_path. retry_after.py: the parser of "Failure design", unchanged (isascii() before isdigit(), because '²'.isdigit() is true). ssh.py: the pinned-key connect of "Fleet automation", with a host that is down or slow as TransientError and a key mismatch, refused login or failing command as TargetError. obs.py: the JSON logs, textfile writer and fsynced audit append of "Observable jobs". errors.py: those errors, ConfigError, Stopped and the exit codes. tlscheck.py: one verified handshake per endpoint, bounded by a semaphore and asyncio.timeout, and on failure a non-verifying one that only reads the certificate. cli.py: the job file and the exit statuses. It prints the results table before it writes the endpoint metrics, so a failed write (a full disk, a missing textfile/) cannot hide what the run did: it logs "endpoint metrics not written" and the run exits 1.

remote.py works on a host, each step over its own SSH connection. The install is a short script in the remote shell with the new PEM on standard input: it keeps a backup of the old file (once per rotation, so a resumed rotation never backs up its own new file), starts the new file as a copy of the old one so owner and mode stay what the server expects, and replaces the old one in one rename. The server finds the old pair or the new one, never half of each, which is why key and certificate share a file. A symlink is refused: the rename would replace the link, not the file the server reads.

src/certrun/remote.py
"""What certrun does on a host, each over its own SSH connection: read the certificate in every
PEM file, replace one file with a backup, put the backup back, and have the server reload.
Blocking (paramiko), so the async code calls these through asyncio.to_thread."""
import shlex
from .config import Config, Endpoint, Host
from .errors import TargetError, TransientError
from .ssh import connect, run
# Runs in the remote login shell with the new PEM on stdin. The old file is kept as BACKUP (once
# per rotation: a resumed rotation must not back up its own new file). The new file starts as a
# copy of the old one, so it keeps the owner and mode the server expects, and replaces it in one
# rename(2): the server finds the old pair or the new one, never half of each. A symlink is
# refused, because the rename would replace the link instead of the file the server reads.
INSTALL = """set -e
f={path} b={backup}
if [ -L "$f" ]; then echo "$f is a symlink: not replacing it" >&2; exit 1; fi
[ -e "$b" ] || cp -p -- "$f" "$b"
cp -p -- "$f" "$f.certrun-new"
cat >"$f.certrun-new"
mv -f -- "$f.certrun-new" "$f"
"""
def deployed(host: Host, cfg: Config) -> dict[str, str | Exception]:
"""The SHA-256 of the certificate in each of this host's PEM files, or why it is unknown."""
found: dict[str, str | Exception] = {}
with connect(host, cfg) as client:
for ep in (e for e in cfg.endpoints if e.host == host.name):
try:
out = run(client, f"openssl x509 -noout -fingerprint -sha256 -in {shlex.quote(ep.pem_path)}", cfg)
found[ep.name] = out.strip().split("=", 1)[1].replace(":", "").lower()
except (TargetError, TransientError) as err:
found[ep.name] = err
except (IndexError, ValueError):
found[ep.name] = TargetError(f"{ep.pem_path}: unexpected openssl output")
return found
def install(host: Host, ep: Endpoint, pem: bytes, backup: str, cfg: Config) -> None:
command = INSTALL.format(path=shlex.quote(ep.pem_path), backup=shlex.quote(backup))
with connect(host, cfg) as client:
run(client, command, cfg, data=pem)
def rollback(host: Host, ep: Endpoint, backup: str, cfg: Config) -> None:
"""Put back the file install() replaced, again in one rename."""
with connect(host, cfg) as client:
run(client, f"mv -f -- {shlex.quote(backup)} {shlex.quote(ep.pem_path)}", cfg)
def forget(host: Host, backup: str, cfg: Config) -> None:
"""The backup holds the old private key: remove it once the new certificate is live."""
with connect(host, cfg) as client:
run(client, f"rm -f -- {shlex.quote(backup)}", cfg)
def reload(host: Host, cfg: Config) -> None:
with connect(host, cfg) as client:
run(client, host.reload, cfg)

decide.py turns what the survey saw into an action, with no network and no side effects. The host comes first: never issue a certificate you cannot deliver. Then any record a stopped run left (next section). Then the file on disk must be what the endpoint serves; if not, a reload is pending or failed, and another certificate would not help. Expiry comes from the leaf's own not_after, not from the verify code: code 10, "certificate has expired", is also what an expired CA above a valid leaf gives.

src/certrun/decide.py
"""decide(): what to do with one endpoint, from what it serves, what its host has on disk and the
record of any rotation a stopped run left behind. No network and no side effects."""
from dataclasses import dataclass
from datetime import datetime
from .config import Config, Endpoint
from .errors import TransientError
from .pending import Pending
from .tlscheck import EXPIRED, Seen
@dataclass
class Result:
endpoint: str
action: str # ok, rotate, finish, rollback (plan); rotated, failed (apply)
detail: str
temporary: bool = False # failed, and the next run may succeed
seen: Seen | None = None
record: Pending | None = None
not_after: datetime | None = None
def failed(ep: Endpoint, err: Exception, seen: Seen | Exception) -> Result:
not_after = seen.not_after if isinstance(seen, Seen) else None # still worth a metric
return Result(ep.name, "failed", str(err), temporary=isinstance(err, TransientError), not_after=not_after)
def decide(ep: Endpoint, seen: Seen | Exception, disk: str | Exception, rec: Pending | None,
cfg: Config) -> Result:
for outcome in (disk, seen): # the host first: no certificate for a host we cannot trust
if isinstance(outcome, Exception):
return failed(ep, outcome, seen)
assert isinstance(seen, Seen) and isinstance(disk, str)
def act(action: str, detail: str) -> Result:
return Result(ep.name, action, detail, seen=seen, record=rec, not_after=seen.not_after)
days = seen.days_left()
if rec and rec.rolled_back:
return act("failed", f"a deploy was rolled back ({rec.rolled_back}); fix the cause,"
f" then delete pending/{ep.name}.json")
if rec and rec.deployed and seen.key == rec.key: # deployed, never checked: check it now
if not seen.problem:
return act("finish", "a stopped run deployed the new certificate; the audit line is missing")
return act("rollback", f"a stopped run deployed a certificate that fails its check: {seen.problem}")
if rec and rec.replacing in (seen.fingerprint, disk):
return act("rotate", "resume the rotation a stopped run began")
if rec:
return act("failed", f"pending/{ep.name}.json replaces {rec.replacing[:12]}, which is gone")
if disk != seen.fingerprint:
return act("failed", f"{ep.host} has {disk[:12]} on disk but serves {seen.fingerprint[:12]}:"
" a reload is pending or failed")
if seen.problem and seen.code != EXPIRED: # a wrong name, an unknown issuer: not ours to fix
return act("failed", f"serves a certificate that does not verify: {seen.problem}")
if seen.problem and not seen.expired: # code 10, but the leaf itself is valid
return act("failed", f"the chain does not verify ({seen.problem}), but the certificate itself"
f" has {days} days left: an issuer's certificate expired")
if seen.expired:
return act("rotate", f"EXPIRED {-days} days ago")
return act("rotate" if days <= cfg.policy.renew_within_days else "ok", f"{days} days left")
deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ ./bin/uv lock --check ./bin/uv tree --no-dev
Using CPython 3.14.4 interpreter at: /usr/bin/python3 Resolved 14 packages in 5ms Using CPython 3.14.4 interpreter at: /usr/bin/python3 Resolved 14 packages in 0.64ms certrun v0.1.0 ├── cryptography v50.0.1 │ └── cffi v2.1.1 │ └── pycparser v3.0 └── paramiko v5.0.0 ├── bcrypt v5.0.0 ├── cryptography v50.0.1 (*) ├── invoke v3.0.3 └── pynacl v1.6.2 └── cffi v2.1.1 (*) (*) Package tree already displayed
$ ./bin/uv sync -q --locked .venv/bin/pytest -v -p no:cacheprovider
… tests/test_ca.py::test_429_waits_as_asked_then_succeeds[ca_server0] PASSED [ 4%] tests/test_ca.py::test_a_lost_answer_is_replayed_not_issued_twice[ca_server0] PASSED [ 8%] tests/test_ca.py::test_the_same_key_with_a_new_csr_is_refused PASSED [ 12%] tests/test_ca.py::test_a_bad_token_is_not_retried PASSED [ 16%] tests/test_ca.py::test_a_redirect_is_refused_and_the_token_stays[ca_server0] PASSED [ 20%] tests/test_ca.py::test_retry_after_parsing PASSED [ 25%] tests/test_ca.py::test_a_retry_after_above_the_limit_gives_up_until_the_next_run PASSED [ 29%] tests/test_cli.py::test_only_a_run_that_handled_every_endpoint_moves_last_success PASSED [ 33%] tests/test_decide.py::test_a_record_is_saved_privately_and_read_back PASSED [ 37%] tests/test_decide.py::test_a_stopped_run_is_resumed_finished_or_rolled_back PASSED [ 41%] tests/test_decide.py::test_disk_and_served_differ_so_nothing_is_issued PASSED [ 45%] tests/test_decide.py::test_an_unreachable_host_is_temporary_a_wrong_key_is_not PASSED [ 50%] tests/test_loader.py::test_the_lab_settings_load_with_paths_relative_to_the_file PASSED [ 54%] tests/test_loader.py::test_one_bad_setting_is_named[string-for-int] PASSED [ 58%] tests/test_loader.py::test_one_bad_setting_is_named[bool-for-int] PASSED [ 62%] tests/test_loader.py::test_one_bad_setting_is_named[typo] PASSED [ 66%] tests/test_loader.py::test_one_bad_setting_is_named[no-such-host] PASSED [ 70%] tests/test_loader.py::test_one_bad_setting_is_named[relative-pem-path] PASSED [ 75%] tests/test_loader.py::test_one_bad_setting_is_named[plain-http] PASSED [ 79%] tests/test_tlscheck.py::test_a_good_certificate_verifies PASSED [ 83%] tests/test_tlscheck.py::test_an_expired_leaf_is_rotated PASSED [ 87%] tests/test_tlscheck.py::test_an_expired_ca_does_not_make_a_valid_leaf_expired PASSED [ 91%] tests/test_tlscheck.py::test_a_certificate_for_another_name_needs_a_person PASSED [ 95%] tests/test_tlscheck.py::test_no_answer_is_temporary PASSED [100%] … ============================== 24 passed in 6.21s ==============================

14 packages: certrun, its seven runtime dependencies, and pytest with its own (the development group). The 24 tests run the CA API and TLS endpoints in-process on free ports (tests/conftest.py); test_an_expired_ca_does_not_make_a_valid_leaf_expired is the code-10 rule, and test_cli.py holds the metrics rules of the end-to-end section. SSH is not faked: the end-to-end runs use real sshd.

Rotation that survives a stop

A rotation spans several systems, and a stop can come between any two steps. The rule that keeps it safe: write down what you are about to do before you do it. pending.py saves a record before the CA is asked: the new private key, the exact request bytes and a random idempotency key, 0600 in the 0700 state directory, fsynced with its directory. Just before the deploy, the record is marked deployed: the new file may be live, but nobody has checked it. The next run takes one of four paths:

The next run finds pending/NAME.json
A record a stopped run left
the key, the request bytes and the idempotency key, saved before the CA was asked
deployed, serves the record key, verifies
finish
remove the backup, write the missing audit line, delete the record
deployed, serves the record key, fails
roll back
mark the record rolled back, put the backup back, reload, check
still the old certificate
resume
resend the same bytes with the same key: the CA answers from its record
rolled back
failed
a person fixes the cause, then deletes the record; nothing is requested again
The record is deleted only after the audit line is on disk, and every rotation starts with a new key: a key is never reused across rotations.
src/certrun/pending.py
"""A rotation in progress, written to disk before the CA is asked: the new private key, the exact
request and its idempotency key, and, from just before the deploy, that the new file may be live
but is not yet checked. A run stopped anywhere after that leaves the record, and the next run
resends the same bytes (the CA answers from its own record) or checks the deployed certificate:
the missing audit line if it verifies, a rollback if it does not.
One file per endpoint, 0600 in a 0700 directory."""
import dataclasses
import hashlib
import json
import os
import secrets
from datetime import UTC, datetime
from pathlib import Path
from cryptography import x509
from cryptography.hazmat.primitives import hashes, serialization
from cryptography.hazmat.primitives.asymmetric import ec
from cryptography.x509.oid import NameOID
from .config import Endpoint
from .tlscheck import DER, SPKI, Seen
PEM = serialization.Encoding.PEM
@dataclasses.dataclass(frozen=True)
class Pending:
endpoint: str
replacing: str # fingerprint of the certificate being replaced
replacing_key: str # and of its public key: the new key must differ
key: str # SHA-256 of the new public key
key_pem: str # the new private key (unencrypted, like the PEM the server loads)
request: str # the request body, resent byte for byte
idempotency_key: str
created: str
deployed: bool = False # the new file may be on the host: deployed, not yet checked
rolled_back: str | None = None # why a deploy of this certificate was undone
def start(state: Path, ep: Endpoint, seen: Seen) -> Pending:
"""A new key and request for ep, saved before anything is sent."""
key = ec.generate_private_key(ec.SECP256R1()) # a new key for every rotation, never reused
subject = x509.Name([x509.NameAttribute(NameOID.COMMON_NAME, ep.server_name)])
csr = (x509.CertificateSigningRequestBuilder().subject_name(subject)
.add_extension(x509.SubjectAlternativeName([x509.DNSName(ep.server_name)]), critical=False)
.sign(key, hashes.SHA256())) # ECDSA signatures are random: a new CSR never repeats
rec = Pending(
endpoint=ep.name, replacing=seen.fingerprint, replacing_key=seen.key,
key=hashlib.sha256(key.public_key().public_bytes(DER, SPKI)).hexdigest(),
key_pem=key.private_bytes(PEM, serialization.PrivateFormat.PKCS8,
serialization.NoEncryption()).decode(),
request=json.dumps({"name": ep.name, "csr": csr.public_bytes(PEM).decode()}),
idempotency_key="certrun-" + secrets.token_hex(16),
created=datetime.now(UTC).isoformat(timespec="seconds"))
save(state, rec)
return rec
def load(state: Path, name: str) -> Pending | None:
try:
return Pending(**json.loads((state / "pending" / f"{name}.json").read_text()))
except FileNotFoundError:
return None
def save(state: Path, rec: Pending) -> None:
folder = state / "pending"
folder.mkdir(mode=0o700, parents=True, exist_ok=True) # parents: 0700 state from install.sh
tmp = folder / f".{rec.endpoint}.tmp"
with os.fdopen(os.open(tmp, os.O_WRONLY | os.O_CREAT | os.O_TRUNC, 0o600), "w") as f:
json.dump(dataclasses.asdict(rec), f)
f.flush()
os.fsync(f.fileno())
os.replace(tmp, folder / f"{rec.endpoint}.json")
sync_dir(folder) # the new name is on disk too, not only the bytes
def drop(state: Path, name: str) -> None:
(state / "pending" / f"{name}.json").unlink()
sync_dir(state / "pending")
def sync_dir(folder: Path) -> None:
fd = os.open(folder, os.O_RDONLY | os.O_DIRECTORY)
try:
os.fsync(fd)
finally:
os.close(fd)

Why the exact bytes: a certificate request (CSR) is signed, and ECDSA signatures are random, so rebuilding it gives new bytes. A CA that follows the IETF Idempotency-Key draft fingerprints the whole request and answers 422 to a known key with another body; the lab CA does, and a test proves that a rebuilt CSR is refused. The key sits unencrypted, like the PEM the server loads, only until the rotation finishes; generating it on the target host, so only the CSR travels, is the production refinement.

ca.py retries a 429, 500, 502, 503 or 504, network errors and timeouts: after exactly the Retry-After the server named, unless it exceeds max_retry_after, which gives up and leaves the endpoint to the next run; otherwise after a jittered backoff. A CA certificate that does not verify, and any other status, is not retried. Redirects are refused: urllib would follow a 302 as a GET and copy the Authorization header to wherever Location points, plain HTTP included (a test proves both), so NoRedirects returns None and the 3xx becomes an error.

src/certrun/ca.py
"""The CA API client. Every attempt of one rotation sends the same Idempotency-Key with the same
request bytes, so a request whose answer was lost is answered from the CA's record instead of
issuing a second certificate. Redirects are refused: the bearer token only ever goes to ca.url."""
import asyncio
import http.client
import json
import logging
import random
import ssl
import urllib.error
import urllib.request
from .config import CA
from .errors import TargetError, TransientError
from .retry_after import parse_retry_after
log = logging.getLogger("certrun.ca")
RETRYABLE = {429, 500, 502, 503, 504}
class NoRedirects(urllib.request.HTTPRedirectHandler):
def redirect_request(self, req, fp, code, msg, headers, newurl): # type: ignore[no-untyped-def]
return None # urllib then raises HTTPError for the 3xx, and the token stays here
def post(ca: CA, token: str, key: str, body: str) -> str:
"""One attempt; blocking. The certificate PEM, or an error that says what should happen next."""
https = urllib.request.HTTPSHandler(context=ssl.create_default_context(cafile=ca.ca_file))
opener = urllib.request.build_opener(NoRedirects, https)
req = urllib.request.Request(
f"{ca.url}/certificates", data=body.encode(), method="POST",
headers={"Content-Type": "application/json", "Authorization": f"Bearer {token}",
"Idempotency-Key": key})
try:
with opener.open(req, timeout=ca.timeout) as resp:
answer = json.load(resp)
except urllib.error.HTTPError as err:
if err.code in RETRYABLE:
raise TransientError(f"HTTP {err.code}", parse_retry_after(err.headers["Retry-After"])) from None
moved = f" to {err.headers['Location']}, not followed" if 300 <= err.code < 400 else ""
raise TargetError(f"the CA answered HTTP {err.code}{moved}") from None
except urllib.error.URLError as err:
if isinstance(err.reason, ssl.SSLError): # the CA's certificate does not verify: not transient
raise TargetError(f"CA TLS: {err.reason}") from None
raise TransientError(repr(err.reason)) from None
except (OSError, http.client.HTTPException, ValueError) as err: # timeouts, resets, a garbled body
raise TransientError(repr(err)) from None
if not isinstance(answer, dict) or not isinstance(answer.get("pem"), str):
raise TargetError("the CA's answer holds no certificate")
return answer["pem"]
async def issue(ca: CA, token: str, key: str, body: str, name: str) -> str:
for attempt in range(1, ca.attempts + 1):
try:
return await asyncio.to_thread(post, ca, token, key, body)
except TransientError as err:
if attempt == ca.attempts:
raise
wait = random.uniform(0, 0.5 * 2**attempt) # full jitter: up to 1 s, 2 s, 4 s
if err.retry_after is not None:
if err.retry_after > ca.max_retry_after: # give up now; the next run tries again
raise TransientError(f"{err}, Retry-After {err.retry_after:.0f} s is too long") from None
wait = err.retry_after # the server said when: not sooner
log.warning("CA call failed, retrying", extra={"fields": {
"endpoint": name, "error": str(err), "attempt": attempt, "wait_s": round(wait, 1)}})
await asyncio.sleep(wait)
raise AssertionError("not reached: attempts is at least 1")

rotate.py resumes the record or starts one, asks the CA, and checks that the certificate carries the recorded key. After the install and the reload it asks the endpoint what it serves, every half second for up to reload_wait seconds. If the new certificate is served and verifies, finish() removes the backup (it holds the old private key), writes the audit line, then deletes the record; a stop in between leaves the record, so the next run finishes the job. If not, roll_back() marks the record rolled back first (so no later run asks the CA again), renames the backup into place, reloads and checks. A stop after the reload but before the check leaves a deployed record, and the next run checks what is served: the audit line if it verifies, roll_back() if not.

src/certrun/rotate.py
"""Rotate one endpoint: never issue twice, never deploy blind. The request is on disk before it is
sent (pending.py); the certificate must match the recorded key; the record says "deployed, not yet
checked" before the deploy, which keeps a backup; the server is reloaded and checked; a failed
check is rolled back; only then is the change audited."""
import asyncio
import dataclasses
import hashlib
import logging
from typing import NoReturn
from cryptography import x509
from . import ca, obs, pending, remote, tlscheck
from .config import Config, Endpoint
from .errors import TargetError, TransientError
from .pending import Pending
from .tlscheck import DER, Seen
log = logging.getLogger("certrun")
async def rotate(ep: Endpoint, rec: Pending | None, seen: Seen, cfg: Config, token: str) -> Seen:
"""rec: the record a stopped run left, resumed as it is; None starts a new rotation."""
rec = rec or pending.start(cfg.paths.state, ep, seen)
if rec.key == rec.replacing_key:
raise TargetError("the new key is the key being replaced; not requesting")
pem = await ca.issue(cfg.ca, token, rec.idempotency_key, rec.request, ep.name)
cert = x509.load_pem_x509_certificate(pem.encode())
if tlscheck.key_id(cert) != rec.key:
raise TargetError("the CA returned a certificate for another key; not deployed")
new = hashlib.sha256(cert.public_bytes(DER)).hexdigest()
host = cfg.hosts[ep.host]
rec = dataclasses.replace(rec, deployed=True) # from here the new file may be live, unchecked
pending.save(cfg.paths.state, rec)
await asyncio.to_thread(remote.install, host, ep, (pem + rec.key_pem).encode(), backup_of(ep, rec), cfg)
await asyncio.to_thread(remote.reload, host, cfg)
try:
now = await served(ep, cfg, new)
why = now.problem or (None if now.fingerprint == new else "the reload did not take effect")
except TransientError as err: # no answer at all after the reload
why = str(err)
if why is None:
await finish(ep, rec, now, cfg)
return now
await roll_back(ep, rec, why, cfg)
async def roll_back(ep: Endpoint, rec: Pending, why: str, cfg: Config) -> NoReturn:
"""Put the backup back, reload, check. Marked first, so no later run asks the CA again for
this rotation: a person decides. Also what the next run does with a failing deployed record."""
pending.save(cfg.paths.state, dataclasses.replace(rec, rolled_back=why))
host = cfg.hosts[ep.host]
await asyncio.to_thread(remote.rollback, host, ep, backup_of(ep, rec), cfg)
await asyncio.to_thread(remote.reload, host, cfg)
back = await served(ep, cfg, rec.replacing)
state = "rolled back" if back.fingerprint == rec.replacing else "ROLLBACK FAILED"
raise TargetError(f"the new certificate failed its check ({why}); {state}, serves {back.fingerprint[:12]}")
def backup_of(ep: Endpoint, rec: Pending) -> str:
return f"{ep.pem_path}.certrun-{rec.idempotency_key[-8:]}" # one backup per rotation
async def served(ep: Endpoint, cfg: Config, want: str) -> Seen:
"""What ep serves once a reload has taken effect: checked every half second until it serves
`want`, for at most policy.reload_wait seconds (a graceful reload takes a moment)."""
loop = asyncio.get_running_loop()
end = loop.time() + cfg.policy.reload_wait
while True:
seen = await tlscheck.check(ep, cfg, asyncio.Semaphore(1))
if seen.fingerprint == want or loop.time() >= end:
return seen
await asyncio.sleep(0.5)
async def finish(ep: Endpoint, rec: Pending, now: Seen, cfg: Config, recovered: bool = False) -> None:
"""The new certificate is live and verified: remove the backup (it holds the old private key),
record the change, then delete the record. Anything that stops this half way leaves the record,
and the next run finishes the job: the backup is never left without a record naming it."""
try:
await asyncio.to_thread(remote.forget, cfg.hosts[ep.host], backup_of(ep, rec), cfg)
except (TargetError, TransientError) as err:
raise type(err)(f"the new certificate is live; the old key file stays until the next run: {err}") from None
obs.audit(cfg.paths.audit, {
"action": "rotate", "endpoint": ep.name, "host": ep.host, "old": rec.replacing,
"new": now.fingerprint, "old_key": rec.replacing_key, "new_key": rec.key,
"not_after": now.not_after.isoformat(), "idempotency_key": rec.idempotency_key,
**({"recovered": True} if recovered else {})})
pending.drop(cfg.paths.state, ep.name)
log.info("rotated, audit line written", extra={"fields": {"endpoint": ep.name, "recovered": recovered}})

Two more modules are in the lesson files. runner.py surveys every endpoint and host at once in a TaskGroup, then acts one endpoint at a time, most urgent first; any exception in one endpoint's work, even an unexpected one, becomes that endpoint's failure, so 70 is left for bugs outside endpoint handling. stop.py handles SIGTERM the asyncio way. "Failure design" raised from a signal.signal handler, but in an event loop that exception would surface inside the loop's own code; so the handler is registered with loop.add_signal_handler and cancels the main task, and CancelledError arrives at the current await. A to_thread call cannot be cancelled, so asyncio.run waits for it, bounded by that call's timeouts (a CA read 2 s; one SSH call about 25 s: connect, banner and login 5 s each, then 5 s per idle read of its output and of its errors) and finally by TimeoutStopSec.

Release, install and the unit

release.sh builds the wheel, installs it and the hash-checked locked dependencies into a relocatable release/venv as copies, signs the wheel with cosign and writes an SBOM. The signing key is made for this release in a temporary directory and deleted when the script ends; only cosign.pub survives, next to release/, not in it. The offline signing of "Release gates" carries the same caveat (a transparency log is the production form). In production the key lives in a KMS or HSM, or signing is keyless: never on the runner.

release.sh
#!/usr/bin/env bash
# Build one release of the certrun core into release/: the wheel and its signature bundle, a venv
# with the locked dependencies and the wheel, and an SBOM of that venv. The signing key is made
# for this one release in a private temporary directory and deleted when the script ends; only
# its public key survives, as cosign.pub next to release/ (not in it), for install.sh to pin.
# In production the key lives in a KMS or HSM, or signing is keyless: never on the runner.
# Needs bin/uv, bin/cosign and bin/syft (lab/get-tools.sh).
set -euo pipefail
cd "$(dirname "$0")"
whl=certrun-0.1.0-py3-none-any.whl
[[ ! -e release ]] || { echo "release.sh: release/ exists; a release is never overwritten" >&2; exit 1; }
keys=$(mktemp -d)
trap 'rm -rf -- "$keys"' EXIT
./bin/uv lock --check -q # the lock still matches pyproject.toml
rm -rf dist
./bin/uv build -q --wheel # dist/$whl
./bin/uv export -q --locked --no-dev --no-emit-project -o dist/requirements.txt # with hashes
# Every file is copied (no links into uv's cache), so the release owns its bytes; the venv is
# relocatable because install.sh moves it to /opt.
./bin/uv venv -q --relocatable --python "$(command -v python3)" release/venv
./bin/uv pip install -q --link-mode=copy --python release/venv/bin/python \
--require-hashes -r dist/requirements.txt
./bin/uv pip install -q --link-mode=copy --python release/venv/bin/python --no-deps "dist/$whl"
cp "dist/$whl" release/
# The password protects the key file for the few seconds it exists; only cosign sees it.
pass=$(openssl rand -hex 16)
COSIGN_PASSWORD=$pass ./bin/cosign generate-key-pair --output-key-prefix "$keys/cosign" 2>/dev/null
COSIGN_PASSWORD=$pass ./bin/cosign sign-blob --yes --key "$keys/cosign.key" \
--signing-config offline-signing.json --bundle "release/$whl.sigstore.json" "release/$whl"
cp "$keys/cosign.pub" cosign.pub
chmod -R u=rwX,go=rX release # readable by the service account; install.sh makes it root's
./bin/syft scan -q dir:release/venv -o cyclonedx-json=release/sbom.cdx.json
printf 'release: %s, %s components in the SBOM; signing key deleted, public key in cosign.pub\n' \
"$whl" "$(jq '.components | length' release/sbom.cdx.json)"

A signature on the wheel says nothing about the installed copies Python will import. verify_core.py compares them with the signed wheel's RECORD and refuses extra files. certrun.sh runs it with the system python3 -I -S, so nothing in the venv (a .pth file in site-packages runs at startup) runs before the check. The gate checks the core's own files only: the interpreter, the dependencies, pyvenv.cfg and site hooks such as a .pth file rely on root ownership alone. The bytecode is root's too: install.sh compiles it next to the source in checked-hash mode (a .pyc is used only while it matches the source beside it, which the gate checked), and the core runs with -B, so it writes none. Under -I, PYTHONDONTWRITEBYTECODE=1 would be ignored; the lab checked both.

verify_core.py
"""Is the installed core the signed wheel? Every file the wheel's RECORD lists for the certrun
package must be installed in VENV with that hash, and nothing else may sit in the package
directory. Run by certrun.sh after cosign has verified the wheel, with the system python -I -S,
so nothing in the venv (a .pth file in site-packages runs at startup) runs before the check.
Standard library only. The dependencies are not checked: root ownership of the release is what
keeps them as installed. Usage: python3 -I -S verify_core.py WHEEL VENV exit 0 same files, 1 not"""
import base64
import csv
import hashlib
import io
import sys
import zipfile
from pathlib import Path
PACKAGE = "certrun/"
def main(wheel: str, venv: str) -> int:
site = next(Path(venv).glob("lib/python3*/site-packages"), None)
if site is None:
print(f"verify_core: no site-packages in {venv}", file=sys.stderr)
return 1
with zipfile.ZipFile(wheel) as zf:
record = next(n for n in zf.namelist() if n.endswith(".dist-info/RECORD"))
rows = csv.reader(io.StringIO(zf.read(record).decode()))
signed = {path: digest for path, digest, _size in rows if path.startswith(PACKAGE)}
problems = []
for path, digest in sorted(signed.items()):
algorithm, _, expected = digest.partition("=")
try:
data = (site / path).read_bytes()
except OSError:
problems.append(f"{path}: missing")
continue
actual = base64.urlsafe_b64encode(hashlib.new(algorithm, data).digest()).rstrip(b"=")
if actual.decode() != expected:
problems.append(f"{path}: differs from the signed wheel")
for file in sorted((site / PACKAGE).rglob("*")):
name = file.relative_to(site).as_posix()
if file.is_file() and "__pycache__" not in file.parts and name not in signed:
problems.append(f"{name}: not in the signed wheel")
for problem in problems:
print(f"verify_core: {problem}", file=sys.stderr)
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1], sys.argv[2]))
deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ ./release.sh ls cosign.*
Using payload from: release/certrun-0.1.0-py3-none-any.whl Signing artifact... Wrote bundle to file release/certrun-0.1.0-py3-none-any.whl.sigstore.json release: certrun-0.1.0-py3-none-any.whl, 31 components in the SBOM; signing key deleted, public key in cosign.pub cosign.pub
$ jq -r ".components[].type" release/sbom.cdx.json | sort | uniq -c jq -r ".components[] | select(.purl // \"\" | startswith(\"pkg:pypi\")) | \"\(.name) \(.version)\"" release/sbom.cdx.json | sort
23 file 8 library bcrypt 5.0.0 certrun 0.1.0 cffi 2.1.1 cryptography 50.0.1 invoke 3.0.3 paramiko 5.0.0 pycparser 3.0 pynacl 1.6.2

ls finds only cosign.pub: the private key is gone. Syft listed 8 packages (library), certrun and its seven locked dependencies, plus 23 file entries it hashed.

The gate is code, and a gate the service account can edit is no gate: owning certrun.sh, the verifier, the public key or the directory holding the release, it could replace any of them. install.sh puts everything that decides what runs under /opt/scr-capstone, root's all the way up, and lets the account write only to /var/lib/scr-capstone. The CA token and its SSH key are root's, readable by its group only. Metrics live outside any home directory, because Ubuntu 26.04 homes are 0750 and node_exporter could not reach inside.

install.sh
#!/usr/bin/env bash
# Install the certrun runner for the service ACCOUNT, as root, from the project directory:
# /opt/scr-capstone root:root, 0755: certrun.sh, the release, the gate and its trust
# anchors, the settings. The account can run all of it, change none.
# /var/lib/scr-capstone state/ (locks, pending rotations, audit; 0700) and textfile/ (0755,
# for node_exporter): the only places the account can write.
# /etc/systemd/system the unit and timer, with User=ACCOUNT. The timer is not enabled.
# Undo: sudo rm -rf /opt/scr-capstone /var/lib/scr-capstone \
# /etc/systemd/system/scr-capstone-certrun@.{service,timer} && sudo systemctl daemon-reload
set -euo pipefail
((EUID == 0)) || { echo "install.sh: run as root" >&2; exit 1; }
account=${1:?usage: install.sh ACCOUNT}
group=$(id -gn -- "$account")
opt=/opt/scr-capstone var=/var/lib/scr-capstone sshd=/run/scr-capstone-sshd
cd -- "$(dirname -- "${BASH_SOURCE[0]}")"
[[ ! -e $opt ]] || { echo "install.sh: $opt exists; remove it first" >&2; exit 1; }
install -d -m 0755 "$opt" "$opt/bin" "$opt/etc"
install -m 0755 certrun.sh "$opt/"
install -m 0644 verify_core.py "$opt/"
install -m 0755 bin/cosign "$opt/bin/"
cp -R release "$opt/release" && chown -R root:root "$opt/release"
# The bytecode, compiled here by root next to the source, so the account cannot write it (and
# certrun.sh runs python -B: nothing is written at run time). checked-hash: a .pyc is used only
# while it matches the source file beside it, and the gate checks that source.
"$opt/release/venv/bin/python" -I -m compileall -q -j 0 --invalidation-mode checked-hash "$opt/release/venv/lib"
install -m 0644 cosign.pub etc/lab.toml "$opt/etc/" # the pinned signing key, the settings
install -m 0644 lab/pki/ca.pem "$opt/etc/lab-ca.pem"
install -m 0644 "$sshd/known_hosts" "$opt/etc/known_hosts"
# Secrets the account must read and must not change: root's, readable by its group only.
install -m 0640 -g "$group" lab/pki/ca-token "$opt/etc/ca-token"
install -m 0640 -g "$group" "$sshd/id_certrun" "$opt/etc/id_certrun"
install -d -m 0755 "$var"
install -d -m 0700 -o "$account" -g "$group" "$var/state"
install -d -m 0755 -o "$account" -g "$group" "$var/textfile"
sed "s/^User=.*/User=$account/" systemd/scr-capstone-certrun@.service \
>/etc/systemd/system/scr-capstone-certrun@.service
install -m 0644 systemd/scr-capstone-certrun@.timer /etc/systemd/system/
systemctl daemon-reload
find "$opt" "$var" -maxdepth 2 -not -path "$opt/release/*" -printf '%M %u:%g %p\n'
deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ sudo ./install.sh "$USER" find /opt/scr-capstone/release/venv -name "*.pyc" -printf "%u\n" | sort | uniq -c stat -c "%U %a %n" /opt/scr-capstone/release/venv/lib/python3.14/site-packages/certrun/__pycache__/runner.*.pyc
drwxr-xr-x root:root /opt/scr-capstone drwxr-xr-x root:root /opt/scr-capstone/bin -rwxr-xr-x root:root /opt/scr-capstone/bin/cosign -rw-r--r-- root:root /opt/scr-capstone/verify_core.py -rwxr-xr-x root:root /opt/scr-capstone/certrun.sh drwxr-xr-x root:root /opt/scr-capstone/etc -rw-r--r-- root:root /opt/scr-capstone/etc/lab.toml -rw-r--r-- root:root /opt/scr-capstone/etc/cosign.pub -rw-r----- root:deploy /opt/scr-capstone/etc/id_certrun -rw-r--r-- root:root /opt/scr-capstone/etc/known_hosts -rw-r----- root:deploy /opt/scr-capstone/etc/ca-token -rw-r--r-- root:root /opt/scr-capstone/etc/lab-ca.pem drwxr-xr-x root:root /opt/scr-capstone/release drwxr-xr-x root:root /var/lib/scr-capstone drwx------ deploy:deploy /var/lib/scr-capstone/state drwxr-xr-x deploy:deploy /var/lib/scr-capstone/textfile 242 root root 644 /opt/scr-capstone/release/venv/lib/python3.14/site-packages/certrun/__pycache__/runner.cpython-314.pyc
$ touch /opt/scr-capstone/release/venv/lib/python3.14/site-packages/certrun/extra.py mv /opt/scr-capstone/release /opt/scr-capstone/mine echo "# hotfix" >>/opt/scr-capstone/certrun.sh
touch: cannot touch '/opt/scr-capstone/release/venv/lib/python3.14/site-packages/certrun/extra.py': Permission denied mv: cannot move '/opt/scr-capstone/release' to '/opt/scr-capstone/mine': Permission denied -bash: line 3: /opt/scr-capstone/certrun.sh: Permission denied

All three fail, the mv too: renaming a directory needs write access to its parent. The 242 .pyc files install.sh compiled are root's, 0644; the lab also checked that no later run wrote one. The threat model: this layout stops the service account, or code running as it, from changing what runs; the gate stops a core nobody signed (an edit in place, a corrupted copy, a development build) from running. Neither protects against root or whoever holds the signing key. The undo is in install.sh's header.

The unit is a template: in scr-capstone-certrun@lab.service, the part after @ (%i) is the destination. Type=oneshot makes systemctl start wait for the run. For a oneshot unit, death by SIGTERM counts as a failure (systemd.service(5), SuccessExitStatus=), so every shutdown would look like a broken job; listing SIGTERM makes it a clean stop. KillMode=mixed sends SIGTERM to certrun.sh alone and SIGKILL to whatever is left once it has exited, which is why the script waits for the core. Restart= wires the contract to the scheduler: 75 is retried in 15 minutes, every other failure waits for a person.

systemd/scr-capstone-certrun@.service
# One apply run of certrun for the destination %i, e.g. scr-capstone-certrun@lab.service.
# Started by scr-capstone-certrun@%i.timer, or by hand with systemctl start.
[Unit]
Description=certrun apply %i: renew expiring TLS certificates
[Service]
Type=oneshot
# install.sh puts the service account here (in production, a system account of its own)
User=certrun
ExecStart=/opt/scr-capstone/certrun.sh apply %i
Environment=CERTRUN_BUDGET=120
# A stop (shutdown, deploy, an operator) ends the run by SIGTERM. For Type=oneshot that counts
# as a failure unless it is listed here, and a failure is what OnFailure= alerts react to.
SuccessExitStatus=SIGTERM
# SIGTERM to certrun.sh only, which forwards it to the core; SIGKILL to whatever is left once
# certrun.sh has exited, or after TimeoutStopSec. A call running in a thread cannot be cancelled,
# so a stop waits for it: 30 s is above the longest, one SSH call of about 25 s (connect, banner
# and auth 5 s each, then 5 s per idle read of its output and of its errors).
KillMode=mixed
TimeoutStopSec=30
# 75 means "temporary, run again later": retry it in 15 minutes. The other failures need a
# person, not a loop.
Restart=on-failure
RestartSec=15min
RestartPreventExitStatus=1 2 70 77 78
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ReadWritePaths=/var/lib/scr-capstone

A timer is a second unit that starts the service on a schedule (systemd.timer(5)): every day at 00:00 and 12:00, catching up a run missed while the host was down, spread over half an hour across runners.

systemd/scr-capstone-certrun@.timer
# Twice a day for the destination %i. In production:
# systemctl enable --now scr-capstone-certrun@lab.timer
[Unit]
Description=certrun apply %i, twice a day
[Timer]
OnCalendar=*-*-* 00,12:00:00
# A run missed while the host was down happens at the next boot; many runners spread their
# requests to the CA over half an hour instead of all arriving at 00:00:00.
Persistent=true
RandomizedDelaySec=30min
[Install]
WantedBy=timers.target
deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ systemd-analyze verify scr-capstone-certrun@lab.service scr-capstone-certrun@lab.timer 2>&1 | grep -F scr-capstone || echo "verify: nothing to report for scr-capstone-certrun@lab" systemd-analyze calendar --iterations=2 "*-*-* 00,12:00:00" systemctl is-enabled scr-capstone-certrun@lab.timer; systemctl is-active scr-capstone-certrun@lab.timer
verify: nothing to report for scr-capstone-certrun@lab Normalized form: *-*-* 00,12:00:00 Next elapse: Tue 2026-09-29 12:00:00 UTC From now: 8h left Iteration #2: Wed 2026-09-30 00:00:00 UTC From now: 20h left disabled inactive

systemd-analyze verify loads the units as systemd would (the grep keeps lines about ours: none), and calendar shows when the timer would fire. It stays disabled and inactive (is-active exits 3): an armed timer could fire mid-lesson and race your manual runs, so the lesson starts the service directly. In production, systemctl enable --now scr-capstone-certrun@lab.timer arms it.

With the release installed, the first failures a person meets:

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ /opt/scr-capstone/certrun.sh rotate lab
usage: certrun.sh plan|apply DESTINATION
$ /opt/scr-capstone/certrun.sh plan prod
{"ts":"2026-09-29T03:14:17Z","level":"error","logger":"certrun.sh","msg":"no settings file /opt/scr-capstone/etc/prod.toml","destination":"prod"}
$ sed "s/^renew_within_days = 30/renew_within_days = \"30\"/" etc/lab.toml | sudo tee /opt/scr-capstone/etc/typo.toml >/dev/null /opt/scr-capstone/certrun.sh plan typo; echo "exit status $?" sudo rm /opt/scr-capstone/etc/typo.toml
{"ts": "2026-09-29T03:14:17.555+00:00", "level": "error", "logger": "certrun", "msg": "config error", "run": "20260929T031417Z-912137", "destination": "typo", "error": "policy.renew_within_days: expected int, got '30'"} exit status 78

A mistyped mode is 2. An unknown destination is 78 from the script. A settings file with renew_within_days = "30" passes the gate, and the core's loader names the setting, the type and the value, then exits 78, which Restart= does not retry.

End to end

Plan first. It looks at everything and changes nothing, not even the metrics:

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ /opt/scr-capstone/certrun.sh plan lab 2>plan.log; echo "exit status $?"
api ok 200 days left billing rotate 12 days left legacy rotate EXPIRED 3 days ago shop rotate 20 days left admin failed serves a certificate that does not verify: Hostname mismatch, certificate is not valid for 'admin.lab.internal'. search failed web2: host key mismatch: pinned SHA256:+E80JU0VjFDBurfzC9sw6N4L2bIMeaIjk3//CSd7iDQ, offered SHA256:gCq4H8S0hfCbGZn7XZPzjPIqN/N/mra0Cl7Hd5FBOmM certrun plan lab: 1 ok, 3 rotate, 2 failed (exit 1) exit status 1
$ head -n 1 plan.log jq -r "[.level, .msg, .endpoint // \"-\"] | join(\" \")" plan.log | sort | uniq -c
{"ts": "2026-09-29T03:14:18.548+00:00", "level": "info", "logger": "certrun", "msg": "ok", "run": "20260929T031417Z-912187", "destination": "lab", "endpoint": "api", "detail": "200 days left"} 1 error failed admin 1 error failed search 1 info ok api 1 info rotate billing 1 info rotate legacy 1 info rotate shop

Three to rotate, one fine, two for a person, so the plan exits 1, as the job would. The search line gives both fingerprints, which is what a person needs to decide. On stderr, one JSON object per event, with the run ID in each.

Now an apply stopped in a CA call: the step waits until the CA has issued legacy's certificate and is still holding the answer, then sends SIGTERM. At a prompt, Bash may also print a job notice.

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ /opt/scr-capstone/certrun.sh apply lab >apply1.out 2>apply1.log & pid=$! timeout 30 bash -c "until grep -q \"issued legacy\" lab/ca.log; do sleep 0.1; done" kill -TERM "$pid" wait "$pid"; echo "certrun exit status: $?"
certrun exit status: 143
$ jq -r "[.logger, .level, .msg, .endpoint // \"\", .error // \"\"] | join(\" \")" apply1.log
certrun error certificate expired before any run renewed it: rotating now legacy certrun.ca warning CA call failed, retrying legacy HTTP 500 certrun warning SIGTERM: cancelling the run certrun warning stopped; a rotation in progress stays in pending/ for the next run certrun.sh warning stopped by SIGTERM
$ ls -d /tmp/certrun-* 2>&1 flock -n /var/lib/scr-capstone/state/lab.lock echo "lock: free" pgrep -fa "[v]env/bin/python -I" || echo "core: not running" ls -l /var/lib/scr-capstone/state/lab/pending jq -cM "{endpoint, idempotency_key, replacing: .replacing[:12], request_bytes: (.request | length)}" /var/lib/scr-capstone/state/lab/pending/legacy.json cat /var/lib/scr-capstone/textfile/certrun-lab-run.prom
ls: cannot access '/tmp/certrun-*': No such file or directory lock: free core: not running total 4 -rw------- 1 deploy deploy 1171 Sep 29 03:14 legacy.json {"endpoint":"legacy","idempotency_key":"certrun-3d4dfec45f85a7cbc9f19d0ad54fe38b","replacing":"3c2c6b215c13","request_bytes":473} # TYPE certrun_last_run_seconds gauge certrun_last_run_seconds{destination="lab"} 1790651661 # TYPE certrun_last_run_status gauge certrun_last_run_status{destination="lab"} 143

legacy went first, being expired. Its first request got a 500; the retry was in flight when SIGTERM arrived, the core stopped, and certrun.sh re-raised: 143. No run directory, the lock is free, no core process, the run metric says 143, and one file remains on purpose: pending/legacy.json, holding the 473 request bytes the CA has already answered.

Next, the unit, started without waiting; the step waits for the web server's second reload request (billing's deploy) and runs systemctl stop there, as a shutdown would.

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ sudo systemctl start --no-block scr-capstone-certrun@lab.service timeout 60 bash -c "until [ \"\$(grep -c SIGHUP lab/tlsfarm.log)\" -ge 2 ]; do sleep 0.1; done" sudo systemctl stop scr-capstone-certrun@lab.service systemctl show -p ActiveState -p Result -p ExecMainStatus scr-capstone-certrun@lab.service
ActiveState=inactive Result=success ExecMainStatus=0
$ journalctl -u scr-capstone-certrun@lab.service -I -o cat --no-pager | grep -v "^{" journalctl -u scr-capstone-certrun@lab.service -I -o cat --no-pager | grep "^{" | jq -r "[.level, .msg, .endpoint // \"\", .error // \"\"] | join(\" \")"
Starting scr-capstone-certrun@lab.service - certrun apply lab: renew expiring TLS certificates... scr-capstone-certrun@lab.service: Deactivated successfully. Stopped scr-capstone-certrun@lab.service - certrun apply lab: renew expiring TLS certificates. info rotated, audit line written legacy warning CA call failed, retrying billing HTTP 429 warning CA call failed, retrying billing TimeoutError('The read operation timed out') warning SIGTERM: cancelling the run warning stopped; a rotation in progress stays in pending/ for the next run warning stopped by SIGTERM
$ timeout 10 bash -c "until [ \"\$(grep -c reloaded lab/tlsfarm.log)\" -ge 2 ]; do sleep 0.1; done" ls /var/lib/scr-capstone/state/lab/pending jq -r .endpoint /var/lib/scr-capstone/state/lab/audit.jsonl jq -r .key /var/lib/scr-capstone/state/lab/pending/billing.json | cut -c1-12 openssl s_client -connect 127.0.0.1:18942 -servername billing.lab.internal </dev/null 2>/dev/null | openssl x509 -noout -pubkey | openssl pkey -pubin -outform DER | sha256sum | cut -c1-12
billing.json legacy 2c4fea390f69 2c4fea390f69

inactive and success, "Deactivated successfully": that is SuccessExitStatus=SIGTERM; without it the same stop leaves the unit failed with result signal (the lab checked both). legacy resumed: the recorded bytes went out again, the CA replayed its answer, and deploy, reload, check and audit followed. billing got a 429, a slow answer (TimeoutError) and a replay, and was stopped after its deploy. Once that reload landed, billing serves a certificate with the key hash in pending/billing.json, and has no audit line. Now let the unit run to the end:

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ sudo systemctl start scr-capstone-certrun@lab.service
Job for scr-capstone-certrun@lab.service failed because the control process exited with error code. See "systemctl status scr-capstone-certrun@lab.service" and "journalctl -xeu scr-capstone-certrun@lab.service" for details.
$ journalctl -u scr-capstone-certrun@lab.service -I -o cat --no-pager | grep -v "^{" journalctl -u scr-capstone-certrun@lab.service -I -o cat --no-pager | grep "^{" | jq -r "[.level, .msg, .endpoint // \"\", .error // \"\"] | join(\" \")"
Starting scr-capstone-certrun@lab.service - certrun apply lab: renew expiring TLS certificates... api ok 200 days left billing rotated a stopped run's deploy, audited now; now 100ceb26b57c legacy ok 90 days left shop failed the new certificate failed its check (unable to get local issuer certificate); rolled back, serves 72ae54e83b74 admin failed serves a certificate that does not verify: Hostname mismatch, certificate is not valid for 'admin.lab.internal'. search failed web2: host key mismatch: pinned SHA256:+E80JU0VjFDBurfzC9sw6N4L2bIMeaIjk3//CSd7iDQ, offered SHA256:gCq4H8S0hfCbGZn7XZPzjPIqN/N/mra0Cl7Hd5FBOmM certrun apply lab: 2 ok, 1 rotated, 3 failed (exit 1) scr-capstone-certrun@lab.service: Main process exited, code=exited, status=1/FAILURE scr-capstone-certrun@lab.service: Failed with result 'exit-code'. Failed to start scr-capstone-certrun@lab.service - certrun apply lab: renew expiring TLS certificates. info rotated, audit line written billing info ok api info rotated billing info ok legacy error failed shop error failed admin error failed search
$ curl -s --cacert lab/pki/ca.pem https://127.0.0.1:18961/v1/stats; echo cat lab/ca.log
{"requests": 7, "issued": {"legacy": 1, "billing": 1, "shop": 1}, "replayed": 2} ca: listening on https://127.0.0.1:18961/v1 ca: POST legacy -> 500 ca: issued legacy, answering in 4 s ca: POST legacy -> 200 Idempotent-Replayed: true ca: POST legacy (4 s late) -> 201 ca: POST billing -> 429 Retry-After: 1 ca: issued billing, answering in 4 s ca: POST billing -> 200 Idempotent-Replayed: true ca: POST billing (4 s late) -> 201 ca: issued shop from the untrusted CA ca: POST shop -> 201

The run exited 1, so the unit entered the failed state (status=1/FAILURE), which is what OnFailure= reacts to, and Restart= did not retry. billing needed no CA request: its live certificate carries the record's key, so the run only wrote the audit line. shop shows the rollback: a certificate from an untrusted issuer failed the post-deploy check, and shop serves its old one again. The CA's log: seven requests, three certificates, two replays; each (4 s late) line is the CA answering a client that had stopped waiting.

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ /opt/scr-capstone/certrun.sh apply lab 2>apply3.log; echo "exit status $?" curl -s --cacert lab/pki/ca.pem https://127.0.0.1:18961/v1/stats; echo
api ok 200 days left billing ok 90 days left legacy ok 90 days left shop failed a deploy was rolled back (unable to get local issuer certificate); fix the cause, then delete pending/shop.json admin failed serves a certificate that does not verify: Hostname mismatch, certificate is not valid for 'admin.lab.internal'. search failed web2: host key mismatch: pinned SHA256:+E80JU0VjFDBurfzC9sw6N4L2bIMeaIjk3//CSd7iDQ, offered SHA256:gCq4H8S0hfCbGZn7XZPzjPIqN/N/mra0Cl7Hd5FBOmM certrun apply lab: 3 ok, 3 failed (exit 1) exit status 1 {"requests": 7, "issued": {"legacy": 1, "billing": 1, "shop": 1}, "replayed": 2}
$ jq -cM "{endpoint, recovered, old_key: .old_key[:12], new_key: .new_key[:12]}" /var/lib/scr-capstone/state/lab/audit.jsonl ls /var/lib/scr-capstone/state/lab/pending (cd /var/lib/scr-capstone/textfile && grep -v "^#" certrun-lab-run.prom certrun-lab.prom) ls -l /var/lib/scr-capstone/textfile
{"endpoint":"legacy","recovered":null,"old_key":"451e6fb854f0","new_key":"9a99a1eb2fb7"} {"endpoint":"billing","recovered":true,"old_key":"2c6ea2bc39f3","new_key":"2c4fea390f69"} shop.json certrun-lab-run.prom:certrun_last_run_seconds{destination="lab"} 1790651676 certrun-lab-run.prom:certrun_last_run_status{destination="lab"} 1 certrun-lab.prom:certrun_last_success_seconds{destination="lab"} 1790651676 certrun-lab.prom:certrun_needs_attention{destination="lab"} 3 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="api"} 1807931635 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="billing"} 1798427665 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="legacy"} 1798427659 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="shop"} 1792379636 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="admin"} 1803611636 certrun-lab.prom:certrun_certificate_expiry_seconds{destination="lab",endpoint="search"} 1791429236 total 8 -rw-r--r-- 1 deploy deploy 175 Sep 29 03:14 certrun-lab-run.prom -rw-r--r-- 1 deploy deploy 989 Sep 29 03:14 certrun-lab.prom

Nothing to do, and the CA's counters did not move: shop's rolled-back record blocks new requests until a person looks. One audit line per change, billing's marked recovered, each new key different from the old. certrun_last_run_seconds and certrun_last_run_status, written by certrun.sh on every apply that got the lock, say whether the job runs and how it last ended. certrun_last_success_seconds is when a run last handled every endpoint (exit 0 or 1); a 75 run leaves it alone, so a CA or host that stays down shows as its age. In "Observable jobs" success was exit 0, but here 1 is normal while any endpoint waits for a person, and a success stuck at 0 could never show a stopped timer; its age is that alert. certrun_needs_attention (3) pages a person, and the per-endpoint expiry catches a certificate near its end whatever the job does.

Last, a core that is not the release: a wheel with one byte appended, then a "hotfix" in an installed file (each needs sudo, which is the point):

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ w=/opt/scr-capstone/release/certrun-0.1.0-py3-none-any.whl sudo cp -p "$w" /tmp/good.whl && printf x | sudo tee -a "$w" >/dev/null /opt/scr-capstone/certrun.sh apply lab; echo "exit status $?" grep "^certrun_last_run_status" /var/lib/scr-capstone/textfile/certrun-lab-run.prom sudo mv /tmp/good.whl "$w"
{"ts":"2026-09-29T03:14:36Z","level":"error","logger":"certrun.sh","msg":"refusing to run: the core is not the signed release","destination":"lab"} WARNING: Skipping tlog verification is an insecure practice that lacks transparency and auditability verification for the blob. Error: failed to verify signature: could not verify message: invalid signature when validating ASN.1 encoded signature error during command execution: failed to verify signature: could not verify message: invalid signature when validating ASN.1 encoded signature exit status 77 certrun_last_run_status{destination="lab"} 77
$ f=/opt/scr-capstone/release/venv/lib/python3.14/site-packages/certrun/runner.py echo "# hotfix" | sudo tee -a "$f" >/dev/null /opt/scr-capstone/certrun.sh plan lab; echo "exit status $?" sudo sed -i "\$d" "$f" /opt/scr-capstone/certrun.sh plan lab >/dev/null 2>&1; echo "after the repair: exit status $?"
{"ts":"2026-09-29T03:14:36Z","level":"error","logger":"certrun.sh","msg":"refusing to run: the core is not the signed release","destination":"lab"} WARNING: Skipping tlog verification is an insecure practice that lacks transparency and auditability verification for the blob. Verified OK verify_core: certrun/runner.py: differs from the signed wheel exit status 77 after the repair: exit status 1

Both 77, recorded in the run metric, and the core never started. cosign refused the changed wheel; with the wheel intact (Verified OK), verify_core.py still refused, naming the edited file. The WARNING is cosign's reminder that the transparency log was skipped. After the repair, the run is back to its honest 1.

Try this

Fix search and shop the right way, before you clean up. Someone with console access to web2 reads its host key fingerprint there (here: ssh-keygen -lf /run/scr-capstone-sshd/host-web2.pub) and finds what web2 offered in the plan. Pinning is a root-owned change: in /opt/scr-capstone/etc/known_hosts, replace the [127.0.0.1]:18952 line with [127.0.0.1]:18952 and the first two fields of host-web2.pub (ssh-ed25519 AAAA...), then run /opt/scr-capstone/certrun.sh apply lab. Expected: search rotated, exit 1, 8 CA requests with "search": 1. The CA's fault for shop was a one-off, so delete /var/lib/scr-capstone/state/lab/pending/shop.json and apply again. Expected: shop rotated, "shop": 2, and certrun_needs_attention 1. Then explain what must change about admin (its certificate names www) before certrun may touch it, and why it does not decide that itself.

Clean up

Stop the lab servers and the hosts, reset the failed unit, remove the installation, and delete the project (everything left in it is yours):

deploy@web01:~/scr-capstone · Ubuntu 26.04 LTS
$ kill "$(cat lab/ca.pid)" "$(cat /srv/scr-capstone/tlsfarm.pid)" sudo systemctl reset-failed scr-capstone-certrun@lab.service sudo lab/hosts.sh stop sudo rm -rf /opt/scr-capstone /var/lib/scr-capstone /etc/systemd/system/scr-capstone-certrun@.service /etc/systemd/system/scr-capstone-certrun@.timer sudo systemctl daemon-reload systemctl list-units --all --no-legend "scr-capstone-*" | wc -l cd ~ && rm -rf scr-capstone
hosts: stopped 0

0 units left. To start over, run this clean-up without its last line and begin at the first terminal: a restarted CA fails the same way again.

Takeaway

Give an unattended job an entry point that owns the process, install whatever decides what runs where the job's account cannot change it, and write each change down before making it, so a retry, crash or stop at any step is finished by the next run instead of repeated or lost.

This is the end of the Scripting track. The CI/CD courses take the release half further. Start with "Software supply chain security" (/courses/ci-supply/) for signing, attestations and verification across a pipeline, then go deeper into SLSA, in-toto and Sigstore in "Software supply chain in depth" (/courses/supplychain-deep/). "Secure CI/CD with GitLab" (/courses/ci-gitlab/) puts gates like these into a GitLab pipeline.

Quick check
01A run was stopped by SIGTERM after the CA had issued legacy's certificate but before certrun heard the answer. The next run starts. What keeps the CA from issuing a second certificate?
Incorrect — The lock only stops two runs at once; the stopped run released it when it died, as the lab's "lock: free" showed.
Incorrect — The lab CA has no such rule, and certrun does not rely on one; a CA that had it would also block real renewals.
Incorrect — A new CSR has new bytes, because ECDSA signatures are random, and a CA that fingerprints requests answers 422; the test with a rebuilt CSR proved it.
Correct — The record was written before the first request, so the resend is byte for byte what the CA already answered; the lab showed "200 Idempotent-Replayed" and one certificate issued.
02The unit is Type=oneshot with KillMode=mixed, and systemctl stop sends SIGTERM to certrun.sh while the core is running. Why does the script forward the signal, wait for the core and re-raise, instead of calling exit 143 in its TERM trap?
Incorrect — exit works in a trap; the problem is what happens to the core once the script has gone.
Correct — With KillMode=mixed, whatever is left after the main process exits gets SIGKILL. And status 143 would be a plain failure, while SuccessExitStatus=SIGTERM makes a death by the signal a clean stop.
Incorrect — 137 (128 + 9) is SIGKILL's status, and a stop requested with systemctl stop never triggers Restart=.
Incorrect — The EXIT trap runs after exit too; the bats test showed cleanup and the 143 metric on the re-raise path as well.
03The lab CA's own certificate expires, while api's certificate still has 200 days left. The TLS check now fails with OpenSSL verify code 10, "certificate has expired". What does certrun do with api?
Incorrect — Code 10 names any certificate in the chain. Reading it as "the leaf expired" is the bug the lab's expired-CA test catches: a new leaf from the same CA would fail the same way.
Incorrect — Nothing about an expired certificate goes away by waiting; certrun keeps temporary for hosts and CAs that do not answer.
Correct — Expiry is read from the leaf's not_after; with 200 days left, code 10 must come from the chain, which only a person can fix.
Incorrect — The renewal window only applies once the chain verifies; a chain that does not verify is reported whatever the dates say.

Related