Back to the course
Test yourself

Advanced scripting for DevSecOps

Final exam · 47 questions · answers explained as you pick
Shell in production
18 questions
01A gate script starts with set -euo pipefail and shopt -s lastpipe, then counts with grep -F 'Failed password' -- "$log" | while IFS= read -r _; do failed=$((failed + 1)); done before it prints a verdict. The log it is given has mode 000. What happens?
Incorrect — That is what the process substitution version did. With lastpipe, grep is still a stage of a pipeline, so pipefail and errexit see its status.
Correct — Under pipefail the pipeline's status is grep's 2, errexit ends the script on that line, and no PASS can be printed: the gate fails closed.
Incorrect — It is the other way round: lastpipe applies when job control is off, the default in scripts, so the loop does run in the current shell.
Incorrect — errexit acts on the status of the whole pipeline, and pipefail makes that status grep's 2, so the script stops at that line.
02A drift check runs diff <(grep -v '^#' -- expected.conf) <(grep -v '^#' -- live.conf), then wait "$!", and treats a status above 1 as an error. One night it reported "no drift" although expected.conf had been deleted and live.conf held only comments. Why did the check pass?
Incorrect — diff opened both /dev/fd paths without trouble and read two empty streams; it never saw the missing file, so its 0 described what it read.
Incorrect — A wait without arguments waits for every child and returns 0; it reports none of their statuses.
Incorrect — grep exits 2 for an error such as a missing file; 1 means that no lines were selected, which is why the check accepts it.
Correct — Open each producer with exec {fd}< <(...), keep one PID per producer and wait for each, as the lesson did for two greps.
03An operator tries the stop logic by hand at an interactive prompt: setsid ./check-host.sh demo 30 0 &, then kill -TERM -- "-$!". the kill fails and the check keeps running, yet the same two lines work inside sweep.sh. What explains the difference?
Correct — In a script setsid need not fork and $! is the new group's leader; at a prompt & already makes a group, setsid forks, and the real group ID is a PID you never saw.
Incorrect — Background jobs of an interactive shell do not ignore SIGTERM; kill failed because no process group had that ID.
Incorrect — The -- already stops kill from reading -PID as an option, in scripts and at a prompt alike.
Incorrect — setsid still creates a new session; the problem is which PID $! holds, not whether the session was made.
04shellcheck -o check-set-e-suppressed,check-extra-masked-returns refresh-naive.sh reports SC2312 on the printf '%s\n' "$(fetch)" line and SC2310 on if ! install_list; then. The team wants install_list to stay correct however a later caller invokes it. Which change does that?
Incorrect — Called in a condition, the function runs with errexit off whatever the header says; the directive hides the bug.
Incorrect — set -E makes the ERR trap inherited; it does not bring errexit back inside a tested function.
Correct — Explicit checks work in any calling context, and a plain body=$(fetch) with its own check also clears SC2312.
Incorrect — The left side of a || list is a tested context too, so errexit stays off inside install_list.
05The fail-closed refresh.sh installs a new blocklist that passed its checks: at least one entry and no invalid line. The single entry is 0.0.0.0/0, and the firewall now blocks every source. What was missing?
Correct — Syntax checks cannot judge meaning; a floor on prefix length, ranges that may never be blocked and a limit on change against the current list can.
Incorrect — The entry pattern accepts any dotted quad with an optional prefix length, 0.0.0.0/0 included; the refresh installed it.
Incorrect — The temp file sits in the list's own directory, so mv is a rename and readers see the old or the new file whole.
Incorrect — An HTML page would contain lines that fail the pattern and be refused; this file was a single well-formed entry.
06The nightly blocklist refresh leaves error: tmp=$(mktemp -- "$list.XXXXXX") exited 1 on stderr, followed by at ./refresh.sh:26 in refresh_list and at ./refresh.sh:36 in main. How should the on-call engineer read it?
Incorrect — The stack reads from the failure outward; main is the name Bash gives the top level, not a function in the script.
Incorrect — Each at line is one frame of a single report; trace.sh reports from the main shell and exits.
Incorrect — mktemp creates the temp file before the download; it failed first, so no install happened.
Correct — caller frame 0 is the failing call and "main" is Bash's name for the top level; the unplanned failure stopped the run before the install, with a non-zero status.
07A timer starts a Type=oneshot unit whose script cleans up on SIGTERM and then re-raises it, a pattern copied from a long-running service that stops cleanly. Stopping the oneshot unit mid-run leaves code=killed, status=15/TERM and result 'signal' in its log, and the unit is marked failed. Why, and what fixes it?
Incorrect — The status is 15/TERM, not 9/KILL, and the result is 'signal', not 'timeout'; the cleanup finished in time.
Correct — systemd.service(5) counts a SIGTERM death as clean except for oneshot units; declaring TERM (or 143 for a script that exits 143) records the stop as a success.
Incorrect — Exit 143 is an ordinary failing status for every unit type unless SuccessExitStatus= names it.
Incorrect — The main process itself died of SIGTERM, as the status line shows; which children were signalled does not change that verdict.
08A container's entrypoint is a Bash script with trap cleanup EXIT and no TERM trap; it runs ./batch-worker in the foreground. On pod deletion the container stops within a second with exit code 143, but the worker's batch is cut off midway. What happened?
Incorrect — That is the no-trap case, which ends after the grace period with 137. Here the EXIT trap gave Bash a handler, so the signal arrived.
Incorrect — Bash does not forward signals to its children; the worker died because PID 1 exited, not from a forwarded signal.
Correct — An EXIT trap makes the stop fast, not graceful: forward SIGTERM to the worker and wait for it, or run the worker under an init such as tini.
Incorrect — The container exited within a second, long before a grace period could run out; more time changes nothing.
09A batch unit uses the default KillMode. Its script sets a flag in a TERM trap and is meant to stop at the next item boundary. Stopped one second into an item, the unit logs upload 1: start, then Terminated ./upload.sh "$item", then status=143/n/a. What does this sequence tell you?
Incorrect — A timeout would show "State 'stop-sigterm' timed out" and a result of 'timeout'; this stop was immediate.
Incorrect — Bash runs the trap after the foreground command returns, which is exactly what makes the item-boundary drain work.
Incorrect — The trap sets a flag and nothing more; the 143 came from the killed child through errexit, before any re-raise could run.
Correct — With mixed the main process alone gets SIGTERM, the upload finishes, and the script re-raises at the boundary.
10A nightly scan runs as a Type=oneshot service with RuntimeMaxSec=30min. One morning systemctl status shows it still "activating" after six hours, and nobody was paged. What is wrong?
Correct — For oneshot units the limit is TimeoutStartSec, which defaults to infinity; the lesson's 2-second test failed the unit with result 'timeout'.
Incorrect — RuntimeMaxSec is wall-clock time; the problem is that it has no effect on this unit type.
Incorrect — Nothing fired: RuntimeMaxSec has no effect on a oneshot unit, so the scan was never signalled.
Incorrect — Time limits belong to the service unit; the timer merely starts it.
11A root job runs write-secure.sh /srv/drop/reports/today.txt. reports is a symlink that the uploads team can repoint, and today it points at a root-owned directory with mode 0755. The script stops with write-secure: refusing /srv/drop/reports: it is a symlink; name the real directory. A teammate calls that overcautious, since the target passes the owner and mode checks. Who is right?
Incorrect — Today's target proves nothing about the next run; whoever can repoint the link picks where a root-owned write lands.
Incorrect — -T governs the last name alone; a symlinked directory earlier in the path is still followed.
Correct — Owner and mode checks can vouch for a real directory; a link on the way hands the choice to whoever controls the link.
Incorrect — mktemp follows a symlinked directory like any path lookup; the refusal is the script's own rule.
12A job writes a report by creating $tmp with mktemp in spool/ and running mv -f -- "$tmp" spool/report. After one run, spool/report is still an old symlink to /home/app/cache, and a .tmp.XXXXXX file has appeared inside /home/app/cache. What happened?
Incorrect — mktemp created the file in spool/ under a new random name; the link played no part until the move.
Correct — Without -T mv treats a directory destination, even one reached through a link, as a place to move into; --no-target-directory replaces the name.
Incorrect — The temp file and the destination are in the same directory, so this was a rename; the problem is how mv chose its target.
Incorrect — -f suppresses prompts; without it mv would still move the file into the directory the link names.
13A wrapper runs ssh svc@db01 "stat -- ${path@Q}". The account svc has /bin/sh (dash) as its login shell, and one path contains a newline. What does the remote side receive?
Incorrect — ${path@Q} wraps the value, so it stays one word; dash misreads the quoting, which corrupts the value rather than splitting it.
Incorrect — That message is dash's printf lacking %q; ${path@Q} is expanded by the local Bash before ssh sends anything.
Incorrect — ssh joins the words into one string for the remote login shell, which parses it again; that shell is dash here.
Correct — It is corruption, not injection; for a POSIX sh, quote with single quotes and '\'', as shlex.quote does.
14A script started with sudo runs exec {key}< /etc/app/signing.key, does its root step, then exec setpriv --reuid="$user" --regid="$gid" --clear-groups --no-new-privs --inh-caps=-all --reset-env -- "$0" --worker. The review says the worker can still read the key. Is that so?
Correct — Changing uid does not close descriptors, and whatever the root part opened stays readable after the drop.
Incorrect — --reset-env clears environment variables and --clear-groups supplementary groups; neither touches open file descriptors.
Incorrect — Permissions are checked when a file is opened; an open descriptor keeps the access it was opened with.
Incorrect — The risk is the inherited descriptor itself; the environment was reset, and the key's content was never in it.
15A Bats test starts ./guarded.sh &, runs its assertions, then kills the job. When an assertion before the kill fails, the CI job hangs until the runner's time limit instead of reporting not ok. What is the cause and the fix?
Incorrect — Bats does not retry tests; a failing assertion ends the test at once.
Incorrect — Bats does not wait for signals, and kill 0 would signal the process group Bats itself runs in.
Correct — It is the inherited-descriptor bug inside the harness, and teardown runs after a failed test too.
Incorrect — The hang happens when the test failed before its kill and wait; the harness waits on a descriptor, not on the test's wait.
16To make healthcheck.sh testable, a teammate proposes CURL=${CURL:-curl} in the script and calling "$CURL", so tests can set CURL=./fakebin/curl. Why does the lesson fake curl with a function, or a fakebin directory put first on PATH for the test command, instead?
Incorrect — run passes the environment through; the lesson sets FAKE_CODE that way.
Correct — A test function exists in the test shell alone, so the fake adds no switch that production could flip.
Incorrect — A variable can hold a path, and "$CURL" runs it; the problem is who else can set it.
Incorrect — ShellCheck accepts "$CURL" args; no gate fails on it.
17An earlier summarize.py anchored its pattern at the end only, Failed password for (?:invalid user )?(?P<user>.*) from (?P<addr>\S+) port \d+ ssh2$, used with search. It meets ... sshd-session[2490]: Failed none for invalid user Failed password for root from 6.6.6.6 port 1 from 203.0.113.45 port 51400 ssh2. What happens, and what fixes it?
Correct — The username supplied the words and an end-only anchor let search start inside it; anchored at both ends, the lab counted nothing for that line.
Incorrect — The greedy user group runs to the last from ... port ... ssh2, so the address is sshd's 203.0.113.45; the bug is that the line counts at all.
Incorrect — search looks anywhere in the line, and the username supplies Failed password for; that is the hole.
Incorrect — \S+ cannot contain spaces, and the captured address is a valid IP; nothing raises.
18A developer adds print(f"warning: {skipped} lines skipped") to the Python core that logwatch.sh calls. That night the wrapper exits 2 with logwatch: cannot read the report, although the log was fine. Why?
Incorrect — print does not change the exit status; the core still returned 0, and the failure came later, in jq.
Incorrect — jq has no such rule; it failed because the captured text was no longer one JSON document.
Incorrect — A here-string passes the whole value, newlines included; jq saw the warning first and stopped there.
Correct — The contract is data on stdout and messages on stderr; one stray line breaks the parser, and the wrapper correctly refuses to guess.
18 questions · explanations appear as you answer
Python automation engineering
16 questions
01blsync.toml sets max_remove_fraction = 0.2, the unit exports BLSYNC_MAX_REMOVE_FRACTION=0.9, and cron runs blsync --max-remove-fraction 0.5 plan. The feed would remove 3 of the 10 current entries. What happens?
Incorrect — The file is the second layer, not the last; the environment and then the flag override it.
Incorrect — The loader does not pick the strictest value; it merges in a fixed order, and the flag's 0.5 is used.
Correct — Layers merge as defaults, file, environment, flags, the later one winning, and the merged value is what gets checked.
Incorrect — Disagreement between layers is normal: later layers override earlier ones, and 78 is for unusable values.
02blsync plan works from a shell in the project directory. As a system service it fails at once, and the journal shows blsync: config: blsync.toml: [Errno 2] No such file or directory: 'blsync.toml' and status=78/CONFIG. What is the fix?
Correct — The default config and plan.json are relative to the working directory, which is / for a system service unless the unit sets one.
Incorrect — A service can read the files its user can reach; the problem is where the relative path is resolved.
Incorrect — The error is ENOENT, not EACCES: the file was looked up in the wrong directory, not refused.
Incorrect — 78 means the configuration is unusable; a retry repeats the same lookup in the same wrong directory.
03The plan job uploads plan.json as a CI artifact, and the apply job runs blsync apply on it without --plan-sha256. Someone with write access to the artifact store adds one valid address that stays inside every limit. Which control would have refused the edit?
Incorrect — The state did not change; the plan's state_sha256 still matches, so this check passes.
Correct — Validation catches malformed or out-of-limit plans; a well-formed edit is caught by the digest of what was reviewed.
Incorrect — One more entry stays inside the limits, so the second run of check_limits passes it.
Incorrect — Plan.load checks shape and address syntax, not origin; a valid address passes.
04In ticketer's run(), the except* PermanentError clause sets status = 1 and the except* TransientError clause sets status = status or 75. post_all raises an ExceptionGroup holding one DeadlineExceeded and one PermanentError. What does run() return?
Incorrect — The second clause runs, but status or 75 keeps the 1 that the first clause set.
Incorrect — DeadlineExceeded subclasses TransientError, so the second clause matches it.
Incorrect — Several except* clauses can run for one group; the transient part is logged by its own clause.
Correct — Each except* clause receives its matching part of the group, and the failure that needs a human outranks the transient one.
05ticketer --budget 60 hung for 40 minutes on one request. A packet capture shows the API's proxy sending the status line and headers one byte every half second. Why did neither the per-call timeout nor the budget stop it?
Incorrect — timeout() returns min(cap, remaining), the smaller value; the size of the timeout was not the problem.
Correct — Each byte reset the silence timer and the body reader never started, so a limit outside the process had to end it.
Incorrect — read_body runs after urlopen returns, and urlopen reads the status line and headers first.
Incorrect — The Deadline class uses the monotonic clock, not SIGALRM; no alarm is involved.
06ticketer derives each idempotency key from the finding's id. A run posted two findings just as the ticket API failed: the tickets were created, but the answers were lost. The API then stayed down for three days, and the next run opened a second ticket for both. The API documents a 24-hour key window. What went wrong?
Correct — A stable key replays within the server's memory of it; after a long outage, a server-side duplicate check on the finding id has to catch it.
Incorrect — A date makes each day's key new, which creates duplicates even inside the window.
Incorrect — The header goes with each attempt of each run; the key was the same, and the server no longer knew it.
Incorrect — The documented window was 24 hours; nothing had to clear it, three days simply outlasted it.
07A log-counting service moves from Ubuntu's python3 to a free-threaded 3.14t build for speed. Four threads update a shared stats.hits += 1. On the old build the totals always matched; now they come out low. What is the diagnosis?
Incorrect — The free-threaded build keeps built-in objects internally consistent; what breaks is a read-modify-write spread over several operations.
Incorrect — Threads in the free-threaded build share one interpreter and its objects; that sharing is why the updates collide.
Correct — The lesson lost updates on the GIL build too once a call sat between the read and the write; a lock fixes both builds.
Incorrect — The threads were joined before the count was read; slower single-thread speed changes timing, not the final total.
08A CPU-bound scorer ran 3 times faster with threads on a free-threaded build in testing. In production, with a vendor C extension imported at start-up, threads give no speed-up, and a warning about the GIL appears once in the log. What is the likely cause?
Incorrect — The default build refuses PYTHON_GIL=0 with a fatal error; it does not ignore it silently.
Incorrect — The free-threaded build runs threads in parallel on several cores without an affinity call.
Incorrect — No such timer exists; the GIL comes back when PYTHON_GIL=1, -X gil=1 or an unmarked extension forces it.
Correct — Py_GIL_DISABLED says how the interpreter was built; whether the GIL is on now is a run-time fact.
09On Ubuntu's default (GIL) build, which of these jobs gets faster when its work is spread over a ThreadPoolExecutor on a 4-CPU host?
Incorrect — Pure-Python bytecode takes turns on the GIL; the lesson's threads matched serial time.
Correct — Large hash updates and file reads release the GIL, so the threads overlap on several cores.
Incorrect — Calls on buffers under 2048 bytes keep the GIL, and the loop around them is Python bytecode.
Incorrect — Contended increments do no parallel work and need a lock for a correct total.
10An asyncio job wraps a blocking urlopen call in asyncio.to_thread and starts 10 of them in a TaskGroup. Each request takes 0.2 s, the host has 4 CPUs, and the run takes about 0.42 s rather than 0.2 s. Why?
Correct — to_thread hands work to the loop's default thread pool; to go wider, set a larger default executor or use an async client.
Incorrect — A TaskGroup starts each task at once; the limit is the thread pool the calls are sent to.
Incorrect — Blocking network waits release the GIL; the rounds come from the size of the pool.
Incorrect — Thread start-up takes microseconds; two rounds of 0.2 s account for the time.
11An early version of run_probe killed the probe's process group with os.killpg(proc.pid, signal.SIGKILL) in its finally block and then called await proc.wait(). A probe stopped by the output cap hung the whole run. Why does the lesson call await proc.communicate() there instead?
Incorrect — The kernel reports a SIGKILL death like any other exit; wait was blocked on the pipe, not on the status.
Incorrect — A late signal would still end the probe, and wait would then return; the stall was on the pipe.
Correct — The pipe the capped reader stopped draining never closed on its own; communicate reads the rest and waits for the process.
Incorrect — communicate sends no signals; SIGKILL to the group already ended each member.
12ssh_fleet.py connects with timeout=5, banner_timeout=5, auth_timeout=5 and passes timeout=DEADLINE + 5 to exec_command, yet it still wraps the command as timeout 10 sh -c ... on the host. What does the host-side timeout add?
Incorrect — The paramiko timeouts apply to each connection; they cover different stalls than the remote timeout does.
Incorrect — timeout does not change how the login shell starts; it limits how long the command runs.
Incorrect — timeout exits 124 when it stops a command, and paramiko's timeouts can be caught; that is not the reason.
Correct — Connect, banner, auth and read timeouts end a stall; a limit on duration has to sit next to the command.
13A nightly fleet run prints 10.0.4.21:22: ERROR host key mismatch (ssh-ed25519) for one host and exits 1. The host was rebuilt yesterday. Which response keeps the fleet's host-key policy intact?
Incorrect — That trusts whatever answers on the network right now, which is trust on first use done by hand.
Correct — Rotation goes through the inventory, not through first contact; until then the unknown key keeps failing closed.
Incorrect — WarningPolicy logs and connects, handing the session to whoever presented the key.
Incorrect — RejectPolicy refuses unknown keys and records nothing, so the run keeps failing, correctly.
14A threaded ingest service must run regular expressions from a rule file it cannot rewrite, so it wraps each match in signal.signal(signal.SIGALRM, handler) and signal.alarm(1). At start-up the worker threads fail with ValueError: signal only works in main thread of the main interpreter. What should it use?
Correct — Both bound the work per match without signals; compare their answers with re on tricky inputs before swapping.
Incorrect — setitimer has finer steps, but the handler still has to be installed from the main thread.
Incorrect — A cap bounds input length, not work: the lesson's exponential pattern took over a second on 25 characters.
Incorrect — re.ASCII narrows \w and its relatives; it does nothing about backtracking.
15A record loader runs raw = json.loads(text) inside try, with except json.JSONDecodeError: reject(text), and then reads raw["target"]. Which input crashes the loader instead of being rejected?
Incorrect — A truncated object raises JSONDecodeError, which the handler catches.
Incorrect — That parses fine and reaches the grammar check, which is the next step's job.
Incorrect — json.loads("") raises JSONDecodeError, so the handler rejects it.
Correct — Deep nesting escapes too, with RecursionError, and [1] parses but is not an object; catch (ValueError, RecursionError) and check the shape.
16A JSON log formatter runs payload.update(record.fields) after it has set ts, level and msg. A record parsed from an upload carries {"user": "bob", "level": "debug", "msg": "login ok"}. What does the log line show?
Incorrect — json.dumps escapes characters, but a merged key has already replaced the real one before the dump.
Incorrect — A dict has one value per key and json.dumps writes one object; no second line appears.
Correct — json.dumps stops line forging, but merging lets input overwrite the reserved keys inside a valid line.
Incorrect — dict.update overwrites existing keys without complaint.
16 questions · explanations appear as you answer
Operate and ship
9 questions
01sweep runs hourly from cron through sweep-cron.sh, and nobody has been mailed for a week. The metric file shows sweep_last_run_seconds from 20 minutes ago, sweep_last_success_seconds from 8 days ago and sweep_last_exit_code 75. What is going on?
Correct — 75 means "a later run can finish", and a week of it means it will not; the freshness gate turns that into a page.
Incorrect — A job that stopped would show an old last-run time; this one ran 20 minutes ago.
Incorrect — A failed run moves last-run and keeps last-success; the gap is the designed signal.
Incorrect — node_exporter keeps no such cache; an unreadable file makes the series disappear.
02In sweep, the first target is under a change freeze (a PermanentError, recorded in the report), and the second target's state file is truncated JSON. What exit status does the run end with?
Incorrect — The crash comes before outcome() runs; except Exception sets 70, and the report's contents are not ranked.
Correct — An unclassified exception means a state nobody planned for, and 70 outranks 1 and 75.
Incorrect — A JSONDecodeError is not classified as transient; no rule covers it, so it is treated as a bug.
Incorrect — 78 is for an unusable configuration before anything was attempted; this file is target state.
03After an incident, check_audit.py audit/sweep.jsonl prints chain over 9 records: intact, but the collector that receives each line as it is written holds 13. What does that show?
Incorrect — Each event is its own line; start and done are two records in both copies.
Incorrect — An edited line with a successor breaks the chain; the local file verified because nothing followed the missing lines.
Incorrect — Concurrent appends without the lock produce a broken chain, not a shorter intact one.
Correct — Removing lines from the end leaves a shorter valid chain; the off-host copy is the evidence.
04Servers get certcheck with uv tool install certcheck-0.1.0-py3-none-any.whl, and the project commits a uv.lock that pins cryptography==50.0.1. After a new cryptography release, a server installed that week behaves differently from staging. Why?
Incorrect — A py3-none-any wheel carries the project's modules alone; dependencies are installed next to it.
Incorrect — uv resolves one lock for each platform the project supports; the server skipped the lock entirely.
Correct — A tool install matches the lock while nothing newer exists; install servers from the hashed lock.
Incorrect — The install used the wheel you named and cryptography's published wheels; nothing was compiled.
05A single-file certcheck.pyz, built with python3 -m zipapp from a directory holding certcheck and its hash-locked dependencies, fails on start with ImportError: cannot import name 'x509' from 'cryptography.hazmat.bindings._rust' (unknown location). What is the cause, and what is a sound choice?
Incorrect — The same archive fails on its own build host; the lesson built and ran it on one machine.
Correct — Zip imports cover source and bytecode; a tool with compiled dependencies needs an unpacked install, pex's cache or an image.
Incorrect — zipapp does not strip files; the extension is in the archive and cannot be loaded from there.
Incorrect — A custom __main__.py fixes the exit status, not imports; this is about loading compiled code.
06uv export --format pylock.toml -o pylock.toml followed by pip install -r pylock.toml stops with The editable requirement file:///srv/certcheck/ (from pylock.toml) cannot be installed when requiring hashes. What is the fix?
Correct — A hashed lock puts pip in hash-checking mode, and a directory has no single file to hash.
Incorrect — --no-deps stops dependency resolution; it does not turn off hash-checking mode.
Incorrect — That removes the verification the lock exists for, to work around a problem with one entry.
Incorrect — pip still refuses the editable entry in hash mode; the project has to stay out of the hashed lock.
07certcheck shipped with 100% branch coverage, yet a bug slipped through: expiring certificates no longer set exit status 1. The test for that path calls main() on an expiring certificate and asserts nothing. What does this show?
Incorrect — Coverage was already 100%; the path ran, so no floor would have noticed.
Incorrect — Branch coverage adds both sides of each branch to line coverage; the lines ran and were counted.
Incorrect — There is no such rule; cli.py's lines were measured, as the report showed.
Correct — A call without an assertion covers the same lines as a real test; the floor finds untested code, reviews find unchecked code.
08A pull request changes one workflow line to uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v7.0.1, a different full SHA with the same comment, and actionlint passes. Why can this still be a problem, and what catches it?
Incorrect — A SHA is immutable, but GitHub serves commits from any fork under the upstream name, and this one is v6.0.3's.
Incorrect — actionlint checks syntax and expressions; it does not resolve tags or check pins.
Correct — check-pins.sh asks upstream which commit the tag names and reports MISMATCH; a format check cannot.
Incorrect — Update bots propose new pins; they do not audit whether an existing SHA matches its comment.
09Deploys run cosign verify --key cosign.pub "$ref" on a digest reference. A test build signed months ago with the same release key is deployed to production by mistake, and verification passes. What would have refused it?
Incorrect — Tags can be moved to any digest; switching to them weakens the check instead of narrowing it.
Correct — A key signature says the key signed that digest, so each image it ever signed passes; the annotation narrows that to releases.
Incorrect — The log records when a signature was made; it does not rank test signatures below release ones.
Incorrect — An attestation says what is inside; it does not say the build was meant for production.
9 questions · explanations appear as you answer
Capstone
4 questions
01A reviewer wants certrun_last_success_seconds to move on exit 0 alone, as sweep's did in "Observable jobs". admin has waited for a person for three weeks, so each apply exits 1, and the alert is time() - certrun_last_success_seconds > 2 days. What would the change do?
Incorrect — certrun.sh writes the run metrics on each apply that gets the lock, 77 and 143 included, and last-success already moves on exit 1.
Incorrect — Exit 1 means every endpoint was handled and some wait for a person; certrun_needs_attention already pages for those.
Correct — With 1 as a normal state, last-success means "a run handled every endpoint"; its age is the stopped-timer alert, and needs_attention pages for admin.
Incorrect — node_exporter serves whatever value the file holds; a stale value is exactly what the age alert measures.
02A timer run of scr-capstone-certrun@lab.service ends 40 seconds into its 120-second budget. The journal shows the core was killed with SIGKILL before the deadline (out of memory?), then status=70/SOFTWARE, and the unit is not restarted. An operator wants that case mapped to 75 like a timeout. What is the right answer?
Correct — certrun.sh calls 137 a deadline when the deadline has passed; before that it is a failure a person must see, and 70 is in RestartPreventExitStatus.
Incorrect — timeout's SIGKILL comes after the budget runs out; 40 seconds into 120 it could not have been timeout.
Incorrect — SuccessExitStatus lists SIGTERM alone; a SIGKILL was not a requested stop, and 143 would misreport it.
Incorrect — RestartPreventExitStatus=1 2 70 77 78 prevents that; 70 waits for a person, by design.
03After a run rolled back shop (the new certificate came from an untrusted issuer), each later apply prints a deploy was rolled back (...); fix the cause, then delete pending/shop.json, and the CA's counters do not move. A colleague proposes a daily cron job that deletes rolled-back records so rotations retry by themselves. What does that lose?
Incorrect — The record holds the new private key and the exact request bytes, and it blocks further CA requests.
Incorrect — The flock on DESTINATION.lock prevents overlapping runs; the record is about the rotation, not concurrency.
Incorrect — The rollback renamed the backup into place on the host; the record in state/ is a separate file.
Correct — The record is marked rolled back so no later run asks the CA again; a person confirms the cause is gone, then deletes it, as the Try this does.
04To let the job write a debug log next to its code, a teammate changes /opt/scr-capstone to certrun:certrun, mode 0755. release/, certrun.sh, verify_core.py and etc/cosign.pub stay root-owned. Does the verify gate still stop a core that nobody signed?
Incorrect — Owning the files is not enough; the account now owns the directory that holds them.
Incorrect — The account can put its own key, its own gate and its own release in place; the check then passes on its terms.
Correct — Renaming needs write access to the parent directory, as the lab's refused mv showed; root's all the way up keeps the gate out of the account's reach.
Incorrect — A full disk is a separate problem; the ownership change hands the account control over what runs.
4 questions · explanations appear as you answer