Signals, timeouts and shutdown under systemd and Kubernetes
Re-raise vs exit, PID 1, forwarding signals, deadlines and drain budgets.
tar -xzf scr-bash-signals.tar.gz, which creates scr-bash-signals/. SHA-256: b3f3d94ff9b0b38a05c34c67e0db8f338ae9e5a3efd35f28dfbfb502e746a71fThis lesson makes a Bash job stop the way its supervisor expects. You will see how systemd judges a stopped service by its exit status, give a batch job a drain budget so the item in flight finishes, find out why a script running as PID 1 in a container can ignore SIGTERM entirely, and put a deadline on a whole run without leaving a watchdog process behind. Each behaviour is executed here: systemd 259 transient units on Ubuntu 26.04, and a PID namespace created with unshare standing in for a container.
Refresher: "Traps, signals and cleanup" in bash-ops covers the EXIT trap, handlers that must exit, re-raising a signal with trap - TERM; kill -s TERM "$$", and the 128+n statuses (130 for SIGINT, 143 for SIGTERM, 137 for SIGKILL). "Temp files, locks and timeouts" covers timeout, its 124 and -k, and "Signals and job control" in Linux essentials covers kill, TERM before KILL and the escalation between them. Here the question is what systemd and Kubernetes do with those statuses and signals. Unpack the lesson files (the box at the top of this page) in your home directory and work in ~/scr-bash-signals as your normal user; sudo is needed only to create system units and PID namespaces.
Exit 143 is not a clean stop to systemd
A worker with one cleanup path, and two ways to answer SIGTERM: exit with 143, or reset the trap and re-raise the signal so the process really dies of SIGTERM.
#!/usr/bin/env bash# A long-running worker with one cleanup path. Usage: worker.sh exit143|reraiseset -euo pipefailstyle=${1:?usage: worker.sh exit143|reraise}work=$(mktemp -d -t worker.XXXXXX)child=cleanup() {trap '' TERM INT # a signal arriving now must not cut the cleanup short[[ -z $child ]] || kill -TERM "$child" 2>/dev/null || truerm -rf -- "$work"echo "worker: cleaned up"}trap cleanup EXITcase $style inexit143) trap 'exit 143' TERM ;; # reports "exited with 143"reraise) trap 'trap - TERM; kill -s TERM "$$"' TERM ;; # dies of SIGTERM, as askedesacecho "worker: started ($style)"sleep 300 &child=$!wait "$child"
systemd-run starts a command as a transient service unit, which behaves like a unit file you wrote under /etc/systemd/system (--uid sets User=, -p sets any other property, --same-dir keeps the working directory). KillMode=mixed makes systemd send SIGTERM to the worker only, and the worker stops its own child; the next section shows what the default does instead. Each step starts the worker, stops it the way systemctl stop does in production, and prints the journal of that run only (-I means the current invocation of the unit). Reading a system unit's journal without sudo needs membership in the adm or systemd-journal group; otherwise put sudo in front of journalctl.
Both workers cleaned up. The first one is recorded as a failure: "Main process exited, code=exited, status=143" and "Failed with result 'exit-code'". A failed unit shows up in systemctl --failed, fires any OnFailure= handler and pages whoever watches those. The re-raised one ends with "Deactivated successfully". The rule is in systemd.service(5) under SuccessExitStatus=: a clean stop is exit status 0 or, except for Type=oneshot, death by SIGHUP, SIGINT, SIGTERM or SIGPIPE. Exit status 143 is just a number to systemd, the same as 1.
If you run a program you cannot change that exits 143 on SIGTERM, declare that status as success in the unit instead:
A failed transient unit stays loaded, and systemd-run refuses to start another unit with the same name. To repeat any step of this lesson, first run sudo systemctl reset-failed 'scr-bash-signals-*'.
The first line of cleanup ignores further TERM and INT signals. Without it, a signal that arrives while the EXIT trap is running makes the TERM handler run exit in the middle of the cleanup, and the rest of it never happens. This job's cleanup takes two seconds, and a SIGTERM arrives after it has started:
#!/usr/bin/env bash# A job whose EXIT cleanup takes 2 s, with a TERM handler that exits.# Usage: slow-cleanup.sh [protect] (protect: ignore TERM while cleaning up)mode=${1:-}trap 'exit 143' TERMcleanup() {[[ $mode == protect ]] && trap '' TERMecho "cleanup: removing work files"sleep 2echo "cleanup: done"}trap cleanup EXITecho "job: working"sleep 1echo "job: finished, exiting 0"
The job had finished and was exiting 0. The late SIGTERM ran exit 143 inside the cleanup: "cleanup: done" never printed and the status became 143. With the signal ignored during cleanup, the cleanup completes and the real status survives. The same thing happens outside a lab. While this lab was being built, the worker units ran with the default KillMode; systemd's SIGTERM sometimes reached the worker's sleep before the worker itself, and in one run two of the three workers had their cleanup cut short and left their temporary directories in /tmp.
Keep one thing in mind when you read 143 in a log: it means "ended by SIGTERM", not "shut down cleanly". A process with no handler at all also ends with 143, without running any cleanup. The next section shows one.
Draining work in flight: KillMode and TimeoutStopSec
A batch job that uploads items one at a time should finish the upload in flight when it is asked to stop, then exit. This version uses the rule "Traps, signals and cleanup" showed as a pitfall: when a signal with a trap arrives while a foreground command runs, Bash runs the trap only after that command returns. Here that is exactly the drain behaviour we want.
#!/usr/bin/env bash# Stand-in for one upload: takes 3 seconds and has no signal handling of its own.echo "upload $1: start"sleep 3echo "upload $1: done"
#!/usr/bin/env bash# Uploads queued items one at a time. On SIGTERM it lets the upload in flight# finish, then stops and re-raises the signal.set -euo pipefailstop_requested=0trap 'stop_requested=1' TERMfor item in 1 2 3 4 5; do./upload.sh "$item" # foreground: Bash runs the TERM trap only after this returnsif (( stop_requested )); thenecho "batch: stop requested, stopping after item $item"trap - TERMkill -s TERM "$$"fidoneecho "batch: queue empty"
systemd decides who receives the stop signal through KillMode=. The default, control-group, sends SIGTERM to every process in the unit's control group at once. mixed sends SIGTERM only to the main process and sends SIGKILL to whatever is left once the main process has exited or TimeoutStopSec= has run out. Here is the same job under each mode, stopped one second into the first upload:
With control-group the upload got SIGTERM directly and died mid-item: "upload 1: start" never has a "done", and Bash reports the child as Terminated. upload.sh had no handler, so its status was 143, and set -e in batch.sh exited with that status before the stop check ran. The unit failed with 143, and no cleanup code in upload.sh ran. With mixed, only batch.sh got the signal; the upload finished, the job stopped at the item boundary and re-raised, and systemd recorded a clean stop.
TimeoutStopSec= is the drain budget: how long systemd waits after SIGTERM before it sends SIGKILL (the default comes from DefaultTimeoutStopSec=, 1 min 30 s on this machine). Set it shorter than the work in flight and the drain is cut off:
One second was not enough for a three-second upload, so systemd killed every process in the unit with SIGKILL and marked it failed with result 'timeout'. Size the budget from the longest item plus cleanup, measured, and make items short or resumable when that number gets large.
PID 1: why a container can ignore SIGTERM
From the Kubernetes documentation (not executed here): when a pod is deleted, the kubelet runs any preStop hook, then the container runtime sends the stop signal (SIGTERM unless the image sets STOPSIGNAL) to process 1 in each container, and after terminationGracePeriodSeconds (30 by default, counted from the start of termination) it sends SIGKILL. What happens to that SIGTERM is decided by a kernel rule described in pid_namespaces(7): the first process in a PID namespace receives a signal from outside only if it has installed a handler for that signal. SIGKILL and SIGSTOP are the exceptions. A default action is not a handler, so a process that relies on the default "terminate" for SIGTERM ignores it as PID 1.
To execute that rule without a container runtime, mini-stop.sh does what a runtime's stop does: it starts the entrypoint as PID 1 of a new PID namespace with unshare, sends SIGTERM from outside, waits a grace period and then sends SIGKILL.
#!/usr/bin/env bash# A stand-in for a container runtime's stop: run the entrypoint as PID 1 of a new PID# namespace, send it SIGTERM from outside, wait a grace period, then SIGKILL.# Usage (as root): mini-stop.sh GRACE_SECONDS ENTRYPOINT [ARGS...]set -euo pipefailgrace=${1:?usage: mini-stop.sh GRACE_SECONDS ENTRYPOINT [ARGS...]}shiftunshare --pid --fork --mount-proc --kill-child -- "$@" &runner=$!sleep 1pid1=$(pgrep -P "$runner") # PID 1 inside the namespace, an ordinary PID out hereecho "runtime: sending SIGTERM to the entrypoint"kill -TERM "$pid1"for (( i = 0; i < grace * 10; i++ )); dokill -0 "$pid1" 2>/dev/null || breaksleep 0.1doneif kill -0 "$pid1" 2>/dev/null; thenecho "runtime: still running after ${grace}s, sending SIGKILL"kill -KILL "$pid1"firc=0wait "$runner" 2>/dev/null || rc=$? # 2>/dev/null: drop Bash's "Killed" job noticeecho "runtime: container exit status $rc"
#!/usr/bin/env bash# Container entrypoint without signal handling: runs the worker in the foreground.echo "entrypoint: PID $$ in its namespace"sleep 300
#!/usr/bin/env bash# Container entrypoint that replaces itself with the worker.echo "entrypoint: PID $$ in its namespace, exec'ing the worker"exec sleep 300
Both entrypoints report PID 1, and both sit through the grace period and are killed: 137, the status Kubernetes shows when a pod took the whole grace period to die. entry-plain.sh has no traps at all, so Bash relies on the default action and the kernel dropped the signal. exec made sleep PID 1, and sleep has no handler either. exec removes the shell in between, but it helps only when the program you exec installs its own SIGTERM handler.
Most real entrypoints are not that plain: they have trap cleanup EXIT. That one line changes the result, because Bash then installs its own SIGTERM handler so that it can run the EXIT trap when it is killed:
#!/usr/bin/env bash# The plain entrypoint plus an EXIT trap, as most real entrypoints have.trap 'echo "entrypoint: EXIT trap ran"' EXITecho "entrypoint: PID $$ in its namespace"sleep 300
This time PID 1 had a handler, so the signal arrived: Bash ran the EXIT trap at once and exited with 143, without waiting for its foreground sleep. That is not a drain. When PID 1 of a namespace exits, the kernel sends SIGKILL to every other process in it, so the worker was killed mid-task. An EXIT trap makes the container stop fast; it does not make it stop gracefully.
Two fixes work. The first is an entrypoint that handles SIGTERM itself and forwards it to its worker:
#!/usr/bin/env bash# Container entrypoint that forwards SIGTERM to its worker and exits with the worker's status.child=# shellcheck disable=SC2329 # called by the TERM trap; ShellCheck 0.11 misses thatforward() {echo "entrypoint: SIGTERM received, forwarding to worker $child"kill -TERM "$child" 2>/dev/null}trap forward TERMecho "entrypoint: PID $$ in its namespace"sleep 300 &child=$!status=0while kill -0 "$child" 2>/dev/null; dostatus=0wait "$child" || status=$? # a trapped signal ends wait early: loop until the worker is gonedoneecho "entrypoint: worker exited with $status"exit "$status" # PID 1 cannot re-raise: the kernel drops the unhandled signal
The second is a small init process as PID 1 whose job is to forward signals and reap children. tini is in the Ubuntu archive but not on a default install (docker run --init injects the same kind of init). Install it, then let it run the unchanged plain entrypoint as its child:
Both stop within a second with 143, and the forwarding entrypoint lets its worker finish its own shutdown first. With tini the entrypoint is PID 2, an ordinary process for which SIGTERM's default action applies. The forwarding entrypoint ends with exit "$status", not a re-raise, for the same reason the plain one failed: kill -s TERM "$$" from PID 1 to itself is dropped by the kernel. Kubernetes has no --init option; the usual choices are an init such as tini as the image's entrypoint, or an entrypoint that handles and forwards SIGTERM like the one above.
Deadlines for the whole run, without orphaned watchdogs
A common way to cap a script's run time is a watchdog subshell that signals the script after a delay. It is unsafe, and here is why:
#!/usr/bin/env bash# UNSAFE deadline pattern, shown to be avoided: a watchdog subshell that signals the script.set -euo pipefail( sleep 600 && kill -TERM "$$" ) &watchdog=$!trap 'kill "$watchdog" 2>/dev/null || true' EXITecho "job: working"sleep 1echo "job: done"
The job finished and its EXIT trap killed the watchdog subshell, but the subshell's sleep 600 survived as an orphan (the second command removes it), as "Processes, pipes and file descriptors in real scripts" showed for any child of a killed process. Its EXIT trap also silently replaces any EXIT trap the script set earlier, because Bash keeps one action per trap.
Put the deadline outside the job instead: timeout in the caller or in the unit's ExecStart=, or a systemd time limit (below). Without --foreground, timeout signals the command's whole process group, so children go too:
#!/usr/bin/env bash# A slow job with two parallel children.sleep 30 &sleep 30 &wait
124 and no sleep left: the deadline reached both children. The rest of timeout's behaviour is in "Temp files, locks and timeouts": -k for a command that ignores SIGTERM (137), and the uutils timeout on Ubuntu 26.04 returning 124 where GNU's returns 137 with -s KILL, so a wrapper that branches on these codes should prefer -k.
In systemd, pick the limit that matches the unit type. RuntimeMaxSec= limits a long-running service, but a timer-driven batch job is usually Type=oneshot, and for those systemd.service(5) says RuntimeMaxSec= has no effect: a oneshot unit is still "starting" until its command exits. Its limit is TimeoutStartSec=, which for oneshot units defaults to infinity. The same four-second observation of slow-job.sh (30 seconds of work) under each:
With RuntimeMaxSec=2 the job was still activating after four seconds, and it would have run for all thirty (the step stopped it by hand). With TimeoutStartSec=2 systemd terminated it after two seconds and recorded the failure with result 'timeout'. The last command clears the failed units this lesson created.
-k) ends the process without running any handler. Keep secrets out of files where you can ("Secrets: where they come from and where they leak" in py-sec ranks the sources): pass them on stdin, read them from $CREDENTIALS_DIRECTORY (systemd LoadCredential=), or, for a service, use RuntimeDirectory=, a private directory under /run (in $RUNTIME_DIRECTORY) that systemd removes when the unit stops. $XDG_RUNTIME_DIR exists only in a user's login session; a system unit with User= and a cron job do not have it, so "$XDG_RUNTIME_DIR/secret" becomes /secret. Do not treat /dev/shm as a secret store: it is a shared, world-writable directory like /tmp.Try this
Run the batch job as a container entrypoint: sudo ./mini-stop.sh 5 ./batch.sh. Before you run it, predict the result from what you saw under systemd and at PID 1: does the re-raise still give a clean stop? Then make the job stop cleanly as PID 1 by changing one line of a copy, batch-pid1.sh, and run the same command with it. Finally confirm with pgrep -a -u "$USER" -x sleep that nothing is left, and remove tini if you do not want to keep it: sudo apt-get purge -y tini.
Expected: the re-raise kill -s TERM "$$" is now sent by PID 1 to itself, so the kernel drops it. The job keeps uploading ("upload 2: start", "upload 2: done", "upload 3: start") until the runtime sends SIGKILL and reports 137. Replacing that line with exit 143 makes the copy stop after "upload 1: done" with status 143. pgrep prints nothing and exits 1.
Takeaway
Decide how your script will be stopped before you write its handler: re-raise under systemd (or declare SuccessExitStatus=), use KillMode=mixed with a measured TimeoutStopSec when the main process must drain its children, forward and exit explicitly when you are PID 1, and put run-time deadlines outside the job.
systemctl stop scan-worker, the journal shows "Main process exited, code=exited, status=143/n/a" and "Failed with result 'exit-code'". The script's handler is trap 'cleanup; exit 143' TERM, and cleanup did run. What is the correct change?./server in the foreground; server installs its own SIGTERM handler and drains in 2 seconds. Which change makes the pod stop in about 2 seconds?./server returns, and server never got the signal.KillMode=mixed and TimeoutStopSec=10. Its main script processes items that take up to 25 seconds and stops at the next item boundary on SIGTERM. What happens when it is stopped mid-item?