HEALTHCHECK and graceful shutdown in containers

Tell the orchestrator when your container is actually healthy, and handle SIGTERM so deploys don't drop requests.

Aug 27, 2024·Updated ·6 min readIntermediate·By SecOpsLog · command-tested

HEALTHCHECK options, defaults, and what each one counts

OptionDefaultMeaning
--interval30stime between checks, measured from the end of the previous one
--timeout30sa probe running longer than this fails, and the probe process is killed with SIGKILL
--retries3consecutive failures before the container is unhealthy; one success resets to healthy
--start-period0sgrace for slow starts: failures are not counted until the first success or the period ends
--start-interval5scheck frequency during the start period (Engine 25.0 or later)
CMD exit coden/a0 healthy, 1 unhealthy; 2 is reserved and must not be used

Up in docker ps means the main process has not exited; it says nothing about whether the process can answer a request, reach its database or has finished loading its model. HEALTHCHECK is the engine's way of asking the container that question on a schedule, and the table is most of what there is to know about it, with two details that catch people: the first check runs one interval after start unless a start period is set (5.1 s after StartedAt with --interval=5s on the engine used here), and a check that hangs is killed at the timeout and counted as a failure, so a probe that opens a database connection can mark a healthy service unhealthy when the database is slow.

A probe that reports on this process, not on the world

Dockerfile
FROM node:22-alpine
# ...
HEALTHCHECK --interval=15s --timeout=3s --start-period=30s --start-interval=2s --retries=3 \
CMD wget -qO- http://127.0.0.1:8080/healthz || exit 1
STOPSIGNAL SIGTERM
ENTRYPOINT ["node", "server.js"] # exec form: node is PID 1 and receives the signal

/healthz should answer for the process itself: the listener is up, the worker pool is alive, the last startup step finished. Whether the database is reachable is a different question with a different consequence; a container marked unhealthy because its dependency is down gets restarted by a restart policy or an orchestrator and comes back to the same dependency being down, now with a restart storm on top. Dependency health belongs in a readiness signal that removes the instance from traffic without killing it, the job of Kubernetes readiness probes and of a load balancer's target health.

bash — observed: the state machine on Engine 28.5.2 (probe: test -f /tmp/healthy; interval 2s, timeout 1s, start-period 4s, start-interval 1s, retries 3)observed
docker run -d --name api p2c-hc-health && docker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" api
starting streak=0
wait_for api healthy 6 # fixture helper: polls .State.Health.Status twice a second
healthy after 1s # a passing probe inside the start period ends it early
docker exec api rm /tmp/healthy && wait_for api unhealthy 12
unhealthy after 6s # three failures, one interval apart
docker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" api
unhealthy streak=3
docker exec api touch /tmp/healthy && wait_for api healthy 6
healthy after 2s
docker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" api
healthy streak=0 # one success resets the streak
docker inspect -f "{{json .State.Health.Log}}" api | jq "length, .[-1].ExitCode"
5
0
the last five probe results live in the inspect output, which is where an intermittent failure is diagnosed. A probe that hangs (CMD sleep 10 with --timeout=1s) logs ExitCode -1 and "Health check exceeded timeout (1s)", and two of those made the container unhealthy in 7s

Shutdown: the signal, the grace period, and the shell that eats both

docker stop and every orchestrator do the same thing: send the stop signal (SIGTERM unless STOPSIGNAL says otherwise) to PID 1, wait a grace period, then SIGKILL. Everything that is graceful about a shutdown happens in that window and is the application's job: stop accepting connections, finish what is in flight, flush, exit 0. The shell form of CMD or ENTRYPOINT (CMD node server.js) runs the command under /bin/sh -c, so the shell is PID 1, and a shell does not forward SIGTERM to its child. The application never hears the signal, the grace period expires, and every deploy ends in a SIGKILL that drops whatever was in progress. The exec form in the Dockerfile above is necessary and, on its own, not sufficient: the kernel does not deliver a default-action signal to PID 1, so a process with no SIGTERM handler (sleep, and many small daemons) is killed after the grace period even in exec form. Node installs a handler and exits on SIGTERM; a process that does not needs --init (or tini as the entrypoint), which adds a minimal PID 1 that forwards signals and reaps children.

bash — observed: docker stop -t 3 against four PID 1 arrangements (the same sleep 300 in each image)observed
docker stop -t 3 shell # CMD sleep 300 && echo done
PID 1: /bin/sh -c sleep 300 && echo done stop took 3.2s exit code 137
docker stop -t 3 exec # CMD ["sleep", "300"]
PID 1: sleep 300 stop took 3.2s exit code 137
docker stop -t 3 init # the same image, docker run --init
PID 1: /sbin/docker-init -- docker-entrypoint.sh sleep 300 stop took 0.1s exit code 143
docker stop -t 3 trap # CMD ["bash","-c","trap 'echo draining; exit 0' TERM; while :; do sleep 1; done"]
PID 1: bash -c trap … stop took 1.1s exit code 0
137 is 128+9: SIGKILL after the grace period. 143 is 128+15: the process died of SIGTERM as soon as tini forwarded it. The trap case exits 0 on its own terms, after its current sleep 1. A shell-form CMD with a single command was exec-ed by busybox sh here, so PID 1 was sleep after all, and it was still killed: the shell is not the only reason
server.js (drain on SIGTERM)
let healthy = true; // read by GET /healthz
const server = app.listen(8080);
process.on('SIGTERM', () => {
healthy = false; // /healthz now returns 503: stop attracting new work
server.close(() => process.exit(0)); // stop accepting; exit when in-flight requests finish
setTimeout(() => process.exit(1), 8000).unref(); // give up before the platform's SIGKILL would
});
compose.yaml
services:
api:
image: registry.acme.dev/shop/api:1.4.0
stop_grace_period: 12s # longer than the 8 s drain above; default is 10 s
depends_on:
db:
condition: service_healthy # wait for db's HEALTHCHECK, not just its start
db:
image: postgres:17
healthcheck:
test: ["CMD-SHELL", "pg_isready -U app -d app"]
interval: 5s
timeout: 3s
retries: 10
start_period: 20s

The drain timeout in the process, the stop_grace_period in Compose and terminationGracePeriodSeconds in Kubernetes are one number seen from three places, and they need to agree: the application gives up slightly before the platform would kill it, so the exit is clean and logged rather than abrupt. depends_on with condition: service_healthy is the other half of the Compose story; plain depends_on orders container start, and an API that starts before Postgres is ready spends its first thirty seconds crash-looping, which the healthcheck on db prevents.

When the healthcheck itself is the outage: recovery in order of preference

SituationDoDo not
a slow dependency makes the probe time out and a restart policy loops the containerdocker update --restart=no <name> to stop the loop, then move the dependency check out of /healthz and into readiness; raise --timeout only if the probe itself is slowraise --retries to 20: the container is still marked unhealthy, later
a new image ships a broken HEALTHCHECK and Compose refuses to start dependantsroll back to the previous digest in compose.yaml (image: …@sha256:<last-good>) and docker compose up -d; fix the probe in the next buildhealthcheck: {disable: true} in the service: service_healthy then means nothing for everyone who depends on it
every deploy ends in exit 137add --init (init: true in Compose) as the immediate fix, then give the process a SIGTERM handler and set stop_grace_period above its drain timeoutraise stop_grace_period alone: a process that never hears the signal waits the whole period and is killed anyway
What was run for this article
Docker Engine 28.5.2 (linux/arm64, Docker Desktop on macOS) with images built from bash:5.2: a flag-file probe with a 2 s interval, a probe that hangs past a 1 s timeout, a 5 s interval with no start period, and the four stop cases. The terminal blocks marked observed are copied from that run, with container names shortened; fifteen state and exit-code assertions are in the fixture script, and the seconds are polled, so a slow machine shows different timings rather than different states. The Node server, the Compose file and the Kubernetes mapping are representative and were not executed. In the recovery table, the --init row is what the stop test showed; the restart-policy and Compose rollback rows follow the documented model.
Kubernetes does not run HEALTHCHECK
The kubelet has its own liveness, readiness and startup probes defined in the Pod spec, and ignores the HEALTHCHECK instruction in the image. An image with a perfect HEALTHCHECK and no probes in its Deployment is, to Kubernetes, a container that is healthy the moment its process starts. Carry the same endpoint over as a readiness probe (traffic) and, only if a restart can actually fix something, a liveness probe; the shutdown side maps to terminationGracePeriodSeconds and a preStop hook.

The two halves meet at deploy time: blue-green and canary switches assume the old pods finish their requests when their endpoints are removed, and Kubernetes probe design is the same readiness-versus-liveness distinction with the platform's names. A Compose file for production is where the healthcheck, the grace period and the restart policy have to be set together rather than discovered one outage at a time.

Go deeper in a courseDocker in depthHealthchecks, signals and PID 1, Compose in production, and what the engine does on stop.View course

Related posts

Quick reference