HEALTHCHECK and graceful shutdown in containers
Tell the orchestrator when your container is actually healthy, and handle SIGTERM so deploys don't drop requests.
HEALTHCHECK options, defaults, and what each one counts
| Option | Default | Meaning |
|---|---|---|
--interval | 30s | time between checks, measured from the end of the previous one |
--timeout | 30s | a probe running longer than this fails, and the probe process is killed with SIGKILL |
--retries | 3 | consecutive failures before the container is unhealthy; one success resets to healthy |
--start-period | 0s | grace for slow starts: failures are not counted until the first success or the period ends |
--start-interval | 5s | check frequency during the start period (Engine 25.0 or later) |
CMD exit code | n/a | 0 healthy, 1 unhealthy; 2 is reserved and must not be used |
Up in docker ps means the main process has not exited; it says nothing about whether the process can answer a request, reach its database or has finished loading its model. HEALTHCHECK is the engine's way of asking the container that question on a schedule, and the table is most of what there is to know about it, with two details that catch people: the first check runs one interval after start unless a start period is set (5.1 s after StartedAt with --interval=5s on the engine used here), and a check that hangs is killed at the timeout and counted as a failure, so a probe that opens a database connection can mark a healthy service unhealthy when the database is slow.
A probe that reports on this process, not on the world
FROM node:22-alpine# ...HEALTHCHECK --interval=15s --timeout=3s --start-period=30s --start-interval=2s --retries=3 \CMD wget -qO- http://127.0.0.1:8080/healthz || exit 1STOPSIGNAL SIGTERMENTRYPOINT ["node", "server.js"] # exec form: node is PID 1 and receives the signal
/healthz should answer for the process itself: the listener is up, the worker pool is alive, the last startup step finished. Whether the database is reachable is a different question with a different consequence; a container marked unhealthy because its dependency is down gets restarted by a restart policy or an orchestrator and comes back to the same dependency being down, now with a restart storm on top. Dependency health belongs in a readiness signal that removes the instance from traffic without killing it, the job of Kubernetes readiness probes and of a load balancer's target health.
docker run -d --name api p2c-hc-health && docker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" apistarting streak=0wait_for api healthy 6 # fixture helper: polls .State.Health.Status twice a secondhealthy after 1s # a passing probe inside the start period ends it earlydocker exec api rm /tmp/healthy && wait_for api unhealthy 12unhealthy after 6s # three failures, one interval apartdocker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" apiunhealthy streak=3docker exec api touch /tmp/healthy && wait_for api healthy 6healthy after 2sdocker inspect -f "{{.State.Health.Status}} streak={{.State.Health.FailingStreak}}" apihealthy streak=0 # one success resets the streakdocker inspect -f "{{json .State.Health.Log}}" api | jq "length, .[-1].ExitCode"50the last five probe results live in the inspect output, which is where an intermittent failure is diagnosed. A probe that hangs (CMD sleep 10 with --timeout=1s) logs ExitCode -1 and "Health check exceeded timeout (1s)", and two of those made the container unhealthy in 7sShutdown: the signal, the grace period, and the shell that eats both
docker stop and every orchestrator do the same thing: send the stop signal (SIGTERM unless STOPSIGNAL says otherwise) to PID 1, wait a grace period, then SIGKILL. Everything that is graceful about a shutdown happens in that window and is the application's job: stop accepting connections, finish what is in flight, flush, exit 0. The shell form of CMD or ENTRYPOINT (CMD node server.js) runs the command under /bin/sh -c, so the shell is PID 1, and a shell does not forward SIGTERM to its child. The application never hears the signal, the grace period expires, and every deploy ends in a SIGKILL that drops whatever was in progress. The exec form in the Dockerfile above is necessary and, on its own, not sufficient: the kernel does not deliver a default-action signal to PID 1, so a process with no SIGTERM handler (sleep, and many small daemons) is killed after the grace period even in exec form. Node installs a handler and exits on SIGTERM; a process that does not needs --init (or tini as the entrypoint), which adds a minimal PID 1 that forwards signals and reaps children.
docker stop -t 3 shell # CMD sleep 300 && echo donePID 1: /bin/sh -c sleep 300 && echo done stop took 3.2s exit code 137docker stop -t 3 exec # CMD ["sleep", "300"]PID 1: sleep 300 stop took 3.2s exit code 137docker stop -t 3 init # the same image, docker run --initPID 1: /sbin/docker-init -- docker-entrypoint.sh sleep 300 stop took 0.1s exit code 143docker stop -t 3 trap # CMD ["bash","-c","trap 'echo draining; exit 0' TERM; while :; do sleep 1; done"]PID 1: bash -c trap … stop took 1.1s exit code 0137 is 128+9: SIGKILL after the grace period. 143 is 128+15: the process died of SIGTERM as soon as tini forwarded it. The trap case exits 0 on its own terms, after its current sleep 1. A shell-form CMD with a single command was exec-ed by busybox sh here, so PID 1 was sleep after all, and it was still killed: the shell is not the only reasonlet healthy = true; // read by GET /healthzconst server = app.listen(8080);process.on('SIGTERM', () => {healthy = false; // /healthz now returns 503: stop attracting new workserver.close(() => process.exit(0)); // stop accepting; exit when in-flight requests finishsetTimeout(() => process.exit(1), 8000).unref(); // give up before the platform's SIGKILL would});
services:api:image: registry.acme.dev/shop/api:1.4.0stop_grace_period: 12s # longer than the 8 s drain above; default is 10 sdepends_on:db:condition: service_healthy # wait for db's HEALTHCHECK, not just its startdb:image: postgres:17healthcheck:test: ["CMD-SHELL", "pg_isready -U app -d app"]interval: 5stimeout: 3sretries: 10start_period: 20s
The drain timeout in the process, the stop_grace_period in Compose and terminationGracePeriodSeconds in Kubernetes are one number seen from three places, and they need to agree: the application gives up slightly before the platform would kill it, so the exit is clean and logged rather than abrupt. depends_on with condition: service_healthy is the other half of the Compose story; plain depends_on orders container start, and an API that starts before Postgres is ready spends its first thirty seconds crash-looping, which the healthcheck on db prevents.
When the healthcheck itself is the outage: recovery in order of preference
| Situation | Do | Do not |
|---|---|---|
| a slow dependency makes the probe time out and a restart policy loops the container | docker update --restart=no <name> to stop the loop, then move the dependency check out of /healthz and into readiness; raise --timeout only if the probe itself is slow | raise --retries to 20: the container is still marked unhealthy, later |
| a new image ships a broken HEALTHCHECK and Compose refuses to start dependants | roll back to the previous digest in compose.yaml (image: …@sha256:<last-good>) and docker compose up -d; fix the probe in the next build | healthcheck: {disable: true} in the service: service_healthy then means nothing for everyone who depends on it |
| every deploy ends in exit 137 | add --init (init: true in Compose) as the immediate fix, then give the process a SIGTERM handler and set stop_grace_period above its drain timeout | raise stop_grace_period alone: a process that never hears the signal waits the whole period and is killed anyway |
The two halves meet at deploy time: blue-green and canary switches assume the old pods finish their requests when their endpoints are removed, and Kubernetes probe design is the same readiness-versus-liveness distinction with the platform's names. A Compose file for production is where the healthcheck, the grace period and the restart policy have to be set together rather than discovered one outage at a time.
Go deeper in a courseDocker in depthHealthchecks, signals and PID 1, Compose in production, and what the engine does on stop.View course