Kubernetes liveness, readiness, startup probes done right

Stop routing traffic to pods that aren't ready and restarts that make outages worse — the three probes explained.

Jun 24, 2025·Updated ·7 min readIntermediate·By SecOpsLog · command-tested

The kubelet knows a container's process is running. It does not know whether the process can serve a request, whether it is deadlocked, or whether it is thirty seconds into a ninety-second cache warm-up. Probes answer those questions, and each of the three answers has a different consequence, which is the only thing that matters when choosing which check goes where.

Three probes, three consequences

ProbeQuestionWhen it failsWhat belongs in it
readinesscan this Pod take traffic right now?the Pod is removed from Service endpoints; nothing is restarteddependency reachability, cache warmed, migrations done, draining
livenessis this process wedged beyond recovery?the container is restarted (after failureThreshold failures)a check of the process itself: event loop alive, no deadlock
startuphas the process finished starting?the container is restarted; until it succeeds, the other two are not runthe same check as liveness, with a long window
Three probes, by what a failure changes

Readiness talks to the Service; liveness talks to the kubelet’s restart logic; startup holds the other two back. The bottom row is the same database check in each of the first two: one removes Pods from traffic while the failover lasts, the other restarts the fleet.

Kubernetes probes by consequence: a failing startup probe keeps the other two waiting and restarts the container when its threshold is exceeded; a failing readiness probe removes the Pod from the Service endpoints without a restart; a failing liveness probe makes the kubelet kill and restart the container. A database check belongs in readiness, where a failover only removes Pods from traffic, never in liveness, where it restarts the whole fleet kubelet: one process per container, three questions, three different consequencesstartupProbealone until it passesfail: keep tryingover threshold: restartpass: hands overreadinessProberuns every periodSecondsfail: out of endpointsno restart, keeps runningpass: back in the ServicelivenessProberuns every periodSecondsfail x failureThreshold:kubelet kills containerrestartPolicy restarts itthe failure path: a dependency check in the wrong probereadiness checks the databaseDB failover: every Pod NotReadyzero endpoints while it lasts,zero restarts; recovers by itselfliveness checks the databaseDB failover: every replica fails 3xkubelet restarts the whole fleet;pools rebuild after the DB is backReadiness may look outward, because its failure means "stop sending me traffic". Liveness may onlylook at the process itself: event loop alive, no deadlock. Startup is liveness with a long window.

The failure that turns a blip into an outage

A liveness handler that checks the database looks thorough and is a fleet-wide restart waiting to happen: the database pauses for a failover, every replica's liveness fails three times, the kubelet restarts all of them at once, and the service that would have recovered in ten seconds now spends two minutes rebuilding connection pools while the incident channel fills with a Kubernetes-shaped explanation of a database event. Readiness is the probe that may look outward, because its failure mode is stop sending me traffic, which is exactly right while a dependency is down. Liveness may only look inward.

deploy.yaml (probes)
readinessProbe:
httpGet: { path: /ready, port: 8080 } # checks dependencies and warm-up
periodSeconds: 5
failureThreshold: 3
livenessProbe:
httpGet: { path: /healthz, port: 8080 } # answers only: is this process responsive
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
terminationGracePeriodSeconds: 20 # a wedged process gets 20 s, not the Pod's 60 s
startupProbe:
httpGet: { path: /healthz, port: 8080 }
periodSeconds: 5
failureThreshold: 36 # up to 180 s to come up before anything else runs

The startup probe is what makes a slow-booting JVM or .NET service survivable without loosening liveness for its whole life: during the failureThreshold × periodSeconds window only the startup probe runs, and once it succeeds the tight liveness settings take over. The run put two Pods with a 20-second boot next to each other under the same liveness probe (three failures, two seconds apart): the one without a startup probe was restarted after 9 s, before it had served anything, and was sitting in CrashLoopBackOff with two restarts a little later; the one with a startup probe was Ready after 12 s with zero restarts. The per-probe terminationGracePeriodSeconds on the liveness and startup probes shortens the shutdown for a process the kubelet is restarting because it is unresponsive; the Pod-level grace period still applies to ordinary terminations, where the process is expected to drain. That number is visible in the restart timing: with terminationGracePeriodSeconds: 5 on the probe, a liveness failure turned into a restarted container in 13 s (three failed probes, then the grace period); a busybox httpd that ignores SIGTERM would otherwise have waited out the Pod's 30 s.

Symptoms of probes doing each other’s job
Every replica restarting within the same minute means liveness is checking a shared dependency. CrashLoopBackOff only on fresh Pods means liveness started before the process could answer and there is no startup probe. A Pod that is Ready while returning errors means readiness is a static 200 and checks nothing the request path needs. Each of these is a one-line change in the manifest and a different one.

How probes behave during a rollout and a shutdown

During a rolling update a new Pod failing readiness is the system working: it stays out of the endpoints until it is ready, and maxUnavailable keeps the old Pods serving; the run rolled out a template that never creates its readiness file, rollout status timed out with the new Pod at ready=false and both old Pods still ready=true in the slice, and kubectl rollout undo brought the Deployment back to two ready endpoints. A new Pod failing liveness during a rollout is a misconfiguration: the kubelet restarts it before it ever joins the Service, and the rollout stalls in CrashLoopBackOff with a healthy previous version still running, which is the moment to check whether a startup probe is missing. On the way out the order reverses: at termination the Pod is removed from endpoints and sent SIGTERM at about the same time, so a preStop hook that sleeps a few seconds, or an application that fails readiness on SIGTERM and keeps serving for the drain period, is what stops in-flight requests from hitting a Pod that is already gone from the load balancer's point of view.

bash — observed: endpoints follow readiness; restarts follow liveness (Kubernetes 1.35.0; probes hit files the fixture removes)observed
kubectl get endpointslices -n shop -l kubernetes.io/service-name=api -o jsonpath="{range .items[*].endpoints[*]}{.targetRef.name} ready={.conditions.ready}{'\n'}{end}"
api-b8dd6b8d-phhbs ready=true
api-b8dd6b8d-sdvcr ready=true
kubectl -n shop exec api-b8dd6b8d-phhbs -- rm /www/ready # the readiness endpoint now answers 404
wait_for api-b8dd6b8d-phhbs "{.status.containerStatuses[0].ready}" false 30 # fixture helper: polls the field once a second
{.status.containerStatuses[0].ready} = false after 4s
kubectl get endpointslices -n shop -l kubernetes.io/service-name=api -o jsonpath="…"
api-b8dd6b8d-phhbs ready=false
api-b8dd6b8d-sdvcr ready=true
kubectl -n shop get pod api-b8dd6b8d-phhbs -o jsonpath="ready={.status.containerStatuses[0].ready} restarts={.status.containerStatuses[0].restartCount}"
ready=false restarts=0
Unhealthy Readiness probe failed: HTTP probe failed with statuscode: 404 # from kubectl get events
one Pod out of rotation after failureThreshold x periodSeconds (2 x 2 s), zero restarts: readiness doing its job while a dependency recovers. touch /www/ready and it was back in the slice within a second
kubectl -n shop exec api-b8dd6b8d-phhbs -- rm /www/healthz # now the liveness endpoint answers 404
wait_restart api-b8dd6b8d-phhbs 1 60 # fixture helper: polls restartCount
restartCount=1 after 13s
Unhealthy Liveness probe failed: HTTP probe failed with statuscode: 404
Killing Container api failed liveness probe, will be restarted
ready=true restarts=1 # the restarted container recreated its files
a RESTARTS column climbing on every replica at once is a liveness probe looking at something shared
bash — observed: a rollout whose new Pod never becomes ready, and the way backobserved
kubectl -n shop patch deploy api --type=json -p='[{"op":"replace","path":"/spec/template/spec/containers/0/command","value":["sh","-c","mkdir -p /www && cd /www && touch healthz && exec httpd -f -p 8080 -h /www"]}]' # no /ready, ever
kubectl -n shop rollout status deploy/api --timeout=30s; echo exit=$?
error: timed out waiting for the condition
exit=1
kubectl get endpointslices -n shop -l kubernetes.io/service-name=api -o jsonpath="…"
api-5b94988cb9-bzx7d ready=false
api-b8dd6b8d-phhbs ready=true
api-b8dd6b8d-sdvcr ready=true
the new Pod never joins the Service and is never restarted (its liveness passes); the old two keep serving. This is the rollout doing what readiness is for
kubectl -n shop rollout undo deploy/api && kubectl -n shop rollout status deploy/api --timeout=120s
deployment "api" successfully rolled out
api-b8dd6b8d-phhbs ready=true
api-b8dd6b8d-sdvcr ready=true
undo restores the previous template; once the NotReady Pod is gone the slice lists the two ready endpoints again

When a probe change takes the service down: recovery in order of preference

SymptomDiagnosisReversal and verification
a rollout stalls, rollout status times out, the new Pods are 0/1 Ready with no restartsreadiness never passes for the new template; kubectl describe pod shows Readiness probe failedkubectl rollout undo deployment/<name>; verify with rollout status and the EndpointSlice showing only ready endpoints (executed)
every replica restarts within the same minuteliveness checks a shared dependency; kubectl get events shows Liveness probe failed on all of them at oncemove the dependency check to readiness and redeploy; until then, raising failureThreshold on the liveness probe buys time without changing what is checked (documentation-backed)
new Pods land in CrashLoopBackOff before ever servingliveness fires during a slow boot; describe pod shows the restart before the first successful probeadd a startupProbe with a window longer than the boot (failureThreshold × periodSeconds) and redeploy; verify restartCount stays 0 while the Pod becomes Ready (executed on the fixture’s 20 s boot)
a Pod is Ready while returning errorsreadiness is a static 200 and checks nothing the request path needsmake the readiness handler check what the request needs and redeploy; there is nothing to roll back, the probe was never doing its job (documentation-backed)
What was run for this article
Kubernetes 1.35.0 on a single-node kind v0.31.0 cluster created and deleted by the fixture (kindest/node:v1.35.0 by digest, kubectl v1.35.0, Docker Engine 28.5.2, linux/arm64). The service is busybox 1.37 httpd behind a Service with two replicas; the probes fetch files the fixture removes and restores. The terminal blocks marked observed are copied from that run (Pod names are from the recorded run; every wait polled the API once a second and printed how long it took). Nineteen assertions: the EndpointSlice before and after a readiness failure, the restart after a liveness failure with a 5 s per-probe grace period, the two slow-boot Pods, and the stalled rollout with its undo. The Node, JVM and gRPC examples, the preStop hook and the terminationGracePeriodSeconds at Pod level are documentation-backed and were not exercised.

Probe types, and the one to avoid

httpGet is the default choice: any 2xx or 3xx is success, and the handler can be as cheap as returning a static 200 for liveness. tcpSocket only proves the port is open, which a deadlocked server with an accept queue still passes. grpc probes speak the gRPC health protocol to services that have no HTTP endpoint. exec runs a command in the container, forks a process every period, and needs a shell or binary that a distroless image does not have; it is the probe type to justify rather than default to. Whatever the type, the check must be cheap and must not itself allocate the resources it is checking for.

Go deeper in a courseKubernetes administrationProbes, rollouts, Services and PodDisruptionBudgets as one lifecycle.View course

Probes decide whether a Pod gets traffic and whether it is restarted. They do not decide whether it can be scheduled or how much it may consume (requests and limits take over there), nor how many replicas a node drain leaves running, the job of a PodDisruptionBudget. An image with no shell changes the probe choice too: the exec type is unavailable there, a limitation distroless images turn into a feature.

Related posts

Quick reference