Engine architecture on Docker 29
CLI, dockerd, containerd, the shim and runc, and what each owns on Docker 29.
During an incident someone types docker ps and it hangs. The web containers on the same host keep answering requests the whole time. That combination is possible because "Docker" on a Linux host is a chain of separate programs, each with its own job and its own failure modes, and the program that serves the API is not the program that keeps your containers alive. This lesson follows one container down that chain on the lab VM so you can name each process, see what it owns on Docker 29, and know where to look first when one of them misbehaves.
Everything here runs on the main lab VM, secopslog-docker, and changes nothing on it. ctr and the containerd paths need sudo; the rest runs as ubuntu, which is in the docker group.
Client and server in one command
docker version is the first thing to paste into any bug report, because it describes both ends of the conversation. The Client block is the docker CLI you typed. The Server block is what the daemon reports about itself and the components it drives: the Engine (dockerd), containerd, runc and docker-init.
Three of those lines matter most in practice. API version: 1.56 (minimum version 1.40) says which REST API versions this daemon accepts. containerd v2.3.6 and runc 1.5.1 are separate projects with their own security advisories, so when a CVE lands against runc you check this line, not the Docker version. docker-init 0.19.0 is the small init process (tini) that Docker injects as PID 1 only when you run a container with --init; "CMD, ENTRYPOINT and PID 1" in Docker for beginners covers when you want that. On an amd64 machine the OS/Arch lines read linux/amd64; the rest is the same.
docker info gives the daemon's runtime view, and its --format templates are a common place for old runbooks to break. A template that names a field the daemon does not have prints what it managed so far and then fails:
There is no CRI field in docker info, so the template stops at it with exit status 1. Ask for fields that exist:
store overlayfs is the containerd image store with its overlayfs snapshotter, the default on a fresh Docker 29 install. cgroup v2 (systemd) means each container gets a systemd scope, which you will see below. live-restore false is the shipped default, and it decides what happens to containers when dockerd restarts.
API version negotiation
The CLI talks to dockerd over HTTP on a Unix socket, and every request path carries an API version (/v1.56/containers/json). On connect, the client pings the daemon, reads the Api-Version header and lowers its own version to the daemon's if the daemon is older. A new CLI therefore works against an older engine. The reverse has a floor: the daemon refuses any version below its minimum. Docker 29.0 raised that minimum to 1.44, which broke many old SDK-based tools; 29.3.0 lowered it back to 1.40, the value this daemon reports.
DOCKER_API_VERSION pins the client instead of negotiating, which is useful for reproducing what an older tool sends. Pinned at 1.44 it still works; pinned at 1.30 the daemon answers with the error in the last step. When a monitoring agent or CI plugin starts failing after an engine upgrade with "client version ... is too old", that agent is speaking an API the daemon no longer accepts, and the fix is a newer agent or SDK. The raw API shows the same negotiation data without the CLI in the way:
That curl worked without sudo for one reason: ubuntu is in the docker group, and the group can read and write the socket. Anyone who can talk to this socket can ask a root daemon to start a privileged container with the host's filesystem mounted, so socket access is root access. "The Docker socket and daemon hardening" in Advanced container security takes that apart and shows the mitigations.
What systemd runs
Three units make up the engine. containerd.service and docker.service are the two daemons. docker.socket is a systemd socket unit: systemd creates /var/run/docker.sock (owned by root:docker, mode srw-rw----) and hands the open socket to dockerd, which is what -H fd:// in ExecStart means. --containerd=/run/containerd/containerd.sock tells dockerd to use the separately managed containerd instead of starting its own. Both daemons are direct children of PID 1:
One container, followed down the chain
pstree -s prints the ancestry of the container's main process (nginx; your PIDs and container ID will differ from this run). Its parent is containerd-shim with -namespace moby -id <container id>, and the shim's parent is systemd, PID 1. Neither dockerd nor containerd is in that line. The shim double-forks at start, so it is reparented to PID 1 and survives restarts of both daemons. It holds the container's stdio pipes, collects its exit status, and serves a ttrpc socket (ttrpc is containerd's lightweight variant of gRPC, the RPC protocol dockerd uses to talk to containerd) that containerd connects to whenever it needs to signal or inspect the task.
runc does the kernel work: it reads the OCI bundle (a config.json plus a root filesystem), creates the namespaces and the cgroup, sets up mounts, capabilities, seccomp and AppArmor, starts the process and exits. pgrep -x runc finds nothing because there is nothing left to find. A failure inside that short life shows up as an error that mentions the OCI runtime when the container is created or started. The cgroup it created is visible as a systemd scope named after the container ID:
What containerd holds on Docker 29
containerd is multi-tenant. Docker keeps its objects in the moby namespace (moby_history holds image history data), and ctr, containerd's debugging client, can read them directly:
The container record names the image (docker.io/library/nginx:1.30-alpine, the fully qualified form of the short name you typed) and the runtime io.containerd.runc.v2, which is the shim you saw in pstree. The task is the running instance; tasks ps lists the same nginx PIDs as pstree. The image list matters most: with the containerd image store, image content lives in containerd's content store and is unpacked into its snapshotter. docker image ls is a view over that store, and both commands list the same image.
The container's root filesystem is an overlay mount that dockerd places under /var/lib/docker/rootfs/overlayfs/<id>, but the read-only lower layers and the writable upper layer are snapshot directories under /var/lib/containerd/. That split has operational consequences (moving data-root does not move images, for one), and "Where Docker keeps data" covers it, including hosts upgraded from older engines that still use the overlay2 graph driver under /var/lib/docker.
dockerd still owns plenty: the API, networks and their firewall rules, volumes, container configuration and log files under /var/lib/docker/containers/, and the builder. BuildKit runs inside dockerd, not as a separate service, and uses containerd as its executor and snapshotter:
The docker driver in that output is the builder every engine has; "Buildx builders, outputs, cache and Bake" shows the separate docker-container builders you can add next to it. Clean up the lab container:
Why containers can outlive dockerd
Since the shim parents each container, nothing in the process tree forces a container to die when dockerd exits. Whether it does is a policy choice. With the default live-restore: false, dockerd stops all containers when it shuts down, so systemctl restart docker restarts every workload on the host. With live-restore: true, dockerd leaves them running and reconnects to their shims when it comes back. The same design explains the hung docker ps from the opening: a stuck API server does not stop processes that the shims supervise. "Configuring the daemon safely" turns live-restore on in the sec lab and shows the shim PID unchanged across a daemon restart. Live-restore covers daemon restarts and patch upgrades, not host reboots, and it does not apply to Swarm services.
The same layering is why Kubernetes nodes run containerd without dockerd. The kubelet talks to containerd through the Container Runtime Interface (CRI), and its containers live in the k8s.io namespace, so docker ps on such a node shows nothing even though sudo ctr -n k8s.io containers ls lists every pod container. Rootless Docker runs this same chain as an ordinary user inside a user namespace, with the socket under $XDG_RUNTIME_DIR; "Rootless Docker" in Advanced container security covers how it differs.
Triage by layer
Read the error before restarting anything. Restarting docker.service without live-restore stops every container on the host, which turns a broken API into an outage.
pstree -s -p on the nginx PID shows systemd,1 then containerd-shim,... then nginx, and pgrep -x runc finds nothing. What explains that tree?client version 1.24 is too old. Minimum supported API version is 1.40. The docker CLI on the host works. What is the right fix?du shows /var/lib/containerd at 40 GB and /var/lib/docker at 2 GB. What is most of that 40 GB?Try this
Work through “Triage by layer” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from engine architecture on docker 29, keep “Triage by layer”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.