Engine architecture on Docker 29

CLI, dockerd, containerd, the shim and runc, and what each owns on Docker 29.

Intermediate12 min · lesson 7 of 24

During an incident someone types docker ps and it hangs. The web containers on the same host keep answering requests the whole time. That combination is possible because "Docker" on a Linux host is a chain of separate programs, each with its own job and its own failure modes, and the program that serves the API is not the program that keeps your containers alive. This lesson follows one container down that chain on the lab VM so you can name each process, see what it owns on Docker 29, and know where to look first when one of them misbehaves.

Everything here runs on the main lab VM, secopslog-docker, and changes nothing on it. ctr and the containerd paths need sudo; the rest runs as ubuntu, which is in the docker group.

Client and server in one command

docker version is the first thing to paste into any bug report, because it describes both ends of the conversation. The Client block is the docker CLI you typed. The Server block is what the daemon reports about itself and the components it drives: the Engine (dockerd), containerd, runc and docker-init.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker version
Client: Docker Engine - Community Version: 29.8.2 API version: 1.56 Go version: go1.26.8 Git commit: 7fc2dff Built: Wed Sep 30 19:32:08 2026 OS/Arch: linux/arm64 Context: default Server: Docker Engine - Community Engine: Version: 29.8.2 API version: 1.56 (minimum version 1.40) Go version: go1.26.8 Git commit: 8af9fe3 Built: Wed Sep 30 19:32:08 2026 OS/Arch: linux/arm64 Experimental: false containerd: Version: v2.3.6 GitCommit: ee2735368117d2eb259779949d5e75cdafec9761 runc: Version: 1.5.1 GitCommit: v1.5.1-0-g8f2685a4 docker-init: Version: 0.19.0 GitCommit: de40ad0

Three of those lines matter most in practice. API version: 1.56 (minimum version 1.40) says which REST API versions this daemon accepts. containerd v2.3.6 and runc 1.5.1 are separate projects with their own security advisories, so when a CVE lands against runc you check this line, not the Docker version. docker-init 0.19.0 is the small init process (tini) that Docker injects as PID 1 only when you run a container with --init; "CMD, ENTRYPOINT and PID 1" in Docker for beginners covers when you want that. On an amd64 machine the OS/Arch lines read linux/amd64; the rest is the same.

docker info gives the daemon's runtime view, and its --format templates are a common place for old runbooks to break. A template that names a field the daemon does not have prints what it managed so far and then fails:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker info --format 'server={{.ServerVersion}} cri={{.CRI}}'
server=29.8.2 cri= template: :1:32: executing "" at <.CRI>: can't evaluate field CRI in type system.dockerInfo

There is no CRI field in docker info, so the template stops at it with exit status 1. Ask for fields that exist:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker info --format 'server {{.ServerVersion}}, store {{.Driver}}, cgroup v{{.CgroupVersion}} ({{.CgroupDriver}}), runtime {{.DefaultRuntime}}, live-restore {{.LiveRestoreEnabled}}'
server 29.8.2, store overlayfs, cgroup v2 (systemd), runtime runc, live-restore false

store overlayfs is the containerd image store with its overlayfs snapshotter, the default on a fresh Docker 29 install. cgroup v2 (systemd) means each container gets a systemd scope, which you will see below. live-restore false is the shipped default, and it decides what happens to containers when dockerd restarts.

API version negotiation

The CLI talks to dockerd over HTTP on a Unix socket, and every request path carries an API version (/v1.56/containers/json). On connect, the client pings the daemon, reads the Api-Version header and lowers its own version to the daemon's if the daemon is older. A new CLI therefore works against an older engine. The reverse has a floor: the daemon refuses any version below its minimum. Docker 29.0 raised that minimum to 1.44, which broke many old SDK-based tools; 29.3.0 lowered it back to 1.40, the value this daemon reports.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker version --format 'client API {{.Client.APIVersion}}, server API {{.Server.APIVersion}}, oldest accepted {{.Server.MinAPIVersion}}'
client API 1.56, server API 1.56, oldest accepted 1.40
$ DOCKER_API_VERSION=1.44 docker version --format 'client now speaks {{.Client.APIVersion}}'
client now speaks 1.44
$ DOCKER_API_VERSION=1.30 docker ps
Error response from daemon: client version 1.30 is too old. Minimum supported API version is 1.40, please upgrade your client to a newer version

DOCKER_API_VERSION pins the client instead of negotiating, which is useful for reproducing what an older tool sends. Pinned at 1.44 it still works; pinned at 1.30 the daemon answers with the error in the last step. When a monitoring agent or CI plugin starts failing after an engine upgrade with "client version ... is too old", that agent is speaking an API the daemon no longer accepts, and the fix is a newer agent or SDK. The raw API shows the same negotiation data without the CLI in the way:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ curl -s --unix-socket /var/run/docker.sock http://localhost/_ping -D - -o /dev/null | grep -E '^(HTTP|Api-Version|Server)' curl -s --unix-socket /var/run/docker.sock http://localhost/v1.56/version | jq -c '{Version, ApiVersion, MinAPIVersion}'
HTTP/1.1 200 OK Api-Version: 1.56 Server: Docker/29.8.2 (linux) {"Version":"29.8.2","ApiVersion":"1.56","MinAPIVersion":"1.40"}

That curl worked without sudo for one reason: ubuntu is in the docker group, and the group can read and write the socket. Anyone who can talk to this socket can ask a root daemon to start a privileged container with the host's filesystem mounted, so socket access is root access. "The Docker socket and daemon hardening" in Advanced container security takes that apart and shows the mitigations.

What systemd runs

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ systemctl list-units --no-pager --plain --no-legend 'docker*' 'containerd*'
docker-5dc318d8c16386327723b2e46cdaec97420bc4a3d747e7477f56e6df63e319da.scope loaded active running libcontainer container 5dc318d8c16386327723b2e46cdaec97420bc4a3d747e7477f56e6df63e319da containerd.service loaded active running containerd container runtime docker.service loaded active running Docker Application Container Engine docker.socket loaded active running Docker Socket for the API
$ systemctl cat docker.service | grep -E '^(ExecStart|Requires)=' ls -l /var/run/docker.sock
Requires=docker.socket ExecStart=/usr/bin/dockerd -H fd:// --containerd=/run/containerd/containerd.sock srw-rw---- 1 root docker 0 Oct 7 23:29 /var/run/docker.sock

Three units make up the engine. containerd.service and docker.service are the two daemons. docker.socket is a systemd socket unit: systemd creates /var/run/docker.sock (owned by root:docker, mode srw-rw----) and hands the open socket to dockerd, which is what -H fd:// in ExecStart means. --containerd=/run/containerd/containerd.sock tells dockerd to use the separately managed containerd instead of starting its own. Both daemons are direct children of PID 1:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ ps -ww -o pid,ppid,user,args -C dockerd,containerd
PID PPID USER COMMAND 16873 1 root /usr/bin/containerd 17027 1 root /usr/bin/dockerd -H fd:// --containerd=/run/containerd/containerd.sock
One docker run on Docker 29
1docker CLI
HTTP over /var/run/docker.sock, API version negotiated
2dockerd
API, networks, volumes, container config, embedded BuildKit
3containerd
content store, snapshotter, containers and tasks (namespace moby)
4containerd-shim-runc-v2
one per container; parent of the container process
5runc
creates namespaces and cgroups, starts the process, exits
dockerd talks to containerd over gRPC; containerd talks to each shim over ttrpc. runc is gone once the container is running, and the shim, not dockerd, is the container's parent.

One container, followed down the chain

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-web nginx:1.30-alpine
0fb09fa51c6c432ec9ed84f5127feee8090ea38f2f6311c8696108a660000fa3
$ pstree -l -s -p -a $(docker inspect -f '{{.State.Pid}}' lab-web)
systemd,1 --switched-root --system --deserialize=48 `-containerd-shim,210628 -namespace moby -id 0fb09fa51c6c432ec9ed84f5127feee8090ea38f2f6311c8696108a660000fa3 -address /run/containerd/containerd.sock `-nginx,210653 |-nginx,210733 |-nginx,210734 |-nginx,210735 `-nginx,210736
$ pgrep -a -x runc || echo "no runc process"
no runc process

pstree -s prints the ancestry of the container's main process (nginx; your PIDs and container ID will differ from this run). Its parent is containerd-shim with -namespace moby -id <container id>, and the shim's parent is systemd, PID 1. Neither dockerd nor containerd is in that line. The shim double-forks at start, so it is reparented to PID 1 and survives restarts of both daemons. It holds the container's stdio pipes, collects its exit status, and serves a ttrpc socket (ttrpc is containerd's lightweight variant of gRPC, the RPC protocol dockerd uses to talk to containerd) that containerd connects to whenever it needs to signal or inspect the task.

runc does the kernel work: it reads the OCI bundle (a config.json plus a root filesystem), creates the namespaces and the cgroup, sets up mounts, capabilities, seccomp and AppArmor, starts the process and exits. pgrep -x runc finds nothing because there is nothing left to find. A failure inside that short life shows up as an error that mentions the OCI runtime when the container is created or started. The cgroup it created is visible as a systemd scope named after the container ID:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ systemctl list-units --no-pager --plain --no-legend 'docker-*.scope' docker inspect -f '{{.Id}}' lab-web
docker-0fb09fa51c6c432ec9ed84f5127feee8090ea38f2f6311c8696108a660000fa3.scope loaded active running libcontainer container 0fb09fa51c6c432ec9ed84f5127feee8090ea38f2f6311c8696108a660000fa3 docker-5dc318d8c16386327723b2e46cdaec97420bc4a3d747e7477f56e6df63e319da.scope loaded active running libcontainer container 5dc318d8c16386327723b2e46cdaec97420bc4a3d747e7477f56e6df63e319da 0fb09fa51c6c432ec9ed84f5127feee8090ea38f2f6311c8696108a660000fa3

What containerd holds on Docker 29

containerd is multi-tenant. Docker keeps its objects in the moby namespace (moby_history holds image history data), and ctr, containerd's debugging client, can read them directly:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ sudo ctr namespaces ls
NAME LABELS moby moby_history
$ ID=$(docker inspect -f '{{.Id}}' lab-web) sudo ctr -n moby containers info $ID | jq -c '{Image, Runtime: .Runtime.Name}' sudo ctr -n moby tasks ps $ID
{"Image":"docker.io/library/nginx:1.30-alpine","Runtime":"io.containerd.runc.v2"} PID INFO 210653 - 210733 - 210734 - 210735 - 210736 -
$ sudo ctr -n moby images ls -q | grep nginx docker image ls nginx:1.30-alpine
docker.io/library/nginx:1.30-alpine IMAGE ID DISK USAGE CONTENT SIZE EXTRA nginx:1.30-alpine 0985e772fb9f 92.9MB 26.9MB U

The container record names the image (docker.io/library/nginx:1.30-alpine, the fully qualified form of the short name you typed) and the runtime io.containerd.runc.v2, which is the shim you saw in pstree. The task is the running instance; tasks ps lists the same nginx PIDs as pstree. The image list matters most: with the containerd image store, image content lives in containerd's content store and is unpacked into its snapshotter. docker image ls is a view over that store, and both commands list the same image.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ ID=$(docker inspect -f '{{.Id}}' lab-web) sudo findmnt -n -o FSTYPE /var/lib/docker/rootfs/overlayfs/$ID sudo findmnt -n -o OPTIONS /var/lib/docker/rootfs/overlayfs/$ID | tr , '\n' | grep -E '^(lowerdir|upperdir)' | cut -c1-95
overlay lowerdir=/var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/1462/fs:/var/lib/ upperdir=/var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/1463/fs

The container's root filesystem is an overlay mount that dockerd places under /var/lib/docker/rootfs/overlayfs/<id>, but the read-only lower layers and the writable upper layer are snapshot directories under /var/lib/containerd/. That split has operational consequences (moving data-root does not move images, for one), and "Where Docker keeps data" covers it, including hosts upgraded from older engines that still use the overlay2 graph driver under /var/lib/docker.

dockerd still owns plenty: the API, networks and their firewall rules, volumes, container configuration and log files under /var/lib/docker/containers/, and the builder. BuildKit runs inside dockerd, not as a separate service, and uses containerd as its executor and snapshotter:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker buildx inspect default | grep -E '^(Driver|BuildKit version)|executor|snapshotter'
Driver: docker BuildKit version: v0.33.1 org.mobyproject.buildkit.worker.executor: containerd org.mobyproject.buildkit.worker.snapshotter: overlayfs

The docker driver in that output is the builder every engine has; "Buildx builders, outputs, cache and Bake" shows the separate docker-container builders you can add next to it. Clean up the lab container:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-web
lab-web

Why containers can outlive dockerd

Since the shim parents each container, nothing in the process tree forces a container to die when dockerd exits. Whether it does is a policy choice. With the default live-restore: false, dockerd stops all containers when it shuts down, so systemctl restart docker restarts every workload on the host. With live-restore: true, dockerd leaves them running and reconnects to their shims when it comes back. The same design explains the hung docker ps from the opening: a stuck API server does not stop processes that the shims supervise. "Configuring the daemon safely" turns live-restore on in the sec lab and shows the shim PID unchanged across a daemon restart. Live-restore covers daemon restarts and patch upgrades, not host reboots, and it does not apply to Swarm services.

The same layering is why Kubernetes nodes run containerd without dockerd. The kubelet talks to containerd through the Container Runtime Interface (CRI), and its containers live in the k8s.io namespace, so docker ps on such a node shows nothing even though sudo ctr -n k8s.io containers ls lists every pod container. Rootless Docker runs this same chain as an ordinary user inside a user namespace, with the socket under $XDG_RUNTIME_DIR; "Rootless Docker" in Advanced container security covers how it differs.

Triage by layer

Read the error before restarting anything. Restarting docker.service without live-restore stops every container on the host, which turns a broken API into an outage.

Quick check
01On a host with one nginx container, pstree -s -p on the nginx PID shows systemd,1 then containerd-shim,... then nginx, and pgrep -x runc finds nothing. What explains that tree?
Incorrect — Nothing crashed. runc always exits once the process is running; that is its normal life cycle.
Correct — The shim holds stdio and the exit status and is reparented to PID 1, so it survives restarts of dockerd and containerd.
Incorrect — systemd only holds the container's cgroup as a scope. The process is started by runc and parented by the shim.
Incorrect — pstree on the host sees every process, including those in container PID namespaces, with their host PIDs.
02After an engine upgrade to Docker 29, an old monitoring agent logs client version 1.24 is too old. Minimum supported API version is 1.40. The docker CLI on the host works. What is the right fix?
Incorrect — That pins the agent at exactly the version the daemon rejects; the error stays.
Incorrect — API versions are served by dockerd, not containerd, and a restart does not change the minimum.
Incorrect — live-restore concerns containers during daemon restarts; it has nothing to do with API versions.
Correct — The CLI works because it negotiates down from 1.56; the agent is stuck below the daemon's floor and needs a newer client.
03A Docker 29 host installed fresh runs out of disk. du shows /var/lib/containerd at 40 GB and /var/lib/docker at 2 GB. What is most of that 40 GB?
Correct — On the containerd image store, images and container layers are containerd's data; dockerd keeps volumes, container config and logs under /var/lib/docker.
Incorrect — json-file logs still live in /var/lib/docker/containers/<id>/, under dockerd's data-root.
Incorrect — Volumes remain dockerd's, under /var/lib/docker/volumes.
Incorrect — runc keeps a few kilobytes of state under /run, not gigabytes on disk.

Try this

Work through “Triage by layer” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from engine architecture on docker 29, keep “Triage by layer”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related