The Docker socket and daemon hardening

Why docker.sock is root, and how to give access without handing it out.

Advanced12 min · lesson 15 of 24
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab adds the learner to the docker group and writes to a file under /root. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate.

A request lands in your queue: "please add the new build engineer to the docker group so they can stop typing sudo". It looks like a convenience setting. On a host running the normal rootful daemon it gives that account root, without a password and without any line in the sudo log. This lesson proves that on a lab VM with a harmless marker file, then shows how to find who already has that access and how to give people what they need without it.

Why the socket is the whole daemon

The docker command is a client. Everything it does is an HTTP request to the Docker Engine API, which dockerd serves on the Unix socket /var/run/docker.sock (and on TCP only if someone configures it). dockerd runs as root, and the local API has no login of its own. Whoever can open the socket can ask for anything the API offers: a container with the host's root filesystem mounted, --privileged, --pid=host, any device. The kernel only checks one thing, the socket file's permissions. Pull the two images the lesson uses first; ubuntu is not in the docker group on the sec VM, so that needs sudo:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker pull -q alpine:3.22 sudo docker pull -q docker:29-cli
docker.io/library/alpine:3.22 docker.io/library/docker:29-cli
$ ls -l /var/run/docker.sock getent group docker id
srw-rw---- 1 root docker 0 Oct 8 03:01 /var/run/docker.sock docker:x:987: uid=1000(ubuntu) gid=1000(ubuntu) groups=1000(ubuntu),4(adm),24(cdrom),27(sudo),30(dip),102(lxd)
$ docker ps
permission denied while trying to connect to the docker API at unix:///var/run/docker.sock

Mode srw-rw----, owner root, group docker. On this VM the docker group is empty and ubuntu is not in it, so the connection is refused. (Profile sec of the lab kit leaves the learner out of the group on purpose; the main VM puts ubuntu in it, as "Installing Docker Engine and running your first container" (Docker for beginners) explains.) Membership of docker therefore means write access to the socket, and write access to the socket means root on the host.

The proof, with a marker file

First create a file only root can read, and confirm that ubuntu cannot:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ echo "lab marker: only root can read this" | sudo tee /root/lab-marker.txt >/dev/null sudo chmod 600 /root/lab-marker.txt sudo ls -l /root/lab-marker.txt
-rw------- 1 root root 36 Oct 8 03:01 /root/lab-marker.txt
$ cat /root/lab-marker.txt
cat: /root/lab-marker.txt: Permission denied

Now do what the request asked for. usermod -aG docker appends the group. Group membership is read at login, so log out of the VM and open a new shell (or start one with newgrp docker) before id shows it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo usermod -aG docker ubuntu getent group docker
docker:x:987:ubuntu
$ id
uid=1000(ubuntu) gid=1000(ubuntu) groups=1000(ubuntu),4(adm),24(cdrom),27(sudo),30(dip),102(lxd),987(docker)

With nothing but that group, start a container that bind-mounts the host's / and use it to read the marker and append a line to it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run --rm --name lab-proof -v /:/host alpine:3.22 sh -c "cat /host/root/lab-marker.txt; echo written by a container >> /host/root/lab-marker.txt"
lab marker: only root can read this
ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo ls -l /root/lab-marker.txt sudo cat /root/lab-marker.txt
-rw------- 1 root root 59 Oct 8 03:01 /root/lab-marker.txt lab marker: only root can read this written by a container

The container printed a file that ubuntu could not read a minute earlier, and appended to it. The host shows the file still owned by root with mode 600, now 59 bytes with the extra line. Nothing exploited a bug. The daemon did exactly what the API promises, as root, for a caller the kernel allowed onto the socket. The same request could have changed any file on the host, which is why this lab touches only its own marker. A socket mounted into a container gives that container the same power, and so does a TCP listener without mutual TLS.

Container hardening flags do not help here. Seccomp, --cap-drop and --read-only restrict the container that makes the request; the new container is created by dockerd with whatever the request specifies. Escapes that go through a mounted socket are covered from the attacker's side in "The big three: privileged, docker.sock and host mounts".

Detection

Three questions cover most of it: which containers were created, and with what; which containers hold the socket; who is in the group.

The daemon publishes an event for every object change. Query a time window and filter for container creation:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker events --since 2m --until 0s --filter type=container --filter event=create \ --format "{{.Time}} {{.Action}} {{.Actor.Attributes.name}} image={{.Actor.Attributes.image}}"

{{.Time}} is a Unix timestamp; the line records that lab-proof was created from alpine:3.22. Events do not list mounts, so pair them with docker inspect while the container exists, and stream them to your log pipeline (docker events --format '{{json .}}' running under systemd, or the daemon's own logging), because docker events only keeps a limited recent history and loses it on restart. For a record of which local process opened the socket, an auditd watch such as auditctl -w /var/run/docker.sock -k docker-sock adds the process name; it is not part of this lab because auditd is not installed on the VM.

Monitoring agents, CI runners and reverse proxies often ask for the socket, frequently with :ro in the belief that it limits them. Mount it that way and sweep for it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run -d --name lab-sockuser -v /var/run/docker.sock:/var/run/docker.sock:ro docker:29-cli sleep 300 >/dev/null docker ps -q | xargs docker inspect --format "{{.Name}}{{range .Mounts}} {{.Source}}:{{.Mode}}{{end}}" | grep docker.sock
/lab-sockuser /var/run/docker.sock:ro
$ docker exec lab-sockuser docker ps --format "{{.Names}} {{.Image}}"
lab-sockuser docker:29-cli

The sweep finds /lab-sockuser with the socket mounted read-only, and the Docker CLI inside it still lists containers through the API. :ro makes the socket file read-only as a file; connecting to a socket and sending requests is not a file write, so :ro does not limit the API at all. That container could have run the marker proof as easily as ubuntu did.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ getent group docker | cut -d: -f4
ubuntu

Run this on every host on a schedule and compare it with an approved list; directory-managed groups need the same check in your identity system. Treat each name as a root account.

Mitigations

For people who administer the host anyway, sudo docker costs a few keystrokes and puts every command in the log with the account that ran it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker rm -f lab-sockuser sudo gpasswd -d ubuntu docker
lab-sockuser Removing user ubuntu from group docker
$ id -nG docker ps
ubuntu adm cdrom sudo dip lxd permission denied while trying to connect to the docker API at unix:///var/run/docker.sock
$ sudo docker ps --format "{{.Names}}" sudo journalctl _COMM=sudo --since -2min --no-pager -o cat | grep "COMMAND=/usr/bin/docker" | tail -1
ubuntu : TTY=/dev/pts/0 ; PWD=/home/ubuntu ; USER=root ; COMMAND=/usr/bin/docker ps --format {{.Names}}

After gpasswd -d and a new login, the socket refuses ubuntu again, and the journal records ubuntu running /usr/bin/docker ps as root. Two caveats: an existing login session keeps the group until it ends, and a sudo rule that allows only /usr/bin/docker still allows docker run -v /:/host, so it is an audit trail, not a sandbox.

Most tools that ask for it need far less. If a tool must read container metadata, put a filtering proxy in front of the socket and give the tool the proxy, never the real socket. Allow only the endpoints the tool needs, for example container listing, and leave exec, POST requests, archive downloads and image export off. Read-only does not mean harmless: GET /containers/{id}/json returns every container's environment variables, which is where -e passwords live ("Runtime secrets, done right"), and GET /containers/{id}/archive downloads files out of any container.

Developers who want to build and run containers without sudo can get a daemon of their own. The rootless daemon's socket still controls that user's whole account, but not the host. "Rootless Docker" sets it up.

Build jobs run code from every branch, and a runner that mounts the host socket gives every job host root. Use ephemeral VMs as runners, a separate build host per trust level, or a daemonless builder: BuildKit in rootless mode (moby/buildkit:rootless), or a remote BuildKit builder reached with docker buildx create --driver remote over mutual TLS. On Kubernetes, the node's containerd or CRI-O socket has the same property, so a pod that mounts it owns the node; build with a rootless BuildKit pod instead.

An authorization plugin ("authorization-plugins" in daemon.json) sees each API request and can allow or deny it, for example refusing bind mounts of / or --privileged. They exist but are rarely used, every rule must cover the whole API, and they have had bypass bugs, the most recent CVE-2026-34040, fixed in Engine 29.3.1. Use one as an extra layer, not as the reason it is acceptable to hand out the socket.

An unauthenticated TCP listener (-H tcp://0.0.0.0:2375) is the socket offered to the network, and scanners find such listeners quickly. If you need remote access, prefer SSH, which reuses accounts you already manage: docker context create remote --docker host=ssh://admin@build01. If the API must be on TCP, require client certificates:

daemon and client flags
# daemon (or the same keys in /etc/docker/daemon.json):
dockerd --tlsverify --tlscacert=ca.pem --tlscert=server-cert.pem --tlskey=server-key.pem -H=0.0.0.0:2376
# client: only holders of a cert signed by ca.pem get in
docker --tlsverify --tlscacert=ca.pem --tlscert=cert.pem --tlskey=key.pem -H=build01:2376 version

Other daemon settings that reduce damage (no-new-privileges, live-restore, icc, log limits) and the safe procedure for editing daemon.json are in "Configuring the daemon safely" (Docker in depth). If you consider userns-remap there, read "User-namespace remapping" first: on Docker 29 it moves the daemon off the containerd image store.

Clean up

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo rm /root/lab-marker.txt

The group membership was already removed in the mitigation step; check with getent group docker.

Quick check
01A monitoring agent runs with -v /var/run/docker.sock:/var/run/docker.sock:ro, --cap-drop ALL and --read-only. An attacker gets code execution inside it. What can they do?
Incorrect — In the lab, a container holding the socket :ro ran docker ps through the API. :ro affects writes to the socket file, not requests sent over it.
Incorrect — Those flags limit this container. The attacker asks dockerd, which runs as root, to create a new container with whatever mounts they choose.
Correct — The socket gives full API access. The new container is created by the root daemon, so the agent container's own restrictions do not apply to it.
Incorrect — dockerd does not know or care how the caller reached the socket. The lab container appended to a root-owned 600 file.
02Which audit finding on a rootful Docker host should you treat with the same urgency as an unknown entry in the sudoers file?
Correct — Group membership gives write access to the socket, and the lab showed that is enough to read and change root-only files.
Incorrect — A read-only root filesystem is a hardening setting, not a risk.
Incorrect — A loopback-only published port is reachable only from the host itself; it does not give anyone control of the daemon.
Incorrect — That is the mitigation working: the commands are logged with the account that ran them.
03Your CI runners mount the host's docker.sock so jobs can build images. Which change removes host-root exposure from untrusted branches?
Incorrect — :ro does not limit API requests, so every job keeps full control of the daemon.
Incorrect — Seccomp filters the runner's own syscalls. The containers the job asks for are created by dockerd as root.
Incorrect — Group membership is the same root-equivalent access to the socket, granted a different way.
Correct — A daemonless or rootless builder, or a disposable machine per job, means a malicious job cannot reach a root daemon on a shared host.

Try this

Work through “Clean up” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from the docker socket and daemon hardening, keep “Clean up”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related