The big three: privileged, docker.sock and host mounts

The settings that strip isolation, how to detect them, and what to grant instead.

Advanced15 min · lesson 22 of 24

Three settings account for most container-to-host takeovers that do not need a kernel bug: the --privileged flag, a bind-mounted Docker socket, and a host directory mounted into the container. Each is granted on purpose, on a docker run line or in a Compose file, so each can also be refused. This lesson measures what each one removes, shows how to find it in a running container, and gives the narrower setting that does the same job.

Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab creates a root-only marker file under /root and shows a container changing it through a host bind mount. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo. Pull the two images the lesson uses first:
ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker pull -q alpine:3.22 sudo docker pull -q docker:29-cli
docker.io/library/alpine:3.22 docker.io/library/docker:29-cli

--privileged: the confinement stack, off

A default container is defined by what it does not get: a small capability set, an active seccomp filter, an AppArmor profile, and a device cgroup that hides the host's disks. --privileged cancels all of it at once. Read a default container and a privileged one side by side:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm alpine:3.22 sh -c "grep -E \"^(CapEff|Seccomp):\" /proc/1/status; ls /dev | tr \"\n\" \" \"; echo"
CapEff: 00000000a80425fb Seccomp: 2 core fd full mqueue null ptmx pts random shm stderr stdin stdout tty urandom zero
$ sudo docker run --rm --privileged alpine:3.22 sh -c "grep -E \"^(CapEff|Seccomp):\" /proc/1/status; ls /dev | grep -E \"sda|vda\" | tr \"\n\" \" \"; echo"
CapEff: 000001ffffffffff Seccomp: 0 sda sda1 sda13 sda15 vda

The default container holds CapEff: 00000000a80425fb (the fourteen default capabilities), runs under Seccomp: 2 (filter active), and sees a /dev with a dozen pseudo-devices. The privileged container holds CapEff: 000001ffffffffff, every one of the 41 capabilities this kernel defines (cap_last_cap is 40), runs under Seccomp: 0 (no filter), and its /dev now contains the host's block devices, here sda and vda with their partitions. The AppArmor profile is gone too, and /sys is mounted read-write:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-priv --privileged alpine:3.22 sleep 300 >/dev/null sudo docker inspect -f "apparmor=[{{.AppArmorProfile}}] securityopt={{.HostConfig.SecurityOpt}}" lab-priv sudo docker rm -f lab-priv >/dev/null
apparmor=[unconfined] securityopt=[label=disable]
$ sudo docker run --rm --privileged alpine:3.22 sh -c "grep sysfs /proc/mounts | head -1"
sysfs /sys sysfs rw,nosuid,nodev,noexec,relatime 0 0

apparmor=[unconfined] and a writable sysfs. securityopt=[label=disable] is the SELinux labelling switch that --privileged also sets; it has no effect on this AppArmor host. An attacker in a privileged container mounts one of those host disks, or writes to /sys, and is on the host, so treat --privileged as a host root shell. The fix is never the whole keyring: grant the one capability and the one device the workload needs. A container that manages network interfaces needs NET_ADMIN, not privileged:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --cap-drop ALL --cap-add NET_ADMIN alpine:3.22 ip link add dummy0 type dummy echo "NET_ADMIN alone, no --privileged: exit $?"
NET_ADMIN alone, no --privileged: exit 0

NET_ADMIN alone, with everything else dropped, creates an interface in the container's own network namespace and nothing more. For a workload that needs a device node, add it with --device /dev/<name> rather than exposing all of /dev. "Capability and kernel escapes" goes through which capabilities are escape-grade and why the default seccomp profile still matters.

Host bind mounts: read-only is not a detail

A bind mount is a window from the container into the host filesystem, and the damage scales with what it frames. -v /:/host is the obvious one. Make a file that only root can read on the host, then mount the whole host into an otherwise ordinary container:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ echo "lab marker: host-only, mode 600" | sudo tee /root/lab-marker.txt >/dev/null sudo chmod 600 /root/lab-marker.txt sudo ls -l /root/lab-marker.txt
-rw------- 1 root root 32 Oct 8 03:03 /root/lab-marker.txt

No privileged flag, no socket, no added capability. The container reads the marker and appends a line to it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm -v /:/host alpine:3.22 sh -c "cat /host/root/lab-marker.txt; echo \"written by a container\" >> /host/root/lab-marker.txt"
lab marker: host-only, mode 600
ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo ls -l /root/lab-marker.txt; sudo cat /root/lab-marker.txt
-rw------- 1 root root 55 Oct 8 03:03 /root/lab-marker.txt lab marker: host-only, mode 600 written by a container

The file is still owned by root with mode 600, and the container both read it and changed it. In a real incident the same mount reads /root/.ssh, drops a file into a directory the host later executes as root (a cron drop-in, a systemd unit), or reads the host's credentials outright. Mounting /var/run is the same story, because the Docker socket lives there. The one thing that does limit it is :ro:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm -v /:/host:ro alpine:3.22 sh -c "echo blocked >> /host/root/lab-marker.txt"
sh: can't create /host/root/lab-marker.txt: Read-only file system

With a read-only bind the write fails with "Read-only file system". That still leaves every host secret readable, so :ro reduces a host bind mount, it does not make one safe. The real fix is to mount only the directory the workload needs, read-only:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-hostro -v /:/host:ro alpine:3.22 sleep 300 >/dev/null sudo docker ps -q | xargs sudo docker inspect --format "{{.Name}}{{range .Mounts}} {{.Source}}->{{.Destination}}({{if .RW}}rw{{else}}ro{{end}}){{end}}"
/lab-hostro /->/host(ro)
$ sudo mkdir -p /srv/lab-app sudo docker run --rm -v /srv/lab-app:/data:ro alpine:3.22 sh -c "echo mounted one dir read-only; ls -ld /data"
mounted one dir read-only drwxr-xr-x 2 root root 4096 Oct 7 21:33 /data

The sweep lists every container's mounts with source, destination and whether each is writable, which is the finding you look for in review: a broad host path, especially a writable one. Here it is lab-hostro with all of / mounted read-only. The scoped version mounts a single application directory read-only, so a compromise of that container reaches that one directory and no further.

A mounted docker.sock: detect it, and treat it as root

A container with /var/run/docker.sock mounted can drive the root daemon, including starting a second container with the host root inside. "The Docker socket and daemon hardening" proves that end to end and covers the mitigations; here the job is to find it, because CI runners, monitoring agents and reverse proxies ask for the socket often, usually with :ro. Sweep for it, then test what :ro buys:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-sockuser -v /var/run/docker.sock:/var/run/docker.sock:ro docker:29-cli sleep 300 >/dev/null sudo docker ps -q | xargs sudo docker inspect --format "{{.Name}}{{range .Mounts}} {{.Source}}:{{.Mode}}{{end}}" | grep docker.sock
/lab-sockuser /var/run/docker.sock:ro
ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker exec lab-sockuser docker ps --format "{{.Names}} {{.Image}}"
lab-sockuser docker:29-cli lab-hostro alpine:3.22

The sweep finds the socket mounted into lab-sockuser, and the Docker CLI inside that container still lists the host's containers through the API. Sending API requests over a socket is not a file write, so :ro does not limit the API, and a container holding the socket, read-only or not, holds host root. The substitutes (a socket proxy scoped to the endpoints a tool needs, a rootless or remote builder for CI) are in "The Docker socket and daemon hardening".

A workload asks for one of the big three. What do you grant?
Not the whole host
grant the scoped substitute
--privileged
The one capability + the one device
e.g. --cap-add NET_ADMIN --device /dev/net/tun
the Docker socket
A filtering proxy, or a rootless/remote builder
never the raw socket, :ro or not
a host bind mount
The single directory it needs, read-only
-v /srv/app/data:/data:ro
In a manifest review the default answer to the big three is no. Each has a narrower setting that does the real job. Deny them by default in CI or an admission policy, and time-box any reviewed exception.

Run the same sweep after every change window: confirm no container went privileged, picked up a broad host mount or gained the socket, and record the reading in the change ticket. "The escape mindset" is the checklist these three sit inside.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker rm -f lab-sockuser lab-hostro >/dev/null sudo rm -f /root/lab-marker.txt sudo rm -rf /srv/lab-app echo cleaned
cleaned
Quick check
01A Compose service runs as root with --cap-drop ALL and no --privileged, and it bind-mounts the host's /etc read-write. A reviewer calls it low risk because every capability is dropped. What is the real exposure?
Incorrect — Container root is host UID 0, and /etc is owned by UID 0, so ordinary owner permissions already allow the reads and writes; no capability is needed. (A non-root UID would get only what host permissions give that UID.)
Incorrect — That understates it; the real prizes are host credentials and host code execution, not disk exhaustion.
Incorrect — A writable host bind alone is enough; it needs no socket and no privileged flag.
Correct — As UID 0 it owns the files, so it reads /etc/shadow and writes into /etc/cron.d with no capability; a host-executed path turns the mount into host code execution.
02docker inspect shows a container started with --privileged. Which description of what that single flag did is correct?
Correct — The lab measured CapEff 0x1ffffffffff, Seccomp 0, apparmor unconfined and host block devices in /dev.
Incorrect — The lab measured Seccomp 0 and apparmor unconfined under privileged; both are off.
Incorrect — Privileged exposes host devices but mounts nothing for you; you still mount a disk yourself.
Incorrect — The capability set became the full 41 and the AppArmor profile was dropped; privileged turns off all of it, not some.
03A monitoring tool mounts /var/run/docker.sock:/var/run/docker.sock:ro. A teammate says the :ro makes it safe to read metadata only. What should you check and conclude?
Incorrect — The in-container UID does not limit the API, and :ro does not either.
Correct — The lab showed a :ro socket container listing containers; :ro controls the file, not the protocol spoken over it.
Incorrect — Create calls are API requests over the socket, not writes to the socket file, so file permissions do not stop them.
Incorrect — A read-only socket mount can issue every API request, including starting a privileged container.

Try this

Work through “A mounted docker.sock: detect it, and treat it as root” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from the big three: privileged, docker.sock and host mounts, keep “A mounted docker.sock: detect it, and treat it as root”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related