The big three: privileged, docker.sock and host mounts
The settings that strip isolation, how to detect them, and what to grant instead.
Three settings account for most container-to-host takeovers that do not need a kernel bug: the --privileged flag, a bind-mounted Docker socket, and a host directory mounted into the container. Each is granted on purpose, on a docker run line or in a Compose file, so each can also be refused. This lesson measures what each one removes, shows how to find it in a running container, and gives the narrower setting that does the same job.
secopslog-docker-sec). The lab creates a root-only marker file under /root and shows a container changing it through a host bind mount. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo. Pull the two images the lesson uses first:--privileged: the confinement stack, off
A default container is defined by what it does not get: a small capability set, an active seccomp filter, an AppArmor profile, and a device cgroup that hides the host's disks. --privileged cancels all of it at once. Read a default container and a privileged one side by side:
The default container holds CapEff: 00000000a80425fb (the fourteen default capabilities), runs under Seccomp: 2 (filter active), and sees a /dev with a dozen pseudo-devices. The privileged container holds CapEff: 000001ffffffffff, every one of the 41 capabilities this kernel defines (cap_last_cap is 40), runs under Seccomp: 0 (no filter), and its /dev now contains the host's block devices, here sda and vda with their partitions. The AppArmor profile is gone too, and /sys is mounted read-write:
apparmor=[unconfined] and a writable sysfs. securityopt=[label=disable] is the SELinux labelling switch that --privileged also sets; it has no effect on this AppArmor host. An attacker in a privileged container mounts one of those host disks, or writes to /sys, and is on the host, so treat --privileged as a host root shell. The fix is never the whole keyring: grant the one capability and the one device the workload needs. A container that manages network interfaces needs NET_ADMIN, not privileged:
NET_ADMIN alone, with everything else dropped, creates an interface in the container's own network namespace and nothing more. For a workload that needs a device node, add it with --device /dev/<name> rather than exposing all of /dev. "Capability and kernel escapes" goes through which capabilities are escape-grade and why the default seccomp profile still matters.
Host bind mounts: read-only is not a detail
A bind mount is a window from the container into the host filesystem, and the damage scales with what it frames. -v /:/host is the obvious one. Make a file that only root can read on the host, then mount the whole host into an otherwise ordinary container:
No privileged flag, no socket, no added capability. The container reads the marker and appends a line to it:
The file is still owned by root with mode 600, and the container both read it and changed it. In a real incident the same mount reads /root/.ssh, drops a file into a directory the host later executes as root (a cron drop-in, a systemd unit), or reads the host's credentials outright. Mounting /var/run is the same story, because the Docker socket lives there. The one thing that does limit it is :ro:
With a read-only bind the write fails with "Read-only file system". That still leaves every host secret readable, so :ro reduces a host bind mount, it does not make one safe. The real fix is to mount only the directory the workload needs, read-only:
The sweep lists every container's mounts with source, destination and whether each is writable, which is the finding you look for in review: a broad host path, especially a writable one. Here it is lab-hostro with all of / mounted read-only. The scoped version mounts a single application directory read-only, so a compromise of that container reaches that one directory and no further.
A mounted docker.sock: detect it, and treat it as root
A container with /var/run/docker.sock mounted can drive the root daemon, including starting a second container with the host root inside. "The Docker socket and daemon hardening" proves that end to end and covers the mitigations; here the job is to find it, because CI runners, monitoring agents and reverse proxies ask for the socket often, usually with :ro. Sweep for it, then test what :ro buys:
The sweep finds the socket mounted into lab-sockuser, and the Docker CLI inside that container still lists the host's containers through the API. Sending API requests over a socket is not a file write, so :ro does not limit the API, and a container holding the socket, read-only or not, holds host root. The substitutes (a socket proxy scoped to the endpoints a tool needs, a rootless or remote builder for CI) are in "The Docker socket and daemon hardening".
Run the same sweep after every change window: confirm no container went privileged, picked up a broad host mount or gained the socket, and record the reading in the change ticket. "The escape mindset" is the checklist these three sit inside.
--cap-drop ALL and no --privileged, and it bind-mounts the host's /etc read-write. A reviewer calls it low risk because every capability is dropped. What is the real exposure?/etc is owned by UID 0, so ordinary owner permissions already allow the reads and writes; no capability is needed. (A non-root UID would get only what host permissions give that UID.)/etc/shadow and writes into /etc/cron.d with no capability; a host-executed path turns the mount into host code execution.docker inspect shows a container started with --privileged. Which description of what that single flag did is correct?/var/run/docker.sock:/var/run/docker.sock:ro. A teammate says the :ro makes it safe to read metadata only. What should you check and conclude?:ro does not either.:ro socket container listing containers; :ro controls the file, not the protocol spoken over it.Try this
Work through “A mounted docker.sock: detect it, and treat it as root” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from the big three: privileged, docker.sock and host mounts, keep “A mounted docker.sock: detect it, and treat it as root”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.