The escape mindset

How a container breakout happens and the control that stops each step.

Advanced12 min · lesson 21 of 24

In January 2024 a set of runc bugs called Leaky Vessels was published. One of them, CVE-2024-21626, let a container set its working directory to a file descriptor runc had left open on the host filesystem, so code in the container could read and write host files as root. No kernel exploit, no clever syscall chain. The runtime handed the container a path out. That is what a container escape is: code that was supposed to stay inside a container gets control of the host underneath it. This lesson is the model you use to reason about every escape, and a short lab that reads, on a default container, the controls that keep that code boxed in.

A container is a process with restrictions around it

A container is an ordinary Linux process on the host. What makes it a container is a set of restrictions the kernel wraps around that process, and each one is a separate boundary an escape has to get past:

An escape is what happens when one of these is missing, loosened or switched off, or when the machinery that applies them has a bug, as Leaky Vessels did. The attacker's goal is always the same: get off the container and onto the host, where the other workloads, the secrets and the daemon's own credentials live. So the first thing an attacker does on landing is take inventory, and that inventory is also your audit. Run the same reads yourself and every answer an attacker would want becomes a finding you can close.

The defaults, measured

This lesson runs on the main lab VM (secopslog-docker). Pull the two images, start a stock container and read the four controls straight from the kernel's view of its first process:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker pull -q nginx:1.30-alpine docker pull -q alpine:3.22
docker.io/library/nginx:1.30-alpine docker.io/library/alpine:3.22
$ docker run -d --name lab-web nginx:1.30-alpine >/dev/null docker exec lab-web grep -E "^(CapEff|Seccomp|NoNewPrivs):" /proc/1/status docker inspect -f "apparmor={{.AppArmorProfile}}" lab-web
CapEff: 00000000a80425fb NoNewPrivs: 0 Seccomp: 2 apparmor=docker-default

CapEff: 00000000a80425fb is the effective capability set as a bitmask; decoded, it is the fourteen capabilities Docker grants by default, and "Capabilities, cap-drop and no-new-privileges" spells them out. None of the fourteen hands over the host on its own the way SYS_ADMIN, SYS_MODULE, SYS_PTRACE or DAC_READ_SEARCH can, and those are held back. The fourteen are still not harmless. NET_RAW was the only precondition for CVE-2020-14386, a kernel memory-corruption bug in raw packet sockets that was published as a container escape, and CHOWN, DAC_OVERRIDE and FOWNER turn any writable host mount into host writes as root. That is why the hardened container below drops all of them. Seccomp: 2 means seccomp is in filter mode; with no security option passed, that filter is Docker's default profile. apparmor=docker-default is the AppArmor profile every container gets unless you override it. NoNewPrivs: 0 is the one control not on by default: no-new-privileges is off until you pass --security-opt no-new-privileges, and the hardened container below turns it on.

Recon as audit

The inventory an attacker runs on a fresh foothold is a handful of read-only questions: who am I, what can I do, is the host reachable from here. Run it against your own container and read the answers as a defender:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker exec lab-web sh -c 'id; grep CapEff /proc/self/status; ls /var/run/docker.sock 2>&1; mount | grep " /host " || echo "no host mount"'
uid=0(root) gid=0(root) groups=0(root),0(root),1(bin),2(daemon),3(sys),4(adm),6(disk),10(wheel),11(floppy),20(dialout),26(tape),27(video) CapEff: 00000000a80425fb ls: /var/run/docker.sock: No such file or directory no host mount

uid=0 means root inside the container, which is the host's root unless a user namespace remaps it ("What root in a container really is"). CapEff is the default set, so none of the capabilities that open a direct path to the host is there. The socket probe prints a "No such file" error, so there is no /var/run/docker.sock mounted in. And there is no host bind under /host. Each of those absences closes a whole class of escape. The commands only read; they change nothing, which is why this is a check you can run on a schedule.

Answer the same questions from the host

You do not need a shell inside a container to learn any of this. docker inspect prints the settings that decide every branch, from outside, which makes the whole checklist scriptable into a nightly sweep:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker inspect lab-web --format "priv={{.HostConfig.Privileged}} caps={{.HostConfig.CapAdd}} apparmor={{.AppArmorProfile}} pid={{.HostConfig.PidMode}} net={{.HostConfig.NetworkMode}}"
priv=false caps=[] apparmor=docker-default pid= net=bridge
$ docker inspect lab-web --format "{{len .Mounts}} mounts{{range .Mounts}}{{println}}{{.Source}} -> {{.Destination}} ({{if .RW}}rw{{else}}ro{{end}}){{end}}"
0 mounts

priv=false, no added capabilities, the default AppArmor profile, no shared PID or network namespace, and 0 mounts. A container started with --privileged, --pid=host, a mounted socket or -v /:/host would show it here instead, and "The big three: privileged, docker.sock and host mounts" is the lesson on finding and refusing those. Reading it from the host matters because a compromised container controls the tools and files you would use inside it: its grep, cat and shell can be replaced, so what it reports about itself is untrusted. The daemon's record of how it was started, and the host's own view of /proc/<pid>, are not under its control. "Container forensics and incident response" uses the same split.

Every escape needs a door already open

The useful thing for a defender is that an escape cannot pick a container at random. Each class needs one specific precondition to be true, and when it is not, that whole class is off the table. So you can walk the preconditions like a checklist and say precisely what a given container could and could not do if it were popped tomorrow.

A container was compromised. Can the attacker reach the host?
Walk the preconditions in order
each one is a setting you control
--privileged set
Seccomp off, AppArmor off, full caps, host devices
the whole confinement stack is gone
docker.sock mounted
Full control of the root daemon
the container asks dockerd for the host
host path bind-mounted
Read and write host files directly
scales with what the mount exposes
escape-grade capability added
A capability-driven path opens
e.g. SYS_ADMIN or SYS_MODULE
none of the above
Only a live kernel or runtime bug is left
the expensive path, like Leaky Vessels
Close the top four doors with configuration and the attacker is left with the last branch, a vulnerability in the kernel or the runtime that no flag prevents.

Stack the controls

No single control is trusted to hold alone, so you stack them and make each layer cost the attacker another step. Run as a non-root user and "root inside" buys nothing ("Run as non-root"). Drop every capability and the capability tricks have nothing to hold ("Capabilities, cap-drop and no-new-privileges"). Keep the default seccomp profile and a large part of the kernel's attack surface is refused before the kernel sees it. Mount the root filesystem read-only and there is nowhere to drop a payload ("Read-only root filesystem"). Here is the same container with the stack on:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-hard \ --user 10001:10001 --cap-drop ALL --security-opt no-new-privileges \ --read-only --tmpfs /tmp alpine:3.22 sleep 300 >/dev/null docker exec lab-hard sh -c "id; grep -E \"^(CapEff|NoNewPrivs):\" /proc/1/status"
uid=10001 gid=10001 groups=10001 CapEff: 0000000000000000 NoNewPrivs: 1
ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker inspect lab-hard --format "user={{.Config.User}} caps_dropped={{.HostConfig.CapDrop}} opts={{.HostConfig.SecurityOpt}} readonly={{.HostConfig.ReadonlyRootfs}}"
user=10001:10001 caps_dropped=[ALL] opts=[no-new-privileges] readonly=true

uid=10001, CapEff: 0000000000000000, and NoNewPrivs: 1. The effective set is empty, so there is nothing to escalate with, and no-new-privileges means no setuid binary in the image can raise it later. docker inspect confirms the same from the daemon's side. Clean up:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-web lab-hard >/dev/null; echo removed
removed
Watch out
Leaky Vessels is the uncomfortable case. A container with the default seccomp profile, every capability dropped and a read-only root filesystem would still have been broken by CVE-2024-21626, because the flaw was in runc's setup code, which runs before the container's restrictions are applied. seccomp and dropped capabilities guard the process after it starts; they have nothing to say about a bug in the machinery that starts it. The control that did cut the damage was running as a non-root user, because the leaked access then landed as an unprivileged account rather than host root. Patch runc, containerd and the engine as a front-line control, and track their CVEs as seriously as you track the ones in your base images. "Capability and kernel escapes" picks up the shared-kernel half of this.
Quick check
01A container runs as a non-root user with --cap-drop ALL and a read-only root filesystem, and it also has /var/run/docker.sock bind-mounted in. How much does that hardening protect the host?
Incorrect — The user and capability hardening is real, but none of it touches the socket. Talking to the Docker API needs no capabilities and no root inside the container.
Incorrect — It is not a file read. The API can start a privileged container with the host root mounted, which is full host root.
Correct — The socket is a line to a root daemon; the request runs in the daemon's context, so the container's own restrictions never apply.
Incorrect — The API grants full control over the socket regardless of the in-container UID.
02You run the recon-as-audit reads on a container: uid=0, the default CapEff, no docker.sock, no host bind mount, Seccomp: 2, and the host kernel is fully patched. Going by the prerequisite checklist, what is the realistic route to the host?
Incorrect — Container root is host UID 0, but namespaces, the default capability set, seccomp and AppArmor still stand between it and the host.
Incorrect — The default set leaves out CAP_SYS_ADMIN, so mounting host disks is not available.
Incorrect — No host path is mounted, so the container's filesystem view holds no host files.
Correct — With privileged, the socket, host mounts and the dangerous capabilities all absent, only the expensive path remains.
03Why does the lesson read a container's privilege settings with docker inspect from the host rather than trusting /proc inside the container?
Incorrect — Both report the capability sets; the reason is trust, not detail.
Correct — In-container grep and cat can be replaced, so their output is untrusted; the daemon's record of how the container started is not under its control, and the checklist is scriptable from it.
Incorrect — /proc/self/status is readable in a default container; the lab reads it directly.
Incorrect — Reading your own /proc/self/status needs no capability; the reason to use the host is trust, not permission.

Try this

Work through “Stack the controls” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from the escape mindset, keep “Stack the controls”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related