The escape mindset
How a container breakout happens and the control that stops each step.
In January 2024 a set of runc bugs called Leaky Vessels was published. One of them, CVE-2024-21626, let a container set its working directory to a file descriptor runc had left open on the host filesystem, so code in the container could read and write host files as root. No kernel exploit, no clever syscall chain. The runtime handed the container a path out. That is what a container escape is: code that was supposed to stay inside a container gets control of the host underneath it. This lesson is the model you use to reason about every escape, and a short lab that reads, on a default container, the controls that keep that code boxed in.
A container is a process with restrictions around it
A container is an ordinary Linux process on the host. What makes it a container is a set of restrictions the kernel wraps around that process, and each one is a separate boundary an escape has to get past:
An escape is what happens when one of these is missing, loosened or switched off, or when the machinery that applies them has a bug, as Leaky Vessels did. The attacker's goal is always the same: get off the container and onto the host, where the other workloads, the secrets and the daemon's own credentials live. So the first thing an attacker does on landing is take inventory, and that inventory is also your audit. Run the same reads yourself and every answer an attacker would want becomes a finding you can close.
The defaults, measured
This lesson runs on the main lab VM (secopslog-docker). Pull the two images, start a stock container and read the four controls straight from the kernel's view of its first process:
CapEff: 00000000a80425fb is the effective capability set as a bitmask; decoded, it is the fourteen capabilities Docker grants by default, and "Capabilities, cap-drop and no-new-privileges" spells them out. None of the fourteen hands over the host on its own the way SYS_ADMIN, SYS_MODULE, SYS_PTRACE or DAC_READ_SEARCH can, and those are held back. The fourteen are still not harmless. NET_RAW was the only precondition for CVE-2020-14386, a kernel memory-corruption bug in raw packet sockets that was published as a container escape, and CHOWN, DAC_OVERRIDE and FOWNER turn any writable host mount into host writes as root. That is why the hardened container below drops all of them. Seccomp: 2 means seccomp is in filter mode; with no security option passed, that filter is Docker's default profile. apparmor=docker-default is the AppArmor profile every container gets unless you override it. NoNewPrivs: 0 is the one control not on by default: no-new-privileges is off until you pass --security-opt no-new-privileges, and the hardened container below turns it on.
Recon as audit
The inventory an attacker runs on a fresh foothold is a handful of read-only questions: who am I, what can I do, is the host reachable from here. Run it against your own container and read the answers as a defender:
uid=0 means root inside the container, which is the host's root unless a user namespace remaps it ("What root in a container really is"). CapEff is the default set, so none of the capabilities that open a direct path to the host is there. The socket probe prints a "No such file" error, so there is no /var/run/docker.sock mounted in. And there is no host bind under /host. Each of those absences closes a whole class of escape. The commands only read; they change nothing, which is why this is a check you can run on a schedule.
Answer the same questions from the host
You do not need a shell inside a container to learn any of this. docker inspect prints the settings that decide every branch, from outside, which makes the whole checklist scriptable into a nightly sweep:
priv=false, no added capabilities, the default AppArmor profile, no shared PID or network namespace, and 0 mounts. A container started with --privileged, --pid=host, a mounted socket or -v /:/host would show it here instead, and "The big three: privileged, docker.sock and host mounts" is the lesson on finding and refusing those. Reading it from the host matters because a compromised container controls the tools and files you would use inside it: its grep, cat and shell can be replaced, so what it reports about itself is untrusted. The daemon's record of how it was started, and the host's own view of /proc/<pid>, are not under its control. "Container forensics and incident response" uses the same split.
Every escape needs a door already open
The useful thing for a defender is that an escape cannot pick a container at random. Each class needs one specific precondition to be true, and when it is not, that whole class is off the table. So you can walk the preconditions like a checklist and say precisely what a given container could and could not do if it were popped tomorrow.
Stack the controls
No single control is trusted to hold alone, so you stack them and make each layer cost the attacker another step. Run as a non-root user and "root inside" buys nothing ("Run as non-root"). Drop every capability and the capability tricks have nothing to hold ("Capabilities, cap-drop and no-new-privileges"). Keep the default seccomp profile and a large part of the kernel's attack surface is refused before the kernel sees it. Mount the root filesystem read-only and there is nowhere to drop a payload ("Read-only root filesystem"). Here is the same container with the stack on:
uid=10001, CapEff: 0000000000000000, and NoNewPrivs: 1. The effective set is empty, so there is nothing to escalate with, and no-new-privileges means no setuid binary in the image can raise it later. docker inspect confirms the same from the daemon's side. Clean up:
--cap-drop ALL and a read-only root filesystem, and it also has /var/run/docker.sock bind-mounted in. How much does that hardening protect the host?uid=0, the default CapEff, no docker.sock, no host bind mount, Seccomp: 2, and the host kernel is fully patched. Going by the prerequisite checklist, what is the realistic route to the host?CAP_SYS_ADMIN, so mounting host disks is not available.docker inspect from the host rather than trusting /proc inside the container?grep and cat can be replaced, so their output is untrusted; the daemon's record of how the container started is not under its control, and the checklist is scriptable from it./proc/self/status is readable in a default container; the lab reads it directly./proc/self/status needs no capability; the reason to use the host is trust, not permission.Try this
Work through “Stack the controls” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from the escape mindset, keep “Stack the controls”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.