How container escapes happen — and how to prevent them
Privileged flags, host mounts, and leaked sockets are the usual doors. Close them before an attacker finds them.
A container is a process with a restricted view of the host: its own PID, mount and network namespaces, a cgroup, a capability set, a seccomp filter. An escape is any way of widening that view back to the host's, and in practice the widening is almost never a kernel exploit. It is a flag somebody set to make a build work, a socket somebody mounted so a job could run docker build, a hostPath somebody added for a log agent. The doors are known, they are few, and each has a control that costs nothing at runtime.
Each door is a spec field. Admission (Pod Security or a policy engine) refuses the field; the runtime settings limit what a process can do if the field slips through; detection watches for the syscalls an escape needs.
Five doors, what each gives an attacker, and what closes it
Configuration escapes
| Door | What it gives a process inside | What closes it |
|---|---|---|
privileged: true | every device, every capability, no seccomp: mount /dev/sda1 /mnt and the node filesystem is yours | Pod Security baseline rejects it; the DaemonSets that need it live in a privileged namespace with a reason |
/var/run/docker.sock or the containerd socket mounted | start a sibling container with --privileged -v /:/host; root on the node by API | no socket mounts in application or CI namespaces; daemonless builds with Buildah or BuildKit rootless |
hostPath of /, /etc, /var/lib/kubelet, /proc | read kubelet credentials, write cron and SSH keys, read every other pod’s volumes | baseline rejects hostPath; agents that need a path get exactly that path, read-only |
CAP_SYS_ADMIN, CAP_SYS_PTRACE, CAP_NET_RAW added | mount, ptrace other processes, forge packets: the ingredients of most namespace breakouts | restricted requires drop: ["ALL"]; add back only NET_BIND_SERVICE |
| root inside plus a writable host mount | files written as uid 0 are root-owned on the node | runAsNonRoot: true with a numeric uid; user namespaces (hostUsers: false) so uid 0 inside is unprivileged outside |
The first two rows are the same door from two sides. A privileged container can mount the node's disk; a container with the Docker socket can ask the daemon to start a privileged container for it. CI runners, log shippers and monitoring agents are where both appear, usually with a comment saying it was temporary. The daemonless builders exist precisely so a build job needs neither: GitLab runner hardening covers the runner side, and kaniko, the tool most older guides recommend here, is no longer maintained.
Find the doors that are open today
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; .securityContext.privileged == true)) | "\(.metadata.namespace)/\(.metadata.name) privileged"'kube-system/calico-node-8xk2q privilegedci/runner-build-14 privilegedkubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.volumes[]?; .hostPath)) | "\(.metadata.namespace)/\(.metadata.name) " + ([.spec.volumes[] | select(.hostPath) | .hostPath.path] | join(","))'logging/fluent-bit-2ptq7 /var/log,/var/lib/docker/containersci/runner-build-14 /var/run/docker.sockkubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; .securityContext.capabilities.add // [] | index("SYS_ADMIN"))) | "\(.metadata.namespace)/\(.metadata.name) CAP_SYS_ADMIN"'the CNI DaemonSet is expected; the CI runner with a socket and privileged is the findingRun this after every platform upgrade, not once. DaemonSets shipped by monitoring and security vendors are how privileged pods reappear in a cluster that had none, and a Helm chart's default values are where the socket mount comes back.
The pod spec that closes them by default
spec:hostUsers: false # user namespace: uid 0 inside maps to an unprivileged host uid (stable since v1.36)automountServiceAccountToken: falsesecurityContext:runAsNonRoot: truerunAsUser: 10001seccompProfile:type: RuntimeDefaultcontainers:- name: appsecurityContext:allowPrivilegeEscalation: falsereadOnlyRootFilesystem: truecapabilities:drop: ["ALL"]volumeMounts:- { name: tmp, mountPath: /tmp }volumes:- { name: tmp, emptyDir: {} }
hostUsers: false is the newest line and the one that changes the arithmetic of the last table row: with a user namespace, a process that is root inside the container is an unprivileged uid on the node, so a writable host mount no longer yields root-owned files, and a capability granted inside the namespace is real only inside it. It needs a runtime with user namespace support (containerd 2.x, CRI-O) and a filesystem that handles id-mapped mounts, so it is a per-cluster capability to confirm before it goes into a chart's defaults. The rest of the spec is the restricted standard written out, and it is what Pod Security Admission enforces per namespace once the audit-mode list is empty.
Detecting the attempt the policy missed
Policy drifts, exceptions accumulate, and a zero-day in the runtime does not care about securityContext. The syscalls an escape needs are distinctive: a mount from inside a container, setns or unshare into another namespace, a write under /proc/sys or /sys, a shell spawned by a process that should never spawn one. Falco rules that watch for those, routed to a channel a human reads, are how you learn that a door was opened after the inventory said they were all shut.
For workloads where even a kernel bug is an unacceptable risk, the next step is not a stricter spec but a different boundary: gVisor or Kata Containers put a second kernel or a user-space kernel between the process and the host, at a cost in compatibility and performance that only multi-tenant or untrusted-code workloads usually justify.
Go deeper in a courseRuntime securityEscape techniques, seccomp and AppArmor profiles, Falco tuning and response on a compromised node.View course