How container escapes happen — and how to prevent them

Privileged flags, host mounts, and leaked sockets are the usual doors. Close them before an attacker finds them.

Jan 28, 2025·Updated ·5 min readAdvanced·By SecOpsLog · documentation-verified

A container is a process with a restricted view of the host: its own PID, mount and network namespaces, a cgroup, a capability set, a seccomp filter. An escape is any way of widening that view back to the host's, and in practice the widening is almost never a kernel exploit. It is a flag somebody set to make a build work, a socket somebody mounted so a job could run docker build, a hostPath somebody added for a log agent. The doors are known, they are few, and each has a control that costs nothing at runtime.

The doors, and the layers that close them

Each door is a spec field. Admission (Pod Security or a policy engine) refuses the field; the runtime settings limit what a process can do if the field slips through; detection watches for the syscalls an escape needs.

Host kernel (shared) — escape target privileged or CAP_SYS_ADMIN hostPID / hostNetwork door #1 hostPath / /var/run/docker.sock write the host FS door #2 No seccomp or AppArmor/SELinux wide syscall surface door #3 Defenses PSS restricted drop caps · readOnly no host mounts One privileged pod with docker.sock is a host root foothold. Assume breakout is possible — shrink the blast radius with PSS + network policy + runtime detection. shared kernel + misconfig → escape → host close doors before you detect them
privileged — capshostPath — sockDefend — PSS

Five doors, what each gives an attacker, and what closes it

Configuration escapes

DoorWhat it gives a process insideWhat closes it
privileged: trueevery device, every capability, no seccomp: mount /dev/sda1 /mnt and the node filesystem is yoursPod Security baseline rejects it; the DaemonSets that need it live in a privileged namespace with a reason
/var/run/docker.sock or the containerd socket mountedstart a sibling container with --privileged -v /:/host; root on the node by APIno socket mounts in application or CI namespaces; daemonless builds with Buildah or BuildKit rootless
hostPath of /, /etc, /var/lib/kubelet, /procread kubelet credentials, write cron and SSH keys, read every other pod’s volumesbaseline rejects hostPath; agents that need a path get exactly that path, read-only
CAP_SYS_ADMIN, CAP_SYS_PTRACE, CAP_NET_RAW addedmount, ptrace other processes, forge packets: the ingredients of most namespace breakoutsrestricted requires drop: ["ALL"]; add back only NET_BIND_SERVICE
root inside plus a writable host mountfiles written as uid 0 are root-owned on the noderunAsNonRoot: true with a numeric uid; user namespaces (hostUsers: false) so uid 0 inside is unprivileged outside

The first two rows are the same door from two sides. A privileged container can mount the node's disk; a container with the Docker socket can ask the daemon to start a privileged container for it. CI runners, log shippers and monitoring agents are where both appear, usually with a comment saying it was temporary. The daemonless builders exist precisely so a build job needs neither: GitLab runner hardening covers the runner side, and kaniko, the tool most older guides recommend here, is no longer maintained.

Find the doors that are open today

bash — cluster-wide inventory, one door per query
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; .securityContext.privileged == true)) | "\(.metadata.namespace)/\(.metadata.name) privileged"'
kube-system/calico-node-8xk2q privileged
ci/runner-build-14 privileged
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.volumes[]?; .hostPath)) | "\(.metadata.namespace)/\(.metadata.name) " + ([.spec.volumes[] | select(.hostPath) | .hostPath.path] | join(","))'
logging/fluent-bit-2ptq7 /var/log,/var/lib/docker/containers
ci/runner-build-14 /var/run/docker.sock
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; .securityContext.capabilities.add // [] | index("SYS_ADMIN"))) | "\(.metadata.namespace)/\(.metadata.name) CAP_SYS_ADMIN"'
the CNI DaemonSet is expected; the CI runner with a socket and privileged is the finding

Run this after every platform upgrade, not once. DaemonSets shipped by monitoring and security vendors are how privileged pods reappear in a cluster that had none, and a Helm chart's default values are where the socket mount comes back.

The pod spec that closes them by default

deploy-hardened.yaml
spec:
hostUsers: false # user namespace: uid 0 inside maps to an unprivileged host uid (stable since v1.36)
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: app
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
volumeMounts:
- { name: tmp, mountPath: /tmp }
volumes:
- { name: tmp, emptyDir: {} }

hostUsers: false is the newest line and the one that changes the arithmetic of the last table row: with a user namespace, a process that is root inside the container is an unprivileged uid on the node, so a writable host mount no longer yields root-owned files, and a capability granted inside the namespace is real only inside it. It needs a runtime with user namespace support (containerd 2.x, CRI-O) and a filesystem that handles id-mapped mounts, so it is a per-cluster capability to confirm before it goes into a chart's defaults. The rest of the spec is the restricted standard written out, and it is what Pod Security Admission enforces per namespace once the audit-mode list is empty.

The socket is root on the node, with extra steps
Mounting the container runtime’s socket into any container is equivalent to giving that container root on the host, because the daemon will start whatever the socket asks for, including a privileged container with the host filesystem mounted. CI runners and build jobs are the usual owners; the fix is a daemonless builder, not a smaller mount.

Detecting the attempt the policy missed

Policy drifts, exceptions accumulate, and a zero-day in the runtime does not care about securityContext. The syscalls an escape needs are distinctive: a mount from inside a container, setns or unshare into another namespace, a write under /proc/sys or /sys, a shell spawned by a process that should never spawn one. Falco rules that watch for those, routed to a channel a human reads, are how you learn that a door was opened after the inventory said they were all shut.

For workloads where even a kernel bug is an unacceptable risk, the next step is not a stricter spec but a different boundary: gVisor or Kata Containers put a second kernel or a user-space kernel between the process and the host, at a cost in compatibility and performance that only multi-tenant or untrusted-code workloads usually justify.

Go deeper in a courseRuntime securityEscape techniques, seccomp and AppArmor profiles, Falco tuning and response on a compromised node.View course

Related posts

Quick reference