Capability and kernel escapes

Why the shared kernel is the boundary and which defaults block the known paths.

Advanced14 min · lesson 23 of 24

Without a user namespace, root in a container is host UID 0 ("What root in a container really is"). What keeps it from acting like root on the host is the set of boundaries around the container process: the capabilities it holds, the seccomp filter over its system calls, and the AppArmor profile over its file and mount operations. A kernel escape is what happens when enough of those boundaries are open at once, or when the one kernel every container shares has a bug. This lesson measures which defaults block the known paths on Docker Engine 29, and it corrects a common misreading of the default seccomp profile.

Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab starts containers with escape-grade capabilities and with seccomp and AppArmor turned off to measure what they expose; it loads no kernel module (it names one that does not exist). If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM docker runs through sudo.

The default container, and the capabilities it does not have

Start with the baseline and try something that needs an escape-grade capability. rmmod asks the kernel to unload a module through the delete_module system call, which needs CAP_SYS_MODULE:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker pull -q alpine:3.22
docker.io/library/alpine:3.22
$ sudo docker run --rm alpine:3.22 sh -c "grep -E \"^(CapEff|Seccomp):\" /proc/1/status; rmmod lab_nope 2>&1"
CapEff: 00000000a80425fb Seccomp: 2 rmmod: can't unload module 'lab_nope': Operation not permitted

CapEff: 00000000a80425fb is the fourteen default capabilities ("Capabilities, cap-drop and no-new-privileges" lists them), and CAP_SYS_MODULE is not among them. Seccomp: 2 is the default filter. rmmod fails with "Operation not permitted": the container has neither the capability nor, as the next step shows, the system call. That one error is two separate refusals at once, which is the thing to understand before you hand any capability back.

CAP_SYS_MODULE: the capability is the lock

Add the capability and run the same command:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --cap-add SYS_MODULE alpine:3.22 sh -c "grep -E \"^(CapEff|Seccomp):\" /proc/1/status; rmmod lab_nope 2>&1"
CapEff: 00000000a80525fb Seccomp: 2 rmmod: can't unload module 'lab_nope': No such file or directory

CapEff is now a80525fb: bit 16, CAP_SYS_MODULE, is set. Seccomp is still 2, the same default filter. And rmmod now fails differently, with "No such file or directory": the delete_module call reached the kernel, and the kernel looked for a module named lab_nope, which does not exist. The syscall was not blocked. Guides often read Docker's list of syscalls the default profile blocks as meaning CAP_SYS_MODULE alone is useless. The profile does not work that way, and has not since at least Engine 17.03: init_module, finit_module and delete_module sit in one rule that is allowed to any container holding CAP_SYS_MODULE. This lab exercised the unload call; the load calls are in the same capability-conditioned rule, so a container with --cap-add SYS_MODULE on the default profile can also load a module. Adding the capability lifts both locks together. The capability is the control; do not grant it, and do not rely on seccomp to save you if you do.

CAP_SYS_ADMIN: the same syscall gate, plus a second layer

CAP_SYS_ADMIN is the broad one; mount is the operation to watch. Add it and look at the sets:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --cap-add SYS_ADMIN alpine:3.22 grep -E "^(CapEff|Seccomp):" /proc/1/status
CapEff: 00000000a82425fb Seccomp: 2

CapEff: 00000000a82425fb, bit 21 set. The default seccomp profile permits mount to a CAP_SYS_ADMIN holder in the same capability-conditioned way it permits the module syscalls. But try the mount on the full default stack:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --cap-add SYS_ADMIN alpine:3.22 sh -c "mkdir -p /mnt/x; mount -t tmpfs none /mnt/x 2>&1"
mount: mounting none on /mnt/x failed: Permission denied

"Permission denied". seccomp allowed the mount call, and the capability is held, yet the mount is refused, because the AppArmor docker-default profile denies mount operations independently ("AppArmor and SELinux" shows this from the profile side). Turn that layer off and the same command succeeds:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --cap-add SYS_ADMIN --security-opt apparmor=unconfined alpine:3.22 sh -c "mkdir -p /mnt/x; mount -t tmpfs none /mnt/x && echo mounted; grep -c tmpfs /proc/mounts >/dev/null && echo tmpfs-present"
mounted tmpfs-present

With --security-opt apparmor=unconfined, the CAP_SYS_ADMIN container mounts a tmpfs. So on this host the default stack stops a CAP_SYS_ADMIN mount with two independent layers, the capability gate and AppArmor, and an attacker needs both open. Dropping the capability closes it, and if the capability is granted the LSM profile is still in the way. That does not make CAP_SYS_ADMIN safe to grant: on a host without AppArmor or SELinux, the capability alone opens the mount path.

Where seccomp is still the only lock

For another class of system calls the default profile refuses the call whatever capabilities you add, because no rule in the allow-list names it. keyctl, add_key and request_key (the kernel keyring, which is not namespaced) are examples. Adding capabilities does not reach them; turning the filter off does, with --security-opt seccomp=unconfined or with --privileged, which runs the container with no seccomp filter at all (Seccomp: 0 below). "seccomp: filter system calls" demonstrates that with add_key. Capability gates and the seccomp filter overlap for mount and the module syscalls; for the rest of the blocked set, seccomp is the only lock.

Privileged removes all of it

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --privileged alpine:3.22 grep -E "^(CapEff|Seccomp):" /proc/1/status
CapEff: 000001ffffffffff Seccomp: 0

CapEff: 000001ffffffffff is every one of the 41 capabilities this kernel defines (cap_last_cap is 40), and Seccomp: 0 means no filter at all. --privileged also drops the AppArmor profile. Every lock in this lesson is open at once, which is why "The big three: privileged, docker.sock and host mounts" treats privileged as a host root shell. A kernel bug reached through an ordinary syscall, meanwhile, needs none of these grants: Dirty Pipe (CVE-2022-0847) needs only pipe, splice and write, which every profile allows, so it was reachable from a default container; many io_uring and netfilter privilege-escalation flaws are similar. Every container shares the one host kernel. You cannot patch that kernel from inside a container; you keep the host kernel and runc current, and you shrink the syscall surface so fewer bugs are reachable.

Detection

docker inspect is the truth of how a container was started. Sweep for the grants that matter, on a schedule or in an admission check:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-web alpine:3.22 sleep 300 >/dev/null sudo docker run -d --name lab-ci --cap-add SYS_ADMIN --security-opt seccomp=unconfined alpine:3.22 sleep 300 >/dev/null sudo docker ps -q | xargs sudo docker inspect --format "{{.Name}} cap_add={{.HostConfig.CapAdd}} securityopt={{.HostConfig.SecurityOpt}} apparmor={{.AppArmorProfile}} priv={{.HostConfig.Privileged}}" sudo docker rm -f lab-web lab-ci >/dev/null
/lab-ci cap_add=[CAP_SYS_ADMIN] securityopt=[seccomp=unconfined] apparmor=docker-default priv=false /lab-web cap_add=[] securityopt=[] apparmor=docker-default priv=false

lab-ci holds CAP_SYS_ADMIN, which the default seccomp profile already lets call mount, and has seccomp off, which opens every call that seccomp alone was blocking (keyctl, add_key, request_key). Neither setting is privileged. Its AppArmorProfile is still docker-default, and that profile is now the only thing between it and a mount, so one more flag, or a host without AppArmor, opens the mount path. The sweep therefore reads AppArmorProfile as well as the security options. For the policy, use an allow-list rather than a list of known-bad capabilities: fail any workload in CI or an admission controller that adds a capability outside an approved list (start from none, plus the few that the drop-all workflow in "Capabilities, cap-drop and no-new-privileges" justified), sets seccomp or AppArmor to unconfined, or sets privileged. Then confirm the host is not exposed to a class you care about:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ stat -fc %T /sys/fs/cgroup uname -r
cgroup2fs 7.0.0-34-generic

Both lines matter for the note below: cgroup2fs is the cgroup mode, and uname -r is the version you track against your distribution's security advisories (this VM runs 7.0.0-34-generic; yours will differ). Every container in this lesson ran with --rm or was removed in the sweep, so there is nothing to clean up.

Watch out
CVE-2022-0492 was a missing capability check in cgroup v1: a process without CAP_SYS_ADMIN could create a new user and cgroup namespace and set the release_agent file there, and the host would run that program as root when a cgroup emptied. Kernels since 5.17 contain the fix, including this VM's 7.0, which is the main reason it does not apply here; this host also runs cgroup v2 (cgroup2fs above), while some distributions (RHEL 8, Amazon Linux 2) still default to v1. On vulnerable kernels Docker's default seccomp and AppArmor profiles blocked stock containers; only containers with seccomp off and no AppArmor or SELinux confinement, such as --privileged ones, were exposed. It is the clearest example of the pattern in this lesson: a kernel bug that only containers with their defences already turned off could reach.
Quick check
01On Docker Engine 29 with the default seccomp profile, you start a container with --cap-add SYS_MODULE and no other flags. A guide says seccomp blocks the module syscalls anyway. What actually happens, per the lab?
Incorrect — The lab showed rmmod reaching the kernel (ENOENT) once the capability was added; the profile permits those syscalls to a holder of the capability.
Correct — The lab's rmmod returned "No such file or directory", meaning the call reached the kernel; the capability is the control, and the rule has been conditioned on it since at least 2017.
Incorrect — The lab's CapEff showed bit 16 set; the capability is granted, and the syscall is permitted with it.
Incorrect — The lab names a non-existent module, so nothing loads; the point is only that the syscall was no longer blocked.
02A --cap-add SYS_ADMIN container on the full default stack tries to mount a tmpfs and gets "Permission denied", but the same command succeeds once --security-opt apparmor=unconfined is added. What does that show?
Incorrect — Disabling AppArmor does not touch seccomp; seccomp permitted the mount for the SYS_ADMIN holder.
Incorrect — CapEff showed CAP_SYS_ADMIN set before AppArmor was touched; the capability was present throughout.
Correct — The capability and seccomp both allowed it; AppArmor was the second, independent layer stopping it.
Incorrect — apparmor=unconfined does not grant privilege; the mount worked because the LSM layer was removed, not because privilege was added.
03A fleet sweep shows /web cap_add=[] securityopt=[] apparmor=docker-default and /ci cap_add=[CAP_SYS_ADMIN] securityopt=[seccomp=unconfined] apparmor=docker-default. Neither is privileged. Which is the real escape risk and why?
Incorrect — Privileged is not required. /ci already holds the capability that lets it call mount; only the AppArmor profile is still refusing the mount.
Incorrect — An empty cap_add is the safe default set of fourteen, not the full set.
Incorrect — seccomp is applied per container; /web still runs under the default profile.
Correct — SYS_ADMIN already passes the default seccomp rule for mount, and seccomp off opens the calls it alone blocked. One more flag, or a host without AppArmor, and the mount path is open.

Try this

Work through “Detection” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from capability and kernel escapes, keep “Detection”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related