Capability and kernel escapes
Why the shared kernel is the boundary and which defaults block the known paths.
Without a user namespace, root in a container is host UID 0 ("What root in a container really is"). What keeps it from acting like root on the host is the set of boundaries around the container process: the capabilities it holds, the seccomp filter over its system calls, and the AppArmor profile over its file and mount operations. A kernel escape is what happens when enough of those boundaries are open at once, or when the one kernel every container shares has a bug. This lesson measures which defaults block the known paths on Docker Engine 29, and it corrects a common misreading of the default seccomp profile.
secopslog-docker-sec). The lab starts containers with escape-grade capabilities and with seccomp and AppArmor turned off to measure what they expose; it loads no kernel module (it names one that does not exist). If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM docker runs through sudo.The default container, and the capabilities it does not have
Start with the baseline and try something that needs an escape-grade capability. rmmod asks the kernel to unload a module through the delete_module system call, which needs CAP_SYS_MODULE:
CapEff: 00000000a80425fb is the fourteen default capabilities ("Capabilities, cap-drop and no-new-privileges" lists them), and CAP_SYS_MODULE is not among them. Seccomp: 2 is the default filter. rmmod fails with "Operation not permitted": the container has neither the capability nor, as the next step shows, the system call. That one error is two separate refusals at once, which is the thing to understand before you hand any capability back.
CAP_SYS_MODULE: the capability is the lock
Add the capability and run the same command:
CapEff is now a80525fb: bit 16, CAP_SYS_MODULE, is set. Seccomp is still 2, the same default filter. And rmmod now fails differently, with "No such file or directory": the delete_module call reached the kernel, and the kernel looked for a module named lab_nope, which does not exist. The syscall was not blocked. Guides often read Docker's list of syscalls the default profile blocks as meaning CAP_SYS_MODULE alone is useless. The profile does not work that way, and has not since at least Engine 17.03: init_module, finit_module and delete_module sit in one rule that is allowed to any container holding CAP_SYS_MODULE. This lab exercised the unload call; the load calls are in the same capability-conditioned rule, so a container with --cap-add SYS_MODULE on the default profile can also load a module. Adding the capability lifts both locks together. The capability is the control; do not grant it, and do not rely on seccomp to save you if you do.
CAP_SYS_ADMIN: the same syscall gate, plus a second layer
CAP_SYS_ADMIN is the broad one; mount is the operation to watch. Add it and look at the sets:
CapEff: 00000000a82425fb, bit 21 set. The default seccomp profile permits mount to a CAP_SYS_ADMIN holder in the same capability-conditioned way it permits the module syscalls. But try the mount on the full default stack:
"Permission denied". seccomp allowed the mount call, and the capability is held, yet the mount is refused, because the AppArmor docker-default profile denies mount operations independently ("AppArmor and SELinux" shows this from the profile side). Turn that layer off and the same command succeeds:
With --security-opt apparmor=unconfined, the CAP_SYS_ADMIN container mounts a tmpfs. So on this host the default stack stops a CAP_SYS_ADMIN mount with two independent layers, the capability gate and AppArmor, and an attacker needs both open. Dropping the capability closes it, and if the capability is granted the LSM profile is still in the way. That does not make CAP_SYS_ADMIN safe to grant: on a host without AppArmor or SELinux, the capability alone opens the mount path.
Where seccomp is still the only lock
For another class of system calls the default profile refuses the call whatever capabilities you add, because no rule in the allow-list names it. keyctl, add_key and request_key (the kernel keyring, which is not namespaced) are examples. Adding capabilities does not reach them; turning the filter off does, with --security-opt seccomp=unconfined or with --privileged, which runs the container with no seccomp filter at all (Seccomp: 0 below). "seccomp: filter system calls" demonstrates that with add_key. Capability gates and the seccomp filter overlap for mount and the module syscalls; for the rest of the blocked set, seccomp is the only lock.
Privileged removes all of it
CapEff: 000001ffffffffff is every one of the 41 capabilities this kernel defines (cap_last_cap is 40), and Seccomp: 0 means no filter at all. --privileged also drops the AppArmor profile. Every lock in this lesson is open at once, which is why "The big three: privileged, docker.sock and host mounts" treats privileged as a host root shell. A kernel bug reached through an ordinary syscall, meanwhile, needs none of these grants: Dirty Pipe (CVE-2022-0847) needs only pipe, splice and write, which every profile allows, so it was reachable from a default container; many io_uring and netfilter privilege-escalation flaws are similar. Every container shares the one host kernel. You cannot patch that kernel from inside a container; you keep the host kernel and runc current, and you shrink the syscall surface so fewer bugs are reachable.
Detection
docker inspect is the truth of how a container was started. Sweep for the grants that matter, on a schedule or in an admission check:
lab-ci holds CAP_SYS_ADMIN, which the default seccomp profile already lets call mount, and has seccomp off, which opens every call that seccomp alone was blocking (keyctl, add_key, request_key). Neither setting is privileged. Its AppArmorProfile is still docker-default, and that profile is now the only thing between it and a mount, so one more flag, or a host without AppArmor, opens the mount path. The sweep therefore reads AppArmorProfile as well as the security options. For the policy, use an allow-list rather than a list of known-bad capabilities: fail any workload in CI or an admission controller that adds a capability outside an approved list (start from none, plus the few that the drop-all workflow in "Capabilities, cap-drop and no-new-privileges" justified), sets seccomp or AppArmor to unconfined, or sets privileged. Then confirm the host is not exposed to a class you care about:
Both lines matter for the note below: cgroup2fs is the cgroup mode, and uname -r is the version you track against your distribution's security advisories (this VM runs 7.0.0-34-generic; yours will differ). Every container in this lesson ran with --rm or was removed in the sweep, so there is nothing to clean up.
CAP_SYS_ADMIN could create a new user and cgroup namespace and set the release_agent file there, and the host would run that program as root when a cgroup emptied. Kernels since 5.17 contain the fix, including this VM's 7.0, which is the main reason it does not apply here; this host also runs cgroup v2 (cgroup2fs above), while some distributions (RHEL 8, Amazon Linux 2) still default to v1. On vulnerable kernels Docker's default seccomp and AppArmor profiles blocked stock containers; only containers with seccomp off and no AppArmor or SELinux confinement, such as --privileged ones, were exposed. It is the clearest example of the pattern in this lesson: a kernel bug that only containers with their defences already turned off could reach.--cap-add SYS_MODULE and no other flags. A guide says seccomp blocks the module syscalls anyway. What actually happens, per the lab?rmmod reaching the kernel (ENOENT) once the capability was added; the profile permits those syscalls to a holder of the capability.rmmod returned "No such file or directory", meaning the call reached the kernel; the capability is the control, and the rule has been conditioned on it since at least 2017.--cap-add SYS_ADMIN container on the full default stack tries to mount a tmpfs and gets "Permission denied", but the same command succeeds once --security-opt apparmor=unconfined is added. What does that show?CAP_SYS_ADMIN set before AppArmor was touched; the capability was present throughout.apparmor=unconfined does not grant privilege; the mount worked because the LSM layer was removed, not because privilege was added./web cap_add=[] securityopt=[] apparmor=docker-default and /ci cap_add=[CAP_SYS_ADMIN] securityopt=[seccomp=unconfined] apparmor=docker-default. Neither is privileged. Which is the real escape risk and why?/ci already holds the capability that lets it call mount; only the AppArmor profile is still refusing the mount.cap_add is the safe default set of fourteen, not the full set./web still runs under the default profile.SYS_ADMIN already passes the default seccomp rule for mount, and seccomp off opens the calls it alone blocked. One more flag, or a host without AppArmor, and the mount path is open.Try this
Work through “Detection” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from capability and kernel escapes, keep “Detection”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.