Linux capabilities: dropping root the right way
Split root into fine-grained capabilities, drop everything you don't need, and stop running containers as full root.
Root on Linux was a single bit for a long time: uid 0 could do everything. Capabilities split "everything" into about forty named privileges, so a process can bind port 443 (CAP_NET_BIND_SERVICE) without also being able to load a kernel module (CAP_SYS_MODULE) or read every file (CAP_DAC_READ_SEARCH). Containers made this practical: a container runtime starts every process with a small set and lets you remove all of it. The part that takes a moment is that a process does not have a capability set; it has five, and which one matters depends on what you are trying to do.
The five sets
| Set | Meaning | Where it comes from |
|---|---|---|
Effective (CapEff) | what the kernel checks right now | permitted, raised by the program or by file capabilities marked +e |
Permitted (CapPrm) | what the process may raise into effective | inherited at exec through file capabilities or ambient |
Inheritable (CapInh) | what may pass through an exec, if the file also allows it | rarely useful on its own; needs matching file inheritable bits |
Bounding (CapBnd) | the ceiling: nothing outside it can ever be gained | the runtime or systemd sets it; cap_drop: ALL shrinks it |
Ambient (CapAmb) | what a non-root process keeps across exec without file capabilities | systemd AmbientCapabilities=, or the runtime for non-root containers |
grep Cap /proc/$(pgrep -x nginx | head -1)/statusCapInh: 0000000000000000CapPrm: 0000000000000400CapEff: 0000000000000400CapBnd: 00000000a80425fbCapAmb: 0000000000000000effective is one capability; the bounding set is still the container default, so this process could gain moredocker run --rm IMG grep CapEff /proc/self/status # a default container, as rootCapEff: 00000000a80425fbdocker run --rm IMG capsh --decode=00000000a80425fb0x00000000a80425fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcapfourteen capabilities the default container starts with — including cap_net_raw, which restricted profiles dropdocker run --rm --cap-drop ALL --cap-add NET_BIND_SERVICE IMG grep CapEff /proc/self/statusCapEff: 0000000000000400docker run --rm IMG capsh --decode=00000000000004000x0000000000000400=cap_net_bind_servicedrop all, add one: the effective set is a single capabilityThree ways to grant one capability, and what each costs
The example everyone meets is a service that must listen on port 80 or 443. Running it as root solves that and hands it every other privilege too. The three narrow answers differ in where the grant lives. File capabilities (setcap cap_net_bind_service+ep /usr/local/bin/api) attach the grant to the binary, so any user running it gets it; they survive cp -p but not a rebuild of the image, and getcap -r / in CI is how you notice one that should not be there. The container runtime (cap_drop: [ALL] then cap_add: [NET_BIND_SERVICE]) attaches the grant to the container; on a non-root container the runtime puts it in the ambient set, so it works without file capabilities. systemd (AmbientCapabilities=CAP_NET_BIND_SERVICE with CapabilityBoundingSet= limiting the ceiling) attaches it to the unit for a service that is not in a container at all.
services:api:image: registry.acme.dev/shop/api:1.4.2user: "65532:65532"cap_drop: [ALL]cap_add: [NET_BIND_SERVICE] # only if it must bind below 1024; listening on 8080 needs nothingsecurity_opt:- no-new-privileges:true # setuid binaries and file capabilities cannot raise privileges
[Service]User=apiCapabilityBoundingSet=CAP_NET_BIND_SERVICEAmbientCapabilities=CAP_NET_BIND_SERVICENoNewPrivileges=yes
The fourth answer is often the right one: listen on 8080 and let the load balancer or ingress own port 443. It needs no capability at all, which is the state most application containers should be in, and it is what cap_drop: [ALL] with no cap_add produces.
The capabilities that are root by another name
From capabilities(7): grants that amount to full privilege
| Capability | What it allows | Why it is equivalent to root |
|---|---|---|
CAP_SYS_ADMIN | mount, namespaces, quotas, dozens of operations | the manual page calls it the new root; a mount is a host filesystem |
CAP_DAC_OVERRIDE | bypass file permission checks | read /etc/shadow, write any file |
CAP_SYS_PTRACE | trace and modify other processes | inject into any process, read its memory and credentials |
CAP_SYS_MODULE | load kernel modules | arbitrary kernel code |
CAP_NET_RAW | raw and packet sockets | ARP spoofing on the node network, which is why restricted profiles drop it |
CAP_SETUID, CAP_SETGID | change to any uid or gid | become root, if root exists in the namespace |
--privileged grants all of them plus device access; a Kubernetes Pod with privileged: true is the same grant on every node it lands on. When an application fails with Operation not permitted after cap_drop: [ALL], the temptation is to add back the one that makes the error go away, and the one that always does is CAP_SYS_ADMIN. The better question is which specific operation failed: strace -f -e trace=%network,%process on the failing call usually names a single capability, and often the answer is a file permission or a port number rather than any capability.
docker run --rm --cap-drop ALL IMG python3 /bind.py 443 # standard <1024 range restored, see the notebind :443: Permission denied (errno 13)exit 1docker run --rm --cap-drop ALL --cap-add NET_BIND_SERVICE IMG python3 /bind.py 443listening on :443docker run --rm --cap-drop ALL IMG python3 /bind.py 8080listening on :8080one capability, named by the failing bind; not SYS_ADMIN, not --privileged. Port 8080 needs noneCapabilities are one of four layers on a process, next to seccomp, the LSM (AppArmor or SELinux) and user namespaces, and each catches something the others miss. The systemd side, where CapabilityBoundingSet= sits with the filesystem and syscall sandboxing directives, is in systemd service hardening; the container side, where a dropped capability set is one of the doors an escape needs, is in container escape paths.