Linux capabilities: dropping root the right way

Split root into fine-grained capabilities, drop everything you don't need, and stop running containers as full root.

Dec 30, 2025·Updated ·4 min readIntermediate·By SecOpsLog · command-tested

Root on Linux was a single bit for a long time: uid 0 could do everything. Capabilities split "everything" into about forty named privileges, so a process can bind port 443 (CAP_NET_BIND_SERVICE) without also being able to load a kernel module (CAP_SYS_MODULE) or read every file (CAP_DAC_READ_SEARCH). Containers made this practical: a container runtime starts every process with a small set and lets you remove all of it. The part that takes a moment is that a process does not have a capability set; it has five, and which one matters depends on what you are trying to do.

The five sets

SetMeaningWhere it comes from
Effective (CapEff)what the kernel checks right nowpermitted, raised by the program or by file capabilities marked +e
Permitted (CapPrm)what the process may raise into effectiveinherited at exec through file capabilities or ambient
Inheritable (CapInh)what may pass through an exec, if the file also allows itrarely useful on its own; needs matching file inheritable bits
Bounding (CapBnd)the ceiling: nothing outside it can ever be gainedthe runtime or systemd sets it; cap_drop: ALL shrinks it
Ambient (CapAmb)what a non-root process keeps across exec without file capabilitiessystemd AmbientCapabilities=, or the runtime for non-root containers
bash — representative: the five sets of a running nginx that has only cap_net_bind_service
grep Cap /proc/$(pgrep -x nginx | head -1)/status
CapInh: 0000000000000000
CapPrm: 0000000000000400
CapEff: 0000000000000400
CapBnd: 00000000a80425fb
CapAmb: 0000000000000000
effective is one capability; the bounding set is still the container default, so this process could gain more
bash — observed: the effective set a container actually gets, decoded (Debian 13 + libcap2-bin)observed
docker run --rm IMG grep CapEff /proc/self/status # a default container, as root
CapEff: 00000000a80425fb
docker run --rm IMG capsh --decode=00000000a80425fb
0x00000000a80425fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap
fourteen capabilities the default container starts with — including cap_net_raw, which restricted profiles drop
docker run --rm --cap-drop ALL --cap-add NET_BIND_SERVICE IMG grep CapEff /proc/self/status
CapEff: 0000000000000400
docker run --rm IMG capsh --decode=0000000000000400
0x0000000000000400=cap_net_bind_service
drop all, add one: the effective set is a single capability

Three ways to grant one capability, and what each costs

The example everyone meets is a service that must listen on port 80 or 443. Running it as root solves that and hands it every other privilege too. The three narrow answers differ in where the grant lives. File capabilities (setcap cap_net_bind_service+ep /usr/local/bin/api) attach the grant to the binary, so any user running it gets it; they survive cp -p but not a rebuild of the image, and getcap -r / in CI is how you notice one that should not be there. The container runtime (cap_drop: [ALL] then cap_add: [NET_BIND_SERVICE]) attaches the grant to the container; on a non-root container the runtime puts it in the ambient set, so it works without file capabilities. systemd (AmbientCapabilities=CAP_NET_BIND_SERVICE with CapabilityBoundingSet= limiting the ceiling) attaches it to the unit for a service that is not in a container at all.

compose.yaml
services:
api:
image: registry.acme.dev/shop/api:1.4.2
user: "65532:65532"
cap_drop: [ALL]
cap_add: [NET_BIND_SERVICE] # only if it must bind below 1024; listening on 8080 needs nothing
security_opt:
- no-new-privileges:true # setuid binaries and file capabilities cannot raise privileges
/etc/systemd/system/api.service (excerpt)
[Service]
User=api
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
NoNewPrivileges=yes

The fourth answer is often the right one: listen on 8080 and let the load balancer or ingress own port 443. It needs no capability at all, which is the state most application containers should be in, and it is what cap_drop: [ALL] with no cap_add produces.

The capabilities that are root by another name

From capabilities(7): grants that amount to full privilege

CapabilityWhat it allowsWhy it is equivalent to root
CAP_SYS_ADMINmount, namespaces, quotas, dozens of operationsthe manual page calls it the new root; a mount is a host filesystem
CAP_DAC_OVERRIDEbypass file permission checksread /etc/shadow, write any file
CAP_SYS_PTRACEtrace and modify other processesinject into any process, read its memory and credentials
CAP_SYS_MODULEload kernel modulesarbitrary kernel code
CAP_NET_RAWraw and packet socketsARP spoofing on the node network, which is why restricted profiles drop it
CAP_SETUID, CAP_SETGIDchange to any uid or gidbecome root, if root exists in the namespace

--privileged grants all of them plus device access; a Kubernetes Pod with privileged: true is the same grant on every node it lands on. When an application fails with Operation not permitted after cap_drop: [ALL], the temptation is to add back the one that makes the error go away, and the one that always does is CAP_SYS_ADMIN. The better question is which specific operation failed: strace -f -e trace=%network,%process on the failing call usually names a single capability, and often the answer is a file permission or a port number rather than any capability.

bash — observed: find the missing capability instead of guessing (a program that binds :443)observed
docker run --rm --cap-drop ALL IMG python3 /bind.py 443 # standard <1024 range restored, see the note
bind :443: Permission denied (errno 13)
exit 1
docker run --rm --cap-drop ALL --cap-add NET_BIND_SERVICE IMG python3 /bind.py 443
listening on :443
docker run --rm --cap-drop ALL IMG python3 /bind.py 8080
listening on :8080
one capability, named by the failing bind; not SYS_ADMIN, not --privileged. Port 8080 needs none
The privileged-port line is a kernel setting, not only a capability
The <1024 rule is `net.ipv4.ip_unprivileged_port_start`, a namespaced sysctl. Its default is 1024, but Docker Desktop and OrbStack set it to 0 in their VM, so a container there binds :443 with no capability at all — the recorded runs saw exactly that on OrbStack, and the bind test above sets the sysctl back to 1024 to show the capability boundary the article is about. On a cluster where the node keeps the default, the capability is what the bind needs; confirm your own nodes with `sysctl net.ipv4.ip_unprivileged_port_start` rather than assuming, and prefer binding above 1024 so the question does not arise.
Inside a user namespace, a capability is only as big as the namespace
With rootless Docker or a Pod running with hostUsers: false, capabilities are held in the container’s user namespace. CAP_SYS_ADMIN there lets the process mount inside its own namespace and does not let it touch host mounts, host devices or other users’ processes. That is the mechanism that makes an escape land unprivileged, and it is also why a tool that needs a capability over a host resource fails in rootless mode even with --cap-add.
What was run for this article
Docker Engine 28.5.2 on linux/arm64 (OrbStack), against a small fixture image (Debian 13-slim pinned by digest, plus libcap2-bin for capsh and python3 for a one-line bind program; build/evidence/linux-capabilities). The observed blocks are copied from that run: a default container's effective set is the 14-capability a80425fb, decoded; --cap-drop ALL --cap-add NET_BIND_SERVICE leaves exactly cap_net_bind_service; with the standard <1024 range restored, binding :443 is refused with no capability and allowed with that one, while :8080 needs none. Seven exit-code and value assertions. The one environment caveat is in the note above: OrbStack's VM disables the privileged-port range, so the fixture sets it back to the kernel default to test the boundary. The nginx five-sets block, file-capability and systemd examples, and the CAP_SYS_ADMIN discussion are representative and documentation-backed.

Capabilities are one of four layers on a process, next to seccomp, the LSM (AppArmor or SELinux) and user namespaces, and each catches something the others miss. The systemd side, where CapabilityBoundingSet= sits with the filesystem and syscall sandboxing directives, is in systemd service hardening; the container side, where a dropped capability set is one of the doors an escape needs, is in container escape paths.

Related posts

Quick reference