Capabilities, cap-drop and no-new-privileges
The pieces of root a container keeps and how to drop them.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/caps.tar.gz && tar -xzf caps.tar.gz, which creates ~/lab/caps/. SHA-256: e9fa31f68ef4143cbc688dcae9b46d8e1373067fd34290b2498a813f10bf85cbA change request asks for two flags on a service that runs as a non-root user: --cap-add NET_BIND_SERVICE so it can listen on port 80, and --cap-add NET_RAW so its health check can ping a dependency. On Docker 29 with a bridge network neither is needed, and the first would not have worked anyway. Reviewing requests like that needs a precise picture of what Linux capabilities are, which ones a container holds, and why the user ID changes what an added capability does. This lesson builds that picture on the main lab VM and ends with the working pattern: drop everything, add back only what an error message names, and set no-new-privileges.
The lesson files (copy them to ~/lab/caps) contain a test image and a Compose file. The image, lab-caps:1, adds capsh and getpcaps from Alpine's libcap-utils and a small Go program that prints its user IDs and capability sets. A second copy of that program is setuid root. It does nothing else, so it shows what a setuid-root binary in an image is worth without exploiting anything:
# lab-caps:1, the capabilities lesson's test image: capsh/getpcaps from libcap-utils, a tiny Go program that# prints its user IDs and capability sets, and a second copy of it with the setuid bit (owned by root).FROM golang:1.27-alpine@sha256:8a5910f31396cd4d89662f56c68b3ae31d374308270a1c3bd96672ee5ed43414 AS buildWORKDIR /srcCOPY whoami.go .RUN CGO_ENABLED=0 go build -o /out/lab-whoami whoami.goFROM alpine:3.22RUN apk add --no-cache libcap-utilsCOPY --from=build /out/lab-whoami /usr/local/bin/lab-whoamiCOPY --from=build /out/lab-whoami /usr/local/bin/lab-suid-whoamiRUN chmod 4755 /usr/local/bin/lab-suid-whoamiUSER 10001:10001
// lab-whoami prints the IDs and capability sets of the running process.// The lab installs a second copy with the setuid bit as lab-suid-whoami.package mainimport ("bufio""fmt""os""strings")func main() {fmt.Printf("uid=%d euid=%d\n", os.Getuid(), os.Geteuid())f, err := os.Open("/proc/self/status")if err != nil {fmt.Println(err)os.Exit(1)}defer f.Close()s := bufio.NewScanner(f)for s.Scan() {l := s.Text()if strings.HasPrefix(l, "CapPrm") || strings.HasPrefix(l, "CapEff") || strings.HasPrefix(l, "NoNewPrivs") {fmt.Println(l)}}}
Five sets per process
Linux splits the privileges of root into capabilities: CAP_CHOWN to change any file's owner, CAP_NET_ADMIN to change interfaces and routes, CAP_SYS_ADMIN for mounts and a long list of administrative operations, and so on, 41 on this kernel. The kernel checks the specific capability when a process attempts a privileged operation. Each process carries five sets of them:
Build the image in ~/lab/caps, then compare a default root container with what capsh reports:
A root container holds a80425fb in its permitted, effective and bounding sets and nothing in the other two. capsh --print spells it out: the Current line lists 14 capabilities with =ep (effective and permitted), the bounding set is the same 14, and the Current IAB line marks the 27 capabilities missing from the bounding set with !, among them cap_sys_admin, cap_sys_module and cap_sys_ptrace. Securebits ends with no-new-privs=0, which matters below. capsh --decode turns any mask back into names; that is how you read a value copied out of /proc or an alert. Some of the 14, and some notable absentees:
The capabilities that are missing from the default set are the ones that most escape techniques need; "Capability and kernel escapes" goes through them. That includes CVE-2022-0492, a missing capability check in cgroup v1: a process without CAP_SYS_ADMIN in the host's user namespace could create a new user namespace, mount a cgroup v1 hierarchy there and set its release_agent, a program the host then runs as root. Docker's defaults blocked it twice: the default seccomp profile refuses unshare to a container without CAP_SYS_ADMIN, and AppArmor's docker-default refuses the mount. It is cgroup v1 only and not applicable on cgroup v2 hosts like this one, and kernels since 5.17 carry the fix.
Reading capabilities from the host
Inside a compromised container, capsh and every other tool can be replaced with one that prints whatever the attacker wants. The host reads the kernel's values for the same process directly. Take the main PID from docker inspect and ask the host:
getpcaps prints the names and /proc/<pid>/status the masks. For a fleet sweep read CapBnd, not CapEff. The effective set depends on the user (the next section shows a non-root container with CapEff 0 and no flags at all), while the bounding set reflects only the run flags: on this engine a default container's bounding set is 00000000a80425fb whatever its user, so any other value means --cap-add, --cap-drop or --privileged, and the larger values deserve the first look. Other runtimes and engine versions use different defaults, so take the baseline from your own hosts. The HostConfig fields at the end of this lesson give the same answer from the Docker side.
A non-root user gets none of them
As UID 10001 (the image's USER), the permitted and effective sets are empty and only the bounding set holds the default 14. --cap-add SYS_ADMIN changes exactly one thing: the bounding set becomes a82425fb, with bit 21, CAP_SYS_ADMIN, added. When a non-root process starts a program, the kernel computes its new permitted set from the program file's capabilities and from the inheritable and ambient sets, and Docker leaves those empty. So a capability added for a non-root container sits in the ceiling where nothing can use it, unless a program in the image can raise the process into it. That is the point of the next demonstration, and the reason the change request's NET_BIND_SERVICE would not have helped.
setuid, the bounding set and no-new-privileges
lab-suid-whoami has the setuid bit (-rwsr-xr-x) and belongs to root. Run by UID 10001, the plain copy has no capabilities. The setuid copy runs with effective UID 0 and, because a program that becomes root at execve gets the full bounding set, with a82425fb in permitted and effective: the SYS_ADMIN that --cap-add put in the ceiling is now usable. Any setuid-root binary in an image, su, mount, or a forgotten debug helper, is that path.
--cap-drop ALL empties the bounding set, so the setuid copy gets no capabilities, but its effective UID still becomes 0. Changing UID through a setuid bit needs no capability, and a process with UID 0 owns every root-owned file it can reach, as the section after next shows. --security-opt no-new-privileges closes the path itself. It sets the kernel's no_new_privs flag on the container's first process; children inherit it and nothing can clear it. With the flag set, execve ignores setuid and setgid bits and does not grant file capabilities, so the setuid copy stays at UID 10001 with nothing. Ordinary programs are unaffected.
Set no-new-privileges on every container. It can also be a daemon-wide default ("no-new-privileges": true in daemon.json), which "The Docker socket and daemon hardening" covers. Before you make it the default, find the images that depend on gaining privileges at exec: sudo or su, other setuid helpers, and binaries with file capabilities. find / -xdev -perm -4000 and getcap -r / inside each image list them. Those images fail under the flag and need fixing first.
ping and ports below 1024 on Docker
Two capabilities in the default set exist mostly because of old habits: NET_RAW "for ping" and NET_BIND_SERVICE "for port 80". Docker sets two network sysctls in every container network namespace that make both unnecessary on bridge networks:
net.ipv4.ping_group_range = 0 2147483647 lets every group ID open ICMP datagram sockets, which ping uses when it has no raw socket, so ping works with every capability dropped and as UID 10001. net.ipv4.ip_unprivileged_port_start = 0 makes every port unprivileged, so UID 10001 with no capabilities listens on port 80. Neither sysctl is a capability check, so the review answer to "add NET_RAW for ping" is no, and dropping NET_RAW removes the packet-forging risk from the table above at no cost. Note that Ubuntu 26.04 sets the same ping range on the host itself; other distributions may not.
Host networking is the exception. A container with --network host sees the host's sysctls, where unprivileged ports start at 1024:
As UID 10001 the bind fails even with --cap-add NET_BIND_SERVICE, for the reason the previous section showed: the capability only reached the bounding set. As root with every other capability dropped, the same command listens on port 81. For a non-root process on the host network, the options are a port at or above 1024, a file capability on the binary (setcap cap_net_bind_service=+ep, which no-new-privileges blocks at exec), or a root process holding only that one capability. On bridge networks, none of this applies.
Drop ALL, then add back what the error names
Root with no capabilities is much weaker, but it is still UID 0:
chown needs CAP_CHOWN and fails. Appending to /etc/motd succeeds, because root owns that file and the owner's write permission involves no capability check. Inside the image that is a nuisance; on a root-owned bind mount it is a write to the host, which "What root in a container really is" demonstrates. Dropping capabilities and running as a non-root user are separate controls, and you want both ("Run as non-root").
For an image you did not write, the workflow is mechanical. Drop everything, start it, and read the first error. The official nginx image starts as root and switches its workers to the nginx user (UID 101):
With nothing, the master process cannot chown its cache directory and exits 1. With CHOWN added, the container is Up, which is exactly the trap: every worker fails setgid(101) and nginx is running without anything to serve requests. Check the logs after every step as well as the status. Adding SETUID and SETGID completes the set. In Compose the same settings are cap_drop, cap_add and security_opt:
services:web:image: nginx:1.30-alpinecap_drop:- ALLcap_add:- CHOWN- SETUID- SETGIDsecurity_opt:- no-new-privileges:true
The page is served. The master process holds three capabilities; its first worker, already switched to UID 101, holds none; and NoNewPrivs is 1. nginx calls setuid and setgid as system calls, which no-new-privileges does not block, so the flag costs this image nothing. Images that switch user with gosu or su-exec work the same way and need the same two capabilities; an image that starts as a non-root USER needs neither.
The audit view from Docker's side reads HostConfig; Docker normalises names to the CAP_ form:
lab-web, started with defaults, shows empty lists and holds 14 capabilities; an audit flags containers like it, with nothing dropped and no no-new-privileges. Clean up:
--cap-add NET_ADMIN so that it can change its routes. Inside the container, ip route add fails with "Operation not permitted". Why?--cap-drop ALL. Running that binary prints euid=0. What does the process now have?--cap-add NET_RAW so a non-root container on a user-defined bridge network can ping its gateway in a health check. What is the right response?Try this
Work through “Drop ALL, then add back what the error names” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from capabilities, cap-drop and no-new-privileges, keep “Drop ALL, then add back what the error names”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.