Capabilities, cap-drop and no-new-privileges

The pieces of root a container keeps and how to drop them.

Advanced16 min · lesson 3 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (3 files, 1 KB): caps.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/caps.tar.gz && tar -xzf caps.tar.gz, which creates ~/lab/caps/. SHA-256: e9fa31f68ef4143cbc688dcae9b46d8e1373067fd34290b2498a813f10bf85cb

A change request asks for two flags on a service that runs as a non-root user: --cap-add NET_BIND_SERVICE so it can listen on port 80, and --cap-add NET_RAW so its health check can ping a dependency. On Docker 29 with a bridge network neither is needed, and the first would not have worked anyway. Reviewing requests like that needs a precise picture of what Linux capabilities are, which ones a container holds, and why the user ID changes what an added capability does. This lesson builds that picture on the main lab VM and ends with the working pattern: drop everything, add back only what an error message names, and set no-new-privileges.

The lesson files (copy them to ~/lab/caps) contain a test image and a Compose file. The image, lab-caps:1, adds capsh and getpcaps from Alpine's libcap-utils and a small Go program that prints its user IDs and capability sets. A second copy of that program is setuid root. It does nothing else, so it shows what a setuid-root binary in an image is worth without exploiting anything:

Dockerfile
# lab-caps:1, the capabilities lesson's test image: capsh/getpcaps from libcap-utils, a tiny Go program that
# prints its user IDs and capability sets, and a second copy of it with the setuid bit (owned by root).
FROM golang:1.27-alpine@sha256:8a5910f31396cd4d89662f56c68b3ae31d374308270a1c3bd96672ee5ed43414 AS build
WORKDIR /src
COPY whoami.go .
RUN CGO_ENABLED=0 go build -o /out/lab-whoami whoami.go
FROM alpine:3.22
RUN apk add --no-cache libcap-utils
COPY --from=build /out/lab-whoami /usr/local/bin/lab-whoami
COPY --from=build /out/lab-whoami /usr/local/bin/lab-suid-whoami
RUN chmod 4755 /usr/local/bin/lab-suid-whoami
USER 10001:10001
whoami.go
// lab-whoami prints the IDs and capability sets of the running process.
// The lab installs a second copy with the setuid bit as lab-suid-whoami.
package main
import (
"bufio"
"fmt"
"os"
"strings"
)
func main() {
fmt.Printf("uid=%d euid=%d\n", os.Getuid(), os.Geteuid())
f, err := os.Open("/proc/self/status")
if err != nil {
fmt.Println(err)
os.Exit(1)
}
defer f.Close()
s := bufio.NewScanner(f)
for s.Scan() {
l := s.Text()
if strings.HasPrefix(l, "CapPrm") || strings.HasPrefix(l, "CapEff") || strings.HasPrefix(l, "NoNewPrivs") {
fmt.Println(l)
}
}
}

Five sets per process

Linux splits the privileges of root into capabilities: CAP_CHOWN to change any file's owner, CAP_NET_ADMIN to change interfaces and routes, CAP_SYS_ADMIN for mounts and a long list of administrative operations, and so on, 41 on this kernel. The kernel checks the specific capability when a process attempts a privileged operation. Each process carries five sets of them:

Build the image in ~/lab/caps, then compare a default root container with what capsh reports:

ubuntu@secopslog-docker:~/lab/caps · Docker 29.8.2
$ docker build -q -t lab-caps:1 .
sha256:d674051011257f391ec53d862d59c375f8ad1ac6122a8a8ad7e1bde43bc8fe26
ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm alpine:3.22 grep Cap /proc/self/status
CapInh: 0000000000000000 CapPrm: 00000000a80425fb CapEff: 00000000a80425fb CapBnd: 00000000a80425fb CapAmb: 0000000000000000
$ docker run --rm --user 0 lab-caps:1 capsh --print
Current: cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap=ep Bounding set =cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap Ambient set = Current IAB: !cap_dac_read_search,!cap_linux_immutable,!cap_net_broadcast,!cap_net_admin,!cap_ipc_lock,!cap_ipc_owner,!cap_sys_module,!cap_sys_rawio,!cap_sys_ptrace,!cap_sys_pacct,!cap_sys_admin,!cap_sys_boot,!cap_sys_nice,!cap_sys_resource,!cap_sys_time,!cap_sys_tty_config,!cap_lease,!cap_audit_control,!cap_mac_override,!cap_mac_admin,!cap_syslog,!cap_wake_alarm,!cap_block_suspend,!cap_audit_read,!cap_perfmon,!cap_bpf,!cap_checkpoint_restore Securebits: 00/0x0/1'b0 (no-new-privs=0) secure-noroot: no (unlocked) secure-no-suid-fixup: no (unlocked) secure-keep-caps: no (unlocked) secure-no-ambient-raise: no (unlocked) uid=0(root) euid=0(root) gid=0(root) groups=0(root),0(root),1(bin),2(daemon),3(sys),4(adm),6(disk),10(wheel),11(floppy),20(dialout),26(tape),27(video) Guessed mode: HYBRID (4)
$ docker run --rm lab-caps:1 capsh --decode=00000000a80425fb
0x00000000a80425fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap

A root container holds a80425fb in its permitted, effective and bounding sets and nothing in the other two. capsh --print spells it out: the Current line lists 14 capabilities with =ep (effective and permitted), the bounding set is the same 14, and the Current IAB line marks the 27 capabilities missing from the bounding set with !, among them cap_sys_admin, cap_sys_module and cap_sys_ptrace. Securebits ends with no-new-privs=0, which matters below. capsh --decode turns any mask back into names; that is how you read a value copied out of /proc or an alert. Some of the 14, and some notable absentees:

The capabilities that are missing from the default set are the ones that most escape techniques need; "Capability and kernel escapes" goes through them. That includes CVE-2022-0492, a missing capability check in cgroup v1: a process without CAP_SYS_ADMIN in the host's user namespace could create a new user namespace, mount a cgroup v1 hierarchy there and set its release_agent, a program the host then runs as root. Docker's defaults blocked it twice: the default seccomp profile refuses unshare to a container without CAP_SYS_ADMIN, and AppArmor's docker-default refuses the mount. It is cgroup v1 only and not applicable on cgroup v2 hosts like this one, and kernels since 5.17 carry the fix.

Reading capabilities from the host

Inside a compromised container, capsh and every other tool can be replaced with one that prints whatever the attacker wants. The host reads the kernel's values for the same process directly. Take the main PID from docker inspect and ask the host:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-web nginx:1.30-alpine
23a48dcaf213114cde556ffbcc60a8475c63350706b9e741274779330a0a7eec
$ P=$(docker inspect -f '{{.State.Pid}}' lab-web) sudo getpcaps "$P" grep -E '^Cap(Eff|Bnd)' /proc/$P/status
646086: cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_net_raw,cap_sys_chroot,cap_mknod,cap_audit_write,cap_setfcap=ep CapEff: 00000000a80425fb CapBnd: 00000000a80425fb

getpcaps prints the names and /proc/<pid>/status the masks. For a fleet sweep read CapBnd, not CapEff. The effective set depends on the user (the next section shows a non-root container with CapEff 0 and no flags at all), while the bounding set reflects only the run flags: on this engine a default container's bounding set is 00000000a80425fb whatever its user, so any other value means --cap-add, --cap-drop or --privileged, and the larger values deserve the first look. Other runtimes and engine versions use different defaults, so take the baseline from your own hosts. The HostConfig fields at the end of this lesson give the same answer from the Docker side.

A non-root user gets none of them

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm lab-caps:1 sh -c 'id; grep Cap /proc/self/status'
uid=10001 gid=10001 groups=10001 CapInh: 0000000000000000 CapPrm: 0000000000000000 CapEff: 0000000000000000 CapBnd: 00000000a80425fb CapAmb: 0000000000000000
$ docker run --rm --cap-add SYS_ADMIN lab-caps:1 grep Cap /proc/self/status
CapInh: 0000000000000000 CapPrm: 0000000000000000 CapEff: 0000000000000000 CapBnd: 00000000a82425fb CapAmb: 0000000000000000

As UID 10001 (the image's USER), the permitted and effective sets are empty and only the bounding set holds the default 14. --cap-add SYS_ADMIN changes exactly one thing: the bounding set becomes a82425fb, with bit 21, CAP_SYS_ADMIN, added. When a non-root process starts a program, the kernel computes its new permitted set from the program file's capabilities and from the inheritable and ambient sets, and Docker leaves those empty. So a capability added for a non-root container sits in the ceiling where nothing can use it, unless a program in the image can raise the process into it. That is the point of the next demonstration, and the reason the change request's NET_BIND_SERVICE would not have helped.

setuid, the bounding set and no-new-privileges

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm lab-caps:1 ls -l /usr/local/bin
total 4608 -rwsr-xr-x 1 root root 2357514 Oct 7 22:04 lab-suid-whoami -rwxr-xr-x 1 root root 2357514 Oct 7 22:04 lab-whoami
$ docker run --rm --cap-add SYS_ADMIN lab-caps:1 sh -c 'lab-whoami; lab-suid-whoami'
uid=10001 euid=10001 CapPrm: 0000000000000000 CapEff: 0000000000000000 NoNewPrivs: 0 uid=10001 euid=0 CapPrm: 00000000a82425fb CapEff: 00000000a82425fb NoNewPrivs: 0
$ docker run --rm --cap-drop ALL lab-caps:1 lab-suid-whoami
uid=10001 euid=0 CapPrm: 0000000000000000 CapEff: 0000000000000000 NoNewPrivs: 0
$ docker run --rm --cap-add SYS_ADMIN --security-opt no-new-privileges lab-caps:1 lab-suid-whoami
uid=10001 euid=10001 CapPrm: 0000000000000000 CapEff: 0000000000000000 NoNewPrivs: 1

lab-suid-whoami has the setuid bit (-rwsr-xr-x) and belongs to root. Run by UID 10001, the plain copy has no capabilities. The setuid copy runs with effective UID 0 and, because a program that becomes root at execve gets the full bounding set, with a82425fb in permitted and effective: the SYS_ADMIN that --cap-add put in the ceiling is now usable. Any setuid-root binary in an image, su, mount, or a forgotten debug helper, is that path.

--cap-drop ALL empties the bounding set, so the setuid copy gets no capabilities, but its effective UID still becomes 0. Changing UID through a setuid bit needs no capability, and a process with UID 0 owns every root-owned file it can reach, as the section after next shows. --security-opt no-new-privileges closes the path itself. It sets the kernel's no_new_privs flag on the container's first process; children inherit it and nothing can clear it. With the flag set, execve ignores setuid and setgid bits and does not grant file capabilities, so the setuid copy stays at UID 10001 with nothing. Ordinary programs are unaffected.

A process running as USER 10001 starts a program
UID 10001 runs a program
permitted and effective are empty before the exec
ordinary binary
Still no capabilities
--cap-add only changed the bounding set
setuid root, no-new-privileges off
Effective UID 0 and the whole bounding set
including every --cap-add, e.g. a82425fb
setuid root, no-new-privileges on
Setuid bit ignored
UID 10001, no capabilities
The bounding set is the ceiling and no-new-privileges blocks the ways up to it. --cap-drop ALL lowers the ceiling to zero but does not stop the UID change.

Set no-new-privileges on every container. It can also be a daemon-wide default ("no-new-privileges": true in daemon.json), which "The Docker socket and daemon hardening" covers. Before you make it the default, find the images that depend on gaining privileges at exec: sudo or su, other setuid helpers, and binaries with file capabilities. find / -xdev -perm -4000 and getcap -r / inside each image list them. Those images fail under the flag and need fixing first.

ping and ports below 1024 on Docker

Two capabilities in the default set exist mostly because of old habits: NET_RAW "for ping" and NET_BIND_SERVICE "for port 80". Docker sets two network sysctls in every container network namespace that make both unnecessary on bridge networks:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm alpine:3.22 cat /proc/sys/net/ipv4/ping_group_range /proc/sys/net/ipv4/ip_unprivileged_port_start
0 2147483647 0
$ docker run --rm --cap-drop ALL --user 10001 alpine:3.22 ping -c1 127.0.0.1
PING 127.0.0.1 (127.0.0.1): 56 data bytes 64 bytes from 127.0.0.1: seq=0 ttl=42 time=0.352 ms --- 127.0.0.1 ping statistics --- 1 packets transmitted, 1 packets received, 0% packet loss round-trip min/avg/max = 0.352/0.352/0.352 ms
$ docker run --rm --cap-drop ALL --user 10001 alpine:3.22 sh -c 'nc -lk -p 80 & sleep 1; netstat -tln'
Active Internet connections (only servers) Proto Recv-Q Send-Q Local Address Foreign Address State tcp 0 0 :::80 :::* LISTEN

net.ipv4.ping_group_range = 0 2147483647 lets every group ID open ICMP datagram sockets, which ping uses when it has no raw socket, so ping works with every capability dropped and as UID 10001. net.ipv4.ip_unprivileged_port_start = 0 makes every port unprivileged, so UID 10001 with no capabilities listens on port 80. Neither sysctl is a capability check, so the review answer to "add NET_RAW for ping" is no, and dropping NET_RAW removes the packet-forging risk from the table above at no cost. Note that Ubuntu 26.04 sets the same ping range on the host itself; other distributions may not.

Host networking is the exception. A container with --network host sees the host's sysctls, where unprivileged ports start at 1024:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --network host --cap-drop ALL --cap-add NET_BIND_SERVICE --user 10001 alpine:3.22 sh -c 'cat /proc/sys/net/ipv4/ip_unprivileged_port_start; nc -l -p 81'
1024 nc: bind: Permission denied
$ docker run --rm --network host --cap-drop ALL --cap-add NET_BIND_SERVICE alpine:3.22 sh -c 'nc -lk -p 81 & sleep 1; netstat -tln | grep ":81 "'
tcp 0 0 :::81 :::* LISTEN

As UID 10001 the bind fails even with --cap-add NET_BIND_SERVICE, for the reason the previous section showed: the capability only reached the bounding set. As root with every other capability dropped, the same command listens on port 81. For a non-root process on the host network, the options are a port at or above 1024, a file capability on the binary (setcap cap_net_bind_service=+ep, which no-new-privileges blocks at exec), or a root process holding only that one capability. On bridge networks, none of this applies.

Drop ALL, then add back what the error names

Root with no capabilities is much weaker, but it is still UID 0:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --cap-drop ALL alpine:3.22 sh -c 'id -u; chown 99 /etc/motd; echo lab >> /etc/motd && ls -ln /etc/motd'
0 chown: /etc/motd: Operation not permitted -rw-r--r-- 1 0 0 288 Oct 7 22:04 /etc/motd

chown needs CAP_CHOWN and fails. Appending to /etc/motd succeeds, because root owns that file and the owner's write permission involves no capability check. Inside the image that is a nuisance; on a root-owned bind mount it is a write to the host, which "What root in a container really is" demonstrates. Dropping capabilities and running as a non-root user are separate controls, and you want both ("Run as non-root").

For an image you did not write, the workflow is mechanical. Drop everything, start it, and read the first error. The official nginx image starts as root and switches its workers to the nginx user (UID 101):

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --name lab-ng0 --cap-drop ALL nginx:1.30-alpine
... 2026/10/07 22:04:47 [emerg] 1#1: chown("/var/cache/nginx/client_temp", 101) failed (1: Operation not permitted) nginx: [emerg] chown("/var/cache/nginx/client_temp", 101) failed (1: Operation not permitted)
$ docker run -d --name lab-ng1 --cap-drop ALL --cap-add CHOWN nginx:1.30-alpine sleep 2 docker ps --filter name=lab-ng1 --format '{{.Names}} {{.Status}}' docker logs lab-ng1 2>&1 | grep -E 'emerg|alert' | head -3
c42c9a0828b84931cedf74f6666a803ee34b6ba180b12a88124874ea9e21d553 lab-ng1 Up 2 seconds 2026/10/07 22:04:47 [emerg] 30#30: setgid(101) failed (1: Operation not permitted) 2026/10/07 22:04:47 [emerg] 31#31: setgid(101) failed (1: Operation not permitted) 2026/10/07 22:04:47 [emerg] 32#32: setgid(101) failed (1: Operation not permitted)

With nothing, the master process cannot chown its cache directory and exits 1. With CHOWN added, the container is Up, which is exactly the trap: every worker fails setgid(101) and nginx is running without anything to serve requests. Check the logs after every step as well as the status. Adding SETUID and SETGID completes the set. In Compose the same settings are cap_drop, cap_add and security_opt:

compose.yaml
services:
web:
image: nginx:1.30-alpine
cap_drop:
- ALL
cap_add:
- CHOWN
- SETUID
- SETGID
security_opt:
- no-new-privileges:true
ubuntu@secopslog-docker:~/lab/caps · Docker 29.8.2
$ docker compose --progress quiet -p lab-caps up -d sleep 2 docker compose -p lab-caps exec web wget -qO- http://127.0.0.1/ | grep "<title>"
<title>Welcome to nginx!</title>
$ P=$(docker inspect -f '{{.State.Pid}}' lab-caps-web-1) sudo getpcaps "$P" $(pgrep -P "$P" | head -1) grep NoNewPrivs /proc/$P/status
648037: cap_chown,cap_setgid,cap_setuid=ep 648097: = NoNewPrivs: 1

The page is served. The master process holds three capabilities; its first worker, already switched to UID 101, holds none; and NoNewPrivs is 1. nginx calls setuid and setgid as system calls, which no-new-privileges does not block, so the flag costs this image nothing. Images that switch user with gosu or su-exec work the same way and need the same two capabilities; an image that starts as a non-root USER needs neither.

The audit view from Docker's side reads HostConfig; Docker normalises names to the CAP_ form:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ for c in lab-web lab-caps-web-1; do docker inspect -f '{{.Name}} add={{.HostConfig.CapAdd}} drop={{.HostConfig.CapDrop}} opt={{.HostConfig.SecurityOpt}}' $c done
/lab-web add=[] drop=[] opt=[] /lab-caps-web-1 add=[CAP_CHOWN CAP_SETGID CAP_SETUID] drop=[ALL] opt=[no-new-privileges:true]

lab-web, started with defaults, shows empty lists and holds 14 capabilities; an audit flags containers like it, with nothing dropped and no no-new-privileges. Clean up:

ubuntu@secopslog-docker:~/lab/caps · Docker 29.8.2
$ docker compose --progress quiet -p lab-caps down docker rm -f lab-web lab-ng0 lab-ng1 docker rmi lab-caps:1
lab-web lab-ng0 lab-ng1 Untagged: lab-caps:1 Deleted: sha256:d674051011257f391ec53d862d59c375f8ad1ac6122a8a8ad7e1bde43bc8fe26
Quick check
01A service runs as USER 10001 with --cap-add NET_ADMIN so that it can change its routes. Inside the container, ip route add fails with "Operation not permitted". Why?
Incorrect — NET_ADMIN applies within the process's own network namespace; a root container with it can change its routes.
Incorrect — Added capabilities are honoured under the default seccomp profile; the lab's root containers show them in CapEff.
Incorrect — Nothing about capabilities depends on a daemon restart.
Correct — Permitted and effective stay empty for UID 10001 unless a setuid or file-capability binary raises them.
02An image contains a setuid-root binary and runs as USER 10001 with --cap-drop ALL. Running that binary prints euid=0. What does the process now have?
Incorrect — With the bounding set empty, the exec grants nothing; the lab shows CapEff 0 with euid 0.
Correct — The UID change needs no capability, and file owner checks need none either. no-new-privileges would have kept euid at 10001.
Incorrect — Owner permission checks still pass for UID 0; the lab appends to a root-owned file with no capabilities.
Incorrect — The kernel has no such rule. UID and capabilities are independent.
03A teammate asks for --cap-add NET_RAW so a non-root container on a user-defined bridge network can ping its gateway in a health check. What is the right response?
Incorrect — ping falls back to an ICMP datagram socket when ping_group_range allows it, which Docker sets.
Incorrect — Reply routing needs no capability; the lab's ping runs with every capability dropped.
Correct — The lab pings as UID 10001 with --cap-drop ALL. Dropping NET_RAW also removes packet forging.
Incorrect — Host networking would remove the network namespace, a far bigger change than the one requested, to solve a problem that does not exist.

Try this

Work through “Drop ALL, then add back what the error names” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from capabilities, cap-drop and no-new-privileges, keep “Drop ALL, then add back what the error names”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related