What root in a container really is

UID 0 inside, the same UID outside, and what that means for mounts.

Advanced11 min · lesson 4 of 24

ls -ln on a directory you own shows a file you cannot edit. Its owner is 0, it appeared a minute ago, and the program that wrote it was a container you started without sudo. Nothing went wrong: without a user namespace, the UID 0 that id reports inside a container is the same UID 0 the host kernel treats as root, and a bind mount hands that root a piece of the host's filesystem. This lesson proves it from both sides on the main lab VM and shows the fields that tell you whether a given container runs as root. Everything happens under ~/lab/rootuser, which the cleanup removes.

A file written by container root

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ mkdir -p ~/lab/rootuser/shared ls -ldn ~/lab/rootuser/shared
drwxr-xr-x 2 1000 1000 4096 Oct 8 00:31 /home/ubuntu/lab/rootuser/shared
$ docker run --rm -v ~/lab/rootuser/shared:/data alpine:3.22 sh -c 'id; echo written-by-container-root > /data/from-root; ls -ln /data'
uid=0(root) gid=0(root) groups=0(root),0(root),1(bin),2(daemon),3(sys),4(adm),6(disk),10(wheel),11(floppy),20(dialout),26(tape),27(video) total 4 -rw-r--r-- 1 0 0 26 Oct 7 19:01 from-root
$ ls -ln ~/lab/rootuser/shared echo more >> ~/lab/rootuser/shared/from-root
total 4 -rw-r--r-- 1 0 0 26 Oct 8 00:31 from-root -bash: line 2: /home/ubuntu/lab/rootuser/shared/from-root: Permission denied

The directory belongs to ubuntu (UID 1000). Inside the container, id says root and the new file is owned by 0. On the host it is still owned by 0, and ubuntu, who owns the directory, gets Permission denied when trying to append to it. The kernel stores only numbers on files and processes. With no user namespace in between (the previous lessons showed the container sharing the host's user namespace), the container's 0 is written to disk as 0, and the host reads 0 as root. Being host UID 0 is not the same as holding all of root's power. The default capability set, seccomp and AppArmor still limit what the process can ask the kernel to do ("Capabilities, cap-drop and no-new-privileges"). The UID decides every ownership and permission check on the host paths the container can reach.

Reading works the same way. The lab puts a marker file in a directory only host root can open, then bind-mounts that directory read-only into a container:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ sudo install -d -m 700 ~/lab/rootuser/root-only sudo sh -c "echo lab-marker > ~ubuntu/lab/rootuser/root-only/marker" cat ~/lab/rootuser/root-only/marker
cat: /home/ubuntu/lab/rootuser/root-only/marker: Permission denied
$ docker run --rm -v ~/lab/rootuser/root-only:/data:ro alpine:3.22 cat /data/marker
lab-marker

ubuntu cannot read the marker. A container started by ubuntu reads it without trouble, because dockerd runs as root and performs the mount, and the process inside is UID 0 holding CAP_DAC_OVERRIDE from the default capability set. Even with every capability dropped, root still passes the owner checks on root-owned files, as "Capabilities, cap-drop and no-new-privileges" showed. So any account that can start containers can read and write root-owned host paths through mounts; "The Docker socket and daemon hardening" treats that as the root-equivalence it is.

The same UID from both sides

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-root alpine:3.22 sleep 600 ps -o user,pid,comm -p $(docker inspect -f '{{.State.Pid}}' lab-root)
158953162e3d5580216650f0a6edef60eab5ce7cb9314044a224f57f053a2b06 USER PID COMMAND root 301210 sleep
$ P=$(docker inspect -f '{{.State.Pid}}' lab-root) docker exec lab-root cat /proc/self/uid_map sudo readlink /proc/1/ns/user /proc/$P/ns/user
0 0 4294967295 user:[4026531837] user:[4026531837]

ps on the host shows the container's sleep running as root. The container's /proc/self/uid_map reads 0 0 4294967295: UIDs starting at 0 inside map to UIDs starting at 0 outside, for the whole 32-bit range. That is the identity map, the map of a process that is in the host's initial user namespace, and the user namespace inode is the host's 4026531837. A remapped container would show a different map, such as the 0 100000 65536 that "Namespaces from first principles" created with unshare.

Numeric IDs cross unchanged

The identity map works for every UID, not just 0. Run as UID 10001 and the same directory refuses the write; run as UID 1000 and the write succeeds:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --user 10001:10001 -v ~/lab/rootuser/shared:/data alpine:3.22 sh -c 'id; touch /data/from-10001'
uid=10001 gid=10001 groups=10001 touch: /data/from-10001: Permission denied
$ docker run --rm --user 1000:1000 -v ~/lab/rootuser/shared:/data alpine:3.22 sh -c 'id; touch /data/from-1000' ls -ln ~/lab/rootuser/shared
uid=1000 gid=1000 groups=1000 total 4 -rw-r--r-- 1 1000 1000 0 Oct 8 00:31 from-1000 -rw-r--r-- 1 0 0 26 Oct 8 00:31 from-root

UID 10001 belongs to no account on this host, so it gets only the "other" permissions of the directory (r-x) and cannot create a file. UID 1000 has no name inside Alpine, but on the host it is ubuntu, the owner, so the file lands as ubuntu's. On most single-user machines and many servers UID 1000 is the first login account. A container told --user 1000 acts as that person on every path mounted into it. Use a dedicated, high, unshared UID such as 10001 or 65532; "Run as non-root" covers how to bake one into an image and keep volumes writable for it.

What the UID decides in a breakout

Container escapes are judged by the UID they land on. CVE-2024-21626 ("Leaky Vessels") is a good example: runc before 1.1.12 leaked a file descriptor for a host directory into the container, and an image whose WORKDIR pointed at /proc/self/fd/<n>, or a docker exec with such a working directory, started a process with its current directory on the host's filesystem. What that process could then change was decided by its UID. As root without a user namespace, it could rewrite host binaries. As UID 10001, it could touch only what UID 10001 may touch, which on a normal host is almost nothing. The lab engine's runc is well past the fix:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker version --format '{{range .Server.Components}}{{if eq .Name "runc"}}runc {{.Version}}{{end}}{{end}}'
runc 1.5.1
A process gets out of its container
Escape to the host filesystem
same bug, different starting UID
root, no user namespace
Host UID 0
can change any file on the host
USER 10001, no user namespace
Host UID 10001
whatever that UID owns, normally nothing
root with userns-remap
Host UID 100000 or similar
an unprivileged subordinate UID that owns no host files
The escape is the same in each case. The UID is fixed when the container is started, by the image's USER, --user and the daemon's user-namespace setting.

Finding root containers

Two sources answer "does this container run as root". docker inspect gives Config.User, which is the image's USER or the --user value, and HostConfig.UsernsMode. ps on the host gives the real numeric UID:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-app --user 10001:10001 alpine:3.22 sleep 600 for c in lab-root lab-app; do docker inspect -f '{{.Name}} user=[{{.Config.User}}] userns=[{{.HostConfig.UsernsMode}}] privileged={{.HostConfig.Privileged}}' $c ps -o user=,pid=,comm= -p $(docker inspect -f '{{.State.Pid}}' $c) done
999befb1a7afc38c39f2126f23480c63e8dc467d67227248c2fed6a2dcd7462f /lab-root user=[] userns=[] privileged=false root 301210 sleep /lab-app user=[10001:10001] userns=[] privileged=false 10001 301630 sleep

An empty user=[] means the image never left root and nobody passed --user. An empty userns=[] means the host's user namespace. lab-root has both and runs as host root; lab-app runs as 10001. Config.User can hold a name that is only resolved inside the image, so for an audit the host-side ps value is the one to trust. Clean up (the directory holds root-owned files, hence sudo):

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-root lab-app sudo rm -rf ~/lab/rootuser
lab-root lab-app

Changing what UID 0 means

Running as a non-root user fixes the problem one image at a time. Two daemon-level options change what container root is for every container on a host. User-namespace remapping (userns-remap in daemon.json) runs containers in a user namespace whose UID 0 maps to a subordinate range from /etc/subuid, such as 100000 to 165535; the exact start depends on the entries already in that file. On Docker 29 it also turns off the containerd image store, and some flags, among them --privileged, need --userns host and are back to real root; "User-namespace remapping" works through that on the sec VM. Rootless Docker runs the daemon itself as an ordinary user inside a user namespace, so even the daemon is not root; "Rootless Docker" covers its networking, storage and limits.

Quick check
01A container runs as root with no user namespace and bind-mounts /srv/app, owned by deploy (UID 1000, mode 755). An attacker inside creates /srv/app/cron.sh. What does the host show?
Correct — Container UID 0 is written to disk as 0, exactly as from-root in the lab.
Incorrect — Docker does no such mapping; ownership is the writing process's UID.
Incorrect — Bind mounts are read-write unless you add :ro; --privileged is not involved.
Incorrect — Remapping is off unless the daemon sets userns-remap; the lab container's uid_map is the identity map.
02Inside a container, cat /proc/self/uid_map prints 0 0 4294967295. What does that tell you?
Incorrect — The third column is the length of the mapped range, not a quota on users.
Incorrect — An offset of 0 to 0 is no remapping; a remapped container shows something like 0 100000 65536.
Correct — It is the identity map of the host's initial user namespace, so container UID 0 is host UID 0.
Incorrect — An unset map is empty and would show nothing; this one maps the full range.
03A team wants --user 1000:1000 for a container that bind-mounts a host data directory. On that host, UID 1000 is the administrator's login account. What is the problem?
Incorrect — Docker reserves no such UID; the lab runs --user 1000 without trouble.
Correct — Numeric IDs cross unchanged, so container UID 1000 is the admin on those paths. Use a dedicated high UID.
Incorrect — They can, wherever the host permissions allow it; UID 1000 created from-1000 in the lab.
Incorrect — There is no translation without a user namespace; 1000 stays 1000.

Try this

Work through “Changing what UID 0 means” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from what root in a container really is, keep “Changing what UID 0 means”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related