User-namespace remapping

userns-remap, subordinate IDs and its limits on Docker 29.

Advanced12 min · lesson 16 of 24
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab edits /etc/docker/daemon.json, changes /etc/subuid and /etc/subgid and restarts dockerd. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM docker runs through sudo.

Someone adds "userns-remap": "default" to daemon.json on a Docker 29 host, restarts the daemon, and every image is gone. docker image ls is empty, CI has to pull everything again, and the storage driver line in docker info has changed. Nothing was deleted. The daemon now uses a different image store in a different directory, and it does so on purpose. This lesson enables user-namespace remapping on the sec VM, walks through what it changes on the host, and shows the limits you have to plan around.

Before: container root is host root

"What root in a container really is" showed the default. Here it is again on the sec VM, with the daemon's storage settings and the number of images it holds:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker pull -q alpine:3.22 ls -l /etc/docker sudo docker info --format "{{.Driver}} {{json .DriverStatus}} root={{.DockerRootDir}}" echo "images: $(sudo docker image ls -q | wc -l)"
docker.io/library/alpine:3.22 total 0 overlayfs [["driver-type","io.containerd.snapshotter.v1"]] root=/var/lib/docker images: 18
$ sudo docker run --rm alpine:3.22 cat /proc/self/uid_map
0 0 4294967295
$ sudo mkdir -p /srv/lab-data sudo docker run --rm -v /srv/lab-data:/data alpine:3.22 touch /data/before-remap ls -ln /srv/lab-data
total 0 -rw-r--r-- 1 0 0 0 Oct 8 03:16 before-remap

There is no /etc/docker/daemon.json yet. The daemon uses the containerd image store (overlayfs, io.containerd.snapshotter.v1, the default for fresh Docker 29 installs) and holds 18 images (the count depends on what you pulled earlier). The uid_map line 0 0 4294967295 is the identity map: container UID 0 is host UID 0, for the whole UID space. So a file that container root writes through a bind mount belongs to host root.

Turn it on

userns-remap is a daemon-wide setting. With it, dockerd still runs as root, but it starts each container in a new user namespace whose UIDs map to a block of unprivileged host UIDs. Change daemon.json the way "Configuring the daemon safely" (Docker in depth) teaches: back up the file, or record {} when there is none, as here; merge the new key into the backup with jq so any settings already in the file (the lab kit writes MTU settings on VPN hosts) survive; validate; restart:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo cp -a /etc/docker/daemon.json /etc/docker/daemon.json.bak 2>/dev/null || echo "{}" | sudo tee /etc/docker/daemon.json.bak >/dev/null jq ". + {\"userns-remap\": \"default\"}" /etc/docker/daemon.json.bak | sudo tee /etc/docker/daemon.json sudo dockerd --validate --config-file /etc/docker/daemon.json sudo systemctl restart docker
{ "userns-remap": "default" } time="2026-10-08T03:16:48.353708040+05:30" level=info msg="User namespaces: ID ranges will be mapped to subuid/subgid ranges of: dockremap" time="2026-10-08T03:16:48.353817956+05:30" level=warning msg="userns remapping enabled, disabling containerd snapshotter" configuration OK

The validation already warns about the main side effect, userns remapping enabled, disabling containerd snapshotter, which the storage section below explains. default tells Docker to create a system account named dockremap and use the subordinate UID and GID ranges assigned to it in /etc/subuid and /etc/subgid. You can name an existing account instead ("userns-remap": "builder"). Check what Docker created:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ id dockremap grep -E "^(ubuntu|dockremap):" /etc/subuid /etc/subgid
uid=104(dockremap) gid=109(dockremap) groups=109(dockremap) /etc/subuid:ubuntu:100000:65536 /etc/subuid:dockremap:100000:65536 /etc/subgid:ubuntu:100000:65536 /etc/subgid:dockremap:100000:65536

Look at the ranges. dockremap got 100000:65536, exactly the range ubuntu already had. Docker's docs say the ranges must not overlap, because a process in one namespace would then have the same host UIDs as processes in another. Here, root in every remapped container would be the same host UID (100000) as UID 1 in the containers that ubuntu runs with rootless Docker or Podman. Always read both files after enabling remap. Now the storage:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker info --format "{{.Driver}} {{json .DriverStatus}} root={{.DockerRootDir}}" sudo docker info --format "{{.SecurityOptions}}" echo "images: $(sudo docker image ls -q | wc -l)" sudo sh -c "ls -ld /var/lib/docker/*.*"
overlay2 [["Backing Filesystem","extfs"],["Supports d_type","true"],["Using metacopy","false"],["Native Overlay Diff","true"],["userxattr","false"]] root=/var/lib/docker/100000.100000 [name=apparmor,profile=default name=seccomp,profile=builtin name=userns name=cgroupns] images: 0 drwx--x--- 12 root 100000 4096 Oct 8 03:16 /var/lib/docker/100000.100000

Three things changed at once. docker info reports overlay2 with graph-driver details instead of the containerd snapshotter. Docker Root Dir moved to /var/lib/docker/100000.100000, a directory named after the mapped UID and GID, so images and containers from before are hidden (not deleted). And the image count is 0. The move to overlay2 is not a side effect of the new directory. Docker 29 made the containerd image store the default for fresh installs, and its release notes say that this does not apply to daemons configured with userns-remap: the containerd store is not available with remapping, as a workaround for moby/moby issue 47377. A remapped Docker 29 daemon always uses the older graph-driver store. Features that need the containerd store, such as keeping multi-platform images and attestations locally, are not available on that host. "Where Docker keeps data" explains the two stores. SecurityOptions now includes name=userns, which is the quickest check that remapping is active.

Fix the overlap before running anything. usermod can replace a subordinate range; pick one that no other account uses, then restart:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo usermod --del-subuids 100000-165535 --add-subuids 165536-231071 \ --del-subgids 100000-165535 --add-subgids 165536-231071 dockremap grep -E "^(ubuntu|dockremap):" /etc/subuid /etc/subgid sudo systemctl restart docker sudo docker info --format "root={{.DockerRootDir}}" sudo docker pull -q alpine:3.22
/etc/subuid:ubuntu:100000:65536 /etc/subuid:dockremap:165536:65536 /etc/subgid:ubuntu:100000:65536 /etc/subgid:dockremap:165536:65536 root=/var/lib/docker/165536.165536 docker.io/library/alpine:3.22

The data directory follows the range, /var/lib/docker/165536.165536, so changing the range later hides everything again. Allocate the ranges in configuration management before you enable remapping, the same on every host, and check them with grep as part of the rollout. The image store is empty in the new directory, which is why the last line pulls alpine:3.22 again.

The mapping, from both sides

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm alpine:3.22 sh -c "id; cat /proc/self/uid_map"
uid=0(root) gid=0(root) groups=0(root),0(root),1(bin),2(daemon),3(sys),4(adm),6(disk),10(wheel),11(floppy),20(dialout),26(tape),27(video) 0 165536 65536
$ sudo docker run -d --name lab-remap alpine:3.22 sleep 300 >/dev/null ps -o user,uid,pid,args -C sleep sudo docker rm -f lab-remap >/dev/null
USER UID PID COMMAND 165536 165536 3266 sleep 300

Inside, the process is uid=0(root). The map says container UIDs 0 to 65535 are host UIDs 165536 to 231071, and from the host the sleep shows as UID 165536, an account that owns nothing. Under userns-remap container UID 0 maps to the first subordinate UID and container UID n to start + n. Rootless Docker maps differently (container root becomes your own UID), as "Rootless Docker" shows.

Root inside the container still holds Docker's default capabilities, but they apply only to objects owned by its user namespace. Operations that need privilege in the host's own user namespace, such as creating device nodes with mknod, fail (Docker's userns-remap documentation lists them; this lab does not run them), and the next writes behave as they do for the same reason.

Host files and bind mounts

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm -v /srv/lab-data:/data alpine:3.22 touch /data/after-remap
touch: /data/after-remap: Permission denied

/srv/lab-data belongs to host root, and container root is host UID 165536, so the write is refused. A database container that worked yesterday with a root-owned data directory fails the same way after you enable remap. Hand the directory to the mapped range rather than turning remapping off:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ START=$(awk -F: "\$1==\"dockremap\"{print \$2}" /etc/subuid) sudo chown "$START:$START" /srv/lab-data sudo docker run --rm -v /srv/lab-data:/data alpine:3.22 touch /data/after-remap ls -ln /srv/lab-data
total 0 -rw-r--r-- 1 165536 165536 0 Oct 8 03:16 after-remap -rw-r--r-- 1 0 0 0 Oct 8 03:16 before-remap

The new file is owned by 165536 on the host. The older before-remap file still belongs to UID 0. Container users other than root shift the same way (container UID 999 becomes host UID 166535), so chown to the UID the container process actually runs as. Named volumes are created under the remapped data root with the right owner, which is usually easier than bind mounts.

Opting out, and what does not work

A single container can leave the remapping with --userns=host. It then runs in the host's user namespace, and its root is host root again, even on a directory that remapped containers could not write to:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo chown root:root /srv/lab-data sudo docker run --rm --userns=host -v /srv/lab-data:/data alpine:3.22 touch /data/userns-host ls -ln /srv/lab-data
total 0 -rw-r--r-- 1 165536 165536 0 Oct 8 03:16 after-remap -rw-r--r-- 1 0 0 0 Oct 8 03:16 before-remap -rw-r--r-- 1 0 0 0 Oct 8 03:16 userns-host

userns-host is owned by UID 0. Docker's docs add a side effect: the image layers are still owned by the remapped range, so programs that check file ownership, such as sudo or setuid binaries, may misbehave in such a container. Some features need the host namespace and are refused while remapping is on:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run --rm --privileged alpine:3.22 true
docker: Error response from daemon: privileged mode is incompatible with user namespaces. You must run the container in the host namespace when running privileged mode Run 'docker run --help' for more information
$ sudo docker run --rm --privileged --userns=host alpine:3.22 cat /proc/self/uid_map
0 0 4294967295
$ sudo docker run --rm --network=host alpine:3.22 true
docker: Error response from daemon: cannot share the host's network namespace when user namespaces are enabled Run 'docker run --help' for more information

--privileged is refused (exit 125) unless you also pass --userns=host, and then the map is the identity map again. Sharing the host network namespace is refused too, and the docs list --pid=host with it. External volume and storage plugins that do not understand the mapping also do not work. Every opt-out brings back host root for that container, so find them:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker run -d --name lab-mapped alpine:3.22 sleep 300 >/dev/null sudo docker run -d --name lab-optout --userns=host alpine:3.22 sleep 300 >/dev/null sudo docker ps -q | xargs sudo docker inspect --format "{{.Name}} userns={{if .HostConfig.UsernsMode}}{{.HostConfig.UsernsMode}}{{else}}remapped{{end}}" sudo docker rm -f lab-mapped lab-optout >/dev/null
/lab-optout userns=host /lab-mapped userns=remapped

An empty HostConfig.UsernsMode means the container is remapped. host means it opted out. Run the sweep on a schedule and require a written reason for every host.

Remapping has limits. All containers on the daemon share one range, so root in one container is the same host UID as root in another; it separates containers from the host, not from each other. It does not protect against a kernel bug. And it does nothing about access to the Docker socket, because the daemon is still root: anyone in the docker group can still start a --userns=host --privileged container, and "The Docker socket and daemon hardening" shows what such access does with a host-root bind mount.

Back to the default

Restore the backup and restart. This lesson restarted the daemon three times within a minute, which trips systemd's start limit for docker.service ("start of the service was attempted too often"); systemctl reset-failed clears it. Then remove the two remapped data roots, which the default daemon never reads:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo cp /etc/docker/daemon.json.bak /etc/docker/daemon.json sudo systemctl reset-failed docker.service sudo systemctl restart docker echo "images: $(sudo docker image ls -q | wc -l)" sudo docker info --format "{{.Driver}} {{json .DriverStatus}} root={{.DockerRootDir}}"
images: 18 overlayfs [["driver-type","io.containerd.snapshotter.v1"]] root=/var/lib/docker
$ sudo rm -rf /srv/lab-data /var/lib/docker/100000.100000 /var/lib/docker/165536.165536 sudo ls /var/lib/docker
buildkit engine-id network rootfs swarm volumes containers image plugins runtimes tmp

The 18 images are back and the containerd store is active again. Without the last step, the remapped data under /var/lib/docker/100000.100000 and /var/lib/docker/165536.165536 would stay on disk indefinitely. Plan remapping for a new host rather than switching an existing one back and forth.

Quick check
01After adding "userns-remap": "default" to a Docker 29 host and restarting, docker image ls is empty and docker info shows overlay2. What happened?
Incorrect — Nothing was deleted. In the lab all 18 images were back as soon as remapping was turned off.
Correct — Docker 29 excludes userns-remap from the containerd image store (moby/moby#47377), and the remapped daemon keeps its data under /var/lib/docker/<uid>.<gid>.
Incorrect — The daemon started fine; docker info answered and listed name=userns in the security options.
Incorrect — Docker did not convert anything; the old images stayed where they were, hidden from the remapped daemon.
02With remapping on and dockremap at 165536:65536, a Postgres container that runs as container UID 999 cannot write to its bind-mounted data directory owned by root. Which fix keeps the protection?
Incorrect — That makes its root host root again, which is the exposure remapping removes.
Incorrect — That lets every user and process on the host write the database files.
Correct — Container UID 999 is host UID 165536 + 999 = 166535, so give the data to that UID.
Incorrect — 165536 is container root. Postgres runs as UID 999 in the container, so it still could not write.
03Why should you compare /etc/subuid and /etc/subgid right after enabling "userns-remap": "default"?
Correct — On the sec VM, dockremap got 100000:65536, the same range as ubuntu, so remapped root would equal ubuntu's rootless UID 1. Ranges must not overlap.
Incorrect — They only hold subordinate ID ranges per account. Opt-outs are found with docker inspect.
Incorrect — The ubuntu entries were unchanged; the problem was a duplicate range, not a removed one.
Incorrect — Image policy has nothing to do with subordinate ID files.

Try this

Work through “Back to the default” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from user-namespace remapping, keep “Back to the default”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related