User-namespace remapping
userns-remap, subordinate IDs and its limits on Docker 29.
secopslog-docker-sec). The lab edits /etc/docker/daemon.json, changes /etc/subuid and /etc/subgid and restarts dockerd. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM docker runs through sudo.Someone adds "userns-remap": "default" to daemon.json on a Docker 29 host, restarts the daemon, and every image is gone. docker image ls is empty, CI has to pull everything again, and the storage driver line in docker info has changed. Nothing was deleted. The daemon now uses a different image store in a different directory, and it does so on purpose. This lesson enables user-namespace remapping on the sec VM, walks through what it changes on the host, and shows the limits you have to plan around.
Before: container root is host root
"What root in a container really is" showed the default. Here it is again on the sec VM, with the daemon's storage settings and the number of images it holds:
There is no /etc/docker/daemon.json yet. The daemon uses the containerd image store (overlayfs, io.containerd.snapshotter.v1, the default for fresh Docker 29 installs) and holds 18 images (the count depends on what you pulled earlier). The uid_map line 0 0 4294967295 is the identity map: container UID 0 is host UID 0, for the whole UID space. So a file that container root writes through a bind mount belongs to host root.
Turn it on
userns-remap is a daemon-wide setting. With it, dockerd still runs as root, but it starts each container in a new user namespace whose UIDs map to a block of unprivileged host UIDs. Change daemon.json the way "Configuring the daemon safely" (Docker in depth) teaches: back up the file, or record {} when there is none, as here; merge the new key into the backup with jq so any settings already in the file (the lab kit writes MTU settings on VPN hosts) survive; validate; restart:
The validation already warns about the main side effect, userns remapping enabled, disabling containerd snapshotter, which the storage section below explains. default tells Docker to create a system account named dockremap and use the subordinate UID and GID ranges assigned to it in /etc/subuid and /etc/subgid. You can name an existing account instead ("userns-remap": "builder"). Check what Docker created:
Look at the ranges. dockremap got 100000:65536, exactly the range ubuntu already had. Docker's docs say the ranges must not overlap, because a process in one namespace would then have the same host UIDs as processes in another. Here, root in every remapped container would be the same host UID (100000) as UID 1 in the containers that ubuntu runs with rootless Docker or Podman. Always read both files after enabling remap. Now the storage:
Three things changed at once. docker info reports overlay2 with graph-driver details instead of the containerd snapshotter. Docker Root Dir moved to /var/lib/docker/100000.100000, a directory named after the mapped UID and GID, so images and containers from before are hidden (not deleted). And the image count is 0. The move to overlay2 is not a side effect of the new directory. Docker 29 made the containerd image store the default for fresh installs, and its release notes say that this does not apply to daemons configured with userns-remap: the containerd store is not available with remapping, as a workaround for moby/moby issue 47377. A remapped Docker 29 daemon always uses the older graph-driver store. Features that need the containerd store, such as keeping multi-platform images and attestations locally, are not available on that host. "Where Docker keeps data" explains the two stores. SecurityOptions now includes name=userns, which is the quickest check that remapping is active.
Fix the overlap before running anything. usermod can replace a subordinate range; pick one that no other account uses, then restart:
The data directory follows the range, /var/lib/docker/165536.165536, so changing the range later hides everything again. Allocate the ranges in configuration management before you enable remapping, the same on every host, and check them with grep as part of the rollout. The image store is empty in the new directory, which is why the last line pulls alpine:3.22 again.
The mapping, from both sides
Inside, the process is uid=0(root). The map says container UIDs 0 to 65535 are host UIDs 165536 to 231071, and from the host the sleep shows as UID 165536, an account that owns nothing. Under userns-remap container UID 0 maps to the first subordinate UID and container UID n to start + n. Rootless Docker maps differently (container root becomes your own UID), as "Rootless Docker" shows.
Root inside the container still holds Docker's default capabilities, but they apply only to objects owned by its user namespace. Operations that need privilege in the host's own user namespace, such as creating device nodes with mknod, fail (Docker's userns-remap documentation lists them; this lab does not run them), and the next writes behave as they do for the same reason.
Host files and bind mounts
/srv/lab-data belongs to host root, and container root is host UID 165536, so the write is refused. A database container that worked yesterday with a root-owned data directory fails the same way after you enable remap. Hand the directory to the mapped range rather than turning remapping off:
The new file is owned by 165536 on the host. The older before-remap file still belongs to UID 0. Container users other than root shift the same way (container UID 999 becomes host UID 166535), so chown to the UID the container process actually runs as. Named volumes are created under the remapped data root with the right owner, which is usually easier than bind mounts.
Opting out, and what does not work
A single container can leave the remapping with --userns=host. It then runs in the host's user namespace, and its root is host root again, even on a directory that remapped containers could not write to:
userns-host is owned by UID 0. Docker's docs add a side effect: the image layers are still owned by the remapped range, so programs that check file ownership, such as sudo or setuid binaries, may misbehave in such a container. Some features need the host namespace and are refused while remapping is on:
--privileged is refused (exit 125) unless you also pass --userns=host, and then the map is the identity map again. Sharing the host network namespace is refused too, and the docs list --pid=host with it. External volume and storage plugins that do not understand the mapping also do not work. Every opt-out brings back host root for that container, so find them:
An empty HostConfig.UsernsMode means the container is remapped. host means it opted out. Run the sweep on a schedule and require a written reason for every host.
Remapping has limits. All containers on the daemon share one range, so root in one container is the same host UID as root in another; it separates containers from the host, not from each other. It does not protect against a kernel bug. And it does nothing about access to the Docker socket, because the daemon is still root: anyone in the docker group can still start a --userns=host --privileged container, and "The Docker socket and daemon hardening" shows what such access does with a host-root bind mount.
Back to the default
Restore the backup and restart. This lesson restarted the daemon three times within a minute, which trips systemd's start limit for docker.service ("start of the service was attempted too often"); systemctl reset-failed clears it. Then remove the two remapped data roots, which the default daemon never reads:
The 18 images are back and the containerd store is active again. Without the last step, the remapped data under /var/lib/docker/100000.100000 and /var/lib/docker/165536.165536 would stay on disk indefinitely. Plan remapping for a new host rather than switching an existing one back and forth.
"userns-remap": "default" to a Docker 29 host and restarting, docker image ls is empty and docker info shows overlay2. What happened?165536:65536, a Postgres container that runs as container UID 999 cannot write to its bind-mounted data directory owned by root. Which fix keeps the protection?/etc/subuid and /etc/subgid right after enabling "userns-remap": "default"?Try this
Work through “Back to the default” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from user-namespace remapping, keep “Back to the default”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.