Rootless Docker and why containers shouldn't run as root

Run the Docker daemon and your containers as a non-root user so a container escape lands unprivileged, plus the subuid, cgroups, and storage gotchas.

Oct 7, 2025·Updated ·6 min readIntermediate·By SecOpsLog · documentation-verified

Whoever can talk to /var/run/docker.sock on a default install is root on that host, because dockerd runs as root and will happily bind-mount / into a container for anyone the socket permits. A kernel bug or a runtime escape ends in the same place. Rootless mode moves the daemon itself into an unprivileged user account: dockerd and every container it starts run inside a user namespace owned by, say, alice, and a process that breaks out of a container is alice, or one of her subordinate ids, on the host.

What you give up is a list, not a feeling, and the Docker docs publish it: specific storage drivers, cgroup v2 as a requirement for resource limits, no privileged ports without a sysctl, and a user-mode network stack that is slower than the kernel's. Whether that trade is worth it depends on the host's job. For a CI builder or a shared development box it usually is; for a node that needs host networking plugins it usually is not.

Where container root lands on the host

The uid map that makes rootless rootless

Container uid 0 is host uid 100000 here because that is the first subordinate id in alice’s /etc/subuid entry; the 65,536-id range is the minimum the docs require. No container uid maps to host uid 0, which is the whole security argument in one row. Simplified: gids follow the same scheme through /etc/subgid.

Rootless Docker uid mapping: container uid 0 becomes an unprivileged subordinate uid on the host (100000 for the first entry), higher container uids map further into the 65,536-id range, and no container uid maps to host uid 0 inside the containeron the hostuid 0 (root)uid 100000unprivileged; first subordinate iduid 1000 (app)uid 101000still unprivilegeduid 65535uid 165535last id of the 65,536 range(no container uid)uid 0 (root)no mapping: an escape lands as aliceWhere the map comes from/etc/subuidalice:100000:65536 (start, count — the daemon runs as alice)dockerdrootlesskit + user namespace; slirp4netns or pasta for networking

The map comes from /etc/subuid and /etc/subgid, which need at least 65,536 subordinate ids for the user that will own the daemon, and from the newuidmap and newgidmap setuid helpers in the uidmap package. Local accounts created by useradd usually get a range automatically; accounts that come from LDAP or SSSD often do not, and the symptom is dockerd-rootless-setuptool.sh refusing to install. Provision the ranges at account creation for build hosts rather than debugging them one at a time.

/etc/subuid
# user:start:count (start must not overlap another user's range)
alice:100000:65536
# container uid 0 -> host uid 100000
# container uid 1000 -> host uid 101000

Install and point the CLI at the user daemon

With Docker installed from packages, dockerd-rootless-setuptool.sh install (in /usr/bin, or in docker-ce-rootless-extras if missing) runs as the ordinary user, creates a user-level systemd unit, and prints the two exports it needs. Without packages, curl -fsSL https://get.docker.com/rootless | sh does the same into ~/bin. The socket lives at $XDG_RUNTIME_DIR/docker.sock, and DOCKER_HOST must point there or the CLI keeps talking to the root daemon's socket, which is the most common way to believe rootless is working when it is not.

install-rootless.sh
# packages already installed: run as the ordinary user, not via sudo
dockerd-rootless-setuptool.sh install
# add to ~/.bashrc (the tool prints these for your uid)
export PATH=/usr/bin:$PATH
export DOCKER_HOST=unix:///run/user/1000/docker.sock
systemctl --user enable --now docker
sudo loginctl enable-linger alice # keep the user daemon running after logout
bash — prove the daemon and the containers are unprivileged
docker info --format "{{.SecurityOptions}} driver={{.Driver}} cgroup={{.CgroupDriver}}/v{{.CgroupVersion}}"
[name=seccomp,profile=builtin name=rootless name=cgroupns] driver=overlay2 cgroup=systemd/v2
docker run --rm alpine id
uid=0(root) gid=0(root) groups=0(root)
docker run -d --name t alpine sleep 300 && ps -o user,pid,cmd -p $(pgrep -f "sleep 300" | head -1)
USER PID CMD
100000 41822 sleep 300
root inside, subordinate uid 100000 outside

The docker info line carries the three things to verify: name=rootless in the security options, a storage driver the rootless daemon supports, and cgroup=systemd with version 2. A Cgroup Driver of none means resource flags such as --memory and --pids-limit are silently ignored, because rootless cgroups need cgroup v2 delegated through systemd.

The limitations that decide placement

Documented rootless limitations and what they mean in practice

LimitationConsequenceWay out
Storage: overlay2 needs kernel 5.11+; else fuse-overlayfs (4.18+), btrfs or vfsOld LTS kernels fall back to fuse-overlayfs, which is slowerInstall fuse-overlayfs, or upgrade the kernel
Resource limits need cgroup v2 + systemd--cpus, --memory, --pids-limit ignored on v1 hostsBoot with cgroup v2; check Cgroup Driver: systemd
Ports below 1024 are privilegeddocker run -p 80:80 fails with cannot expose privileged portnet.ipv4.ip_unprivileged_port_start=0, or CAP_NET_BIND_SERVICE on rootlesskit
User-mode network stack (slirp4netns or pasta)Lower throughput; -p does not propagate source IPs by defaultAccept it, enable source-IP propagation, or --net=host (Engine v29.5+)
No AppArmor, checkpoint, overlay networks or SCTP portsSwarm overlay and CRIU-based workflows do not applyKeep those workloads on rootful hosts
--cap-add only affects the container’s user namespaceOperations that need capabilities in the initial namespace still failThat is the point; design around it

Two of these have changed recently enough that older advice is wrong. docker run --net=host used to be namespaced inside RootlessKit, so a container listening on the host network was not reachable from the real host; since Engine v29.5 it shares the host namespace, which also bypasses the slow user-mode stack at the cost of network isolation. And the storage story is no longer "fuse-overlayfs or nothing": on kernel 5.11 and later the rootless daemon uses overlay2 natively. The --cap-add row is the one to read twice if you are used to Linux capabilities on rootful hosts: a capability granted inside the user namespace is real inside it and meaningless outside it.

Rootless protects the host, not the workload
An escape from a rootless container lands as an unprivileged host uid, which is the property you bought. It does nothing about a container that is itself compromised: seccomp, dropped capabilities, a read-only root filesystem and a non-root USER in the image are still the controls for that, and they are the same in both modes.
Go deeper in a courseDocker hardeningRootless mode, seccomp, capabilities and minimal production images.View course
Where rootless earns its friction
Good fit
CI build agents running untrusted Dockerfiles
Shared developer hosts
Anything where docker.sock access equalled root
Laptops and workstations
Keep rootful
Hosts that need host-network plugins
Swarm overlay networking
Workloads that need --privileged anyway
Kubernetes nodes (containerd runs as root)

Kubernetes reaches the same goal a different way: hostUsers: false in a Pod spec gives that Pod its own user namespace, and the feature is stable as of v1.36, so a cluster on a current release can apply the uid-mapping idea per Pod without changing the runtime. The decision is the same one as above, made per workload instead of per host: which processes must never be able to become root outside their own namespace.

Related posts

Quick reference