Rootless Docker and why containers shouldn't run as root
Run the Docker daemon and your containers as a non-root user so a container escape lands unprivileged, plus the subuid, cgroups, and storage gotchas.
Whoever can talk to /var/run/docker.sock on a default install is root on that host, because dockerd runs as root and will happily bind-mount / into a container for anyone the socket permits. A kernel bug or a runtime escape ends in the same place. Rootless mode moves the daemon itself into an unprivileged user account: dockerd and every container it starts run inside a user namespace owned by, say, alice, and a process that breaks out of a container is alice, or one of her subordinate ids, on the host.
What you give up is a list, not a feeling, and the Docker docs publish it: specific storage drivers, cgroup v2 as a requirement for resource limits, no privileged ports without a sysctl, and a user-mode network stack that is slower than the kernel's. Whether that trade is worth it depends on the host's job. For a CI builder or a shared development box it usually is; for a node that needs host networking plugins it usually is not.
Where container root lands on the host
Container uid 0 is host uid 100000 here because that is the first subordinate id in alice’s /etc/subuid entry; the 65,536-id range is the minimum the docs require. No container uid maps to host uid 0, which is the whole security argument in one row. Simplified: gids follow the same scheme through /etc/subgid.
The map comes from /etc/subuid and /etc/subgid, which need at least 65,536 subordinate ids for the user that will own the daemon, and from the newuidmap and newgidmap setuid helpers in the uidmap package. Local accounts created by useradd usually get a range automatically; accounts that come from LDAP or SSSD often do not, and the symptom is dockerd-rootless-setuptool.sh refusing to install. Provision the ranges at account creation for build hosts rather than debugging them one at a time.
# user:start:count (start must not overlap another user's range)alice:100000:65536# container uid 0 -> host uid 100000# container uid 1000 -> host uid 101000
Install and point the CLI at the user daemon
With Docker installed from packages, dockerd-rootless-setuptool.sh install (in /usr/bin, or in docker-ce-rootless-extras if missing) runs as the ordinary user, creates a user-level systemd unit, and prints the two exports it needs. Without packages, curl -fsSL https://get.docker.com/rootless | sh does the same into ~/bin. The socket lives at $XDG_RUNTIME_DIR/docker.sock, and DOCKER_HOST must point there or the CLI keeps talking to the root daemon's socket, which is the most common way to believe rootless is working when it is not.
# packages already installed: run as the ordinary user, not via sudodockerd-rootless-setuptool.sh install# add to ~/.bashrc (the tool prints these for your uid)export PATH=/usr/bin:$PATHexport DOCKER_HOST=unix:///run/user/1000/docker.socksystemctl --user enable --now dockersudo loginctl enable-linger alice # keep the user daemon running after logout
docker info --format "{{.SecurityOptions}} driver={{.Driver}} cgroup={{.CgroupDriver}}/v{{.CgroupVersion}}"[name=seccomp,profile=builtin name=rootless name=cgroupns] driver=overlay2 cgroup=systemd/v2docker run --rm alpine iduid=0(root) gid=0(root) groups=0(root)docker run -d --name t alpine sleep 300 && ps -o user,pid,cmd -p $(pgrep -f "sleep 300" | head -1)USER PID CMD100000 41822 sleep 300root inside, subordinate uid 100000 outsideThe docker info line carries the three things to verify: name=rootless in the security options, a storage driver the rootless daemon supports, and cgroup=systemd with version 2. A Cgroup Driver of none means resource flags such as --memory and --pids-limit are silently ignored, because rootless cgroups need cgroup v2 delegated through systemd.
The limitations that decide placement
Documented rootless limitations and what they mean in practice
| Limitation | Consequence | Way out |
|---|---|---|
Storage: overlay2 needs kernel 5.11+; else fuse-overlayfs (4.18+), btrfs or vfs | Old LTS kernels fall back to fuse-overlayfs, which is slower | Install fuse-overlayfs, or upgrade the kernel |
| Resource limits need cgroup v2 + systemd | --cpus, --memory, --pids-limit ignored on v1 hosts | Boot with cgroup v2; check Cgroup Driver: systemd |
| Ports below 1024 are privileged | docker run -p 80:80 fails with cannot expose privileged port | net.ipv4.ip_unprivileged_port_start=0, or CAP_NET_BIND_SERVICE on rootlesskit |
| User-mode network stack (slirp4netns or pasta) | Lower throughput; -p does not propagate source IPs by default | Accept it, enable source-IP propagation, or --net=host (Engine v29.5+) |
| No AppArmor, checkpoint, overlay networks or SCTP ports | Swarm overlay and CRIU-based workflows do not apply | Keep those workloads on rootful hosts |
--cap-add only affects the container’s user namespace | Operations that need capabilities in the initial namespace still fail | That is the point; design around it |
Two of these have changed recently enough that older advice is wrong. docker run --net=host used to be namespaced inside RootlessKit, so a container listening on the host network was not reachable from the real host; since Engine v29.5 it shares the host namespace, which also bypasses the slow user-mode stack at the cost of network isolation. And the storage story is no longer "fuse-overlayfs or nothing": on kernel 5.11 and later the rootless daemon uses overlay2 natively. The --cap-add row is the one to read twice if you are used to Linux capabilities on rootful hosts: a capability granted inside the user namespace is real inside it and meaningless outside it.
Kubernetes reaches the same goal a different way: hostUsers: false in a Pod spec gives that Pod its own user namespace, and the feature is stable as of v1.36, so a cluster on a current release can apply the uid-mapping idea per Pod without changing the runtime. The decision is the same one as above, made per workload instead of per host: which processes must never be able to become root outside their own namespace.