Rootless Docker

The daemon as an ordinary user: setup, networking, storage and limits.

Advanced17 min · lesson 14 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (1 files, 1 KB): rootless.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/rootless.tar.gz && tar -xzf rootless.tar.gz, which creates ~/lab/rootless/. SHA-256: 5759a7e73517cb68bad62a4f96f2e334bcdee2ab0e37edcc3f0214f3e8b1e05d
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab stops the system Docker daemon, sets a file capability on rootlesskit and adds a systemd drop-in. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and unpack the lesson files into ~/lab/rootless. To get a clean VM back, run ./setup/create-lab.sh --profile sec --recreate.

On a normal Docker host, look at a container process from the host side and you will find it owned by root. Here is a container started through the system daemon on the sec VM (the learner ubuntu is not in the docker group there, so it uses sudo):

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker pull -q alpine:3.22 sudo docker run -d --name lab-rootful alpine:3.22 sleep 300 >/dev/null ps -o user,pid,args -C sleep sudo docker exec lab-rootful cat /proc/self/attr/current sudo docker rm -f lab-rootful
docker.io/library/alpine:3.22 USER PID COMMAND root 2111 sleep 300 docker-default (enforce) lab-rootful

The sleep inside the container is host UID 0, and the only thing keeping it in check is the AppArmor profile docker-default plus seccomp and dropped capabilities. If anything gets past those, it lands as root on the host. Rootless mode changes the identity underneath. The daemon, containerd and every container run as an ordinary login account, inside a user namespace that RootlessKit sets up. Container root becomes your UID, and the other container UIDs become numbers from your subordinate range that own nothing on the host. This lesson installs it on the sec VM with Docker 29.8.2 and measures what works and what does not.

Prerequisites, and the Ubuntu 26.04 catch

Rootless needs newuidmap and newgidmap (package uidmap), at least 65,536 subordinate UIDs and GIDs for your account in /etc/subuid and /etc/subgid, cgroup v2 with systemd for resource limits, and a systemd user session (dbus-user-session). The Docker apt packages put the setup tool and RootlessKit in docker-ce-rootless-extras. The lab VM has all of them:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ grep ^ubuntu: /etc/subuid /etc/subgid command -v newuidmap newgidmap rootlesskit dockerd-rootless-setuptool.sh command -v slirp4netns || echo "slirp4netns: not installed"
/etc/subuid:ubuntu:100000:65536 /etc/subgid:ubuntu:100000:65536 /usr/bin/newuidmap /usr/bin/newgidmap /usr/bin/rootlesskit /usr/bin/dockerd-rootless-setuptool.sh slirp4netns: not installed

Ubuntu 24.04 and later add one more condition. The kernel setting kernel.apparmor_restrict_unprivileged_userns=1 means an unprivileged process that creates a user namespace gets no privileges inside it unless an AppArmor profile grants userns. Try it with unshare:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ cat /proc/sys/kernel/apparmor_restrict_unprivileged_userns unshare --user --map-root-user id
1 unshare: write failed /proc/self/uid_map: Operation not permitted

Creating the namespace is allowed, but the process is put under a restrictive profile and cannot write its own UID map. RootlessKit would hit the same wall without a profile of its own. The deb install path is covered, because Ubuntu's apparmor package ships a profile for /usr/bin/rootlesskit:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ cat /etc/apparmor.d/rootlesskit
# This profile allows everything and only exists to give the # application a name instead of having the label "unconfined" abi <abi/5.0>, include <tunables/global> profile rootlesskit /usr/bin/rootlesskit flags=(unconfined) { userns, @{exec_path} mr, # Site-specific additions and overrides. See local/README for details. include if exists <local/rootlesskit> }

flags=(unconfined) and the userns, rule are the important parts: the binary gets a name and permission to create user namespaces, and no other restriction. If you install the static binaries into ~/bin with the get.docker.com/rootless script, no packaged profile matches that path. Docker's troubleshooting page prescribes a profile for it, written as root and loaded with an AppArmor restart:

/etc/apparmor.d/home.<user>.bin.rootlesskit
abi <abi/4.0>,
include <tunables/global>
"/home/<user>/bin/rootlesskit" flags=(unconfined) {
userns,
include if exists <local/home.<user>.bin.rootlesskit>
}
# then: sudo systemctl restart apparmor.service

Install with dockerd-rootless-setuptool.sh

Docker's docs suggest disabling the system daemon first; otherwise the setup tool wants --force and you keep a root daemon running next to the rootless one. Then enable lingering, so systemd keeps your user manager (and with it the daemon) running when you are not logged in and starts it at boot:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo systemctl disable --now docker.service docker.socket sudo rm -f /var/run/docker.sock
Synchronizing state of docker.service with SysV service script with /usr/lib/systemd/systemd-sysv-install. Executing: /usr/lib/systemd/systemd-sysv-install disable docker Removed '/etc/systemd/system/sockets.target.wants/docker.socket'. Removed '/etc/systemd/system/multi-user.target.wants/docker.service'. Disabling 'docker.service', but its triggering units are still active: docker.socket
$ sudo loginctl enable-linger ubuntu loginctl show-user ubuntu -p Linger -p RuntimePath
RuntimePath=/run/user/1000 Linger=yes

RuntimePath is the directory systemd gives your session, /run/user/1000. A real login (multipass shell, SSH, the console) goes through pam_systemd, which exports it as XDG_RUNTIME_DIR and starts the user bus. Switching accounts with sudo -iu or su skips that, and systemctl --user then fails with "Failed to connect to bus"; log in as the account directly instead. Now run the installer as yourself, without sudo. The sed only strips the colour codes the script prints:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ dockerd-rootless-setuptool.sh install 2>&1 | sed "s/\x1b\[[0-9;]*m//g"
[INFO] Creating /home/ubuntu/.config/systemd/user/docker.service [INFO] starting systemd service docker.service + systemctl --user start docker.service + sleep 3 + systemctl --user --no-pager --full status docker.service ● docker.service - Docker Application Container Engine (Rootless) Loaded: loaded (/home/ubuntu/.config/systemd/user/docker.service; disabled; preset: enabled) Active: active (running) since Thu 2026-10-08 03:08:36 IST; 3s ago ... ├─2647 rootlesskit --state-dir=/run/user/1000/dockerd-rootless --net=gvisor-tap-vsock --mtu=65520 --slirp4netns-sandbox=auto --slirp4netns-seccomp=auto --disable-host-loopback --port-driver=builtin --copy-up=/etc --copy-up=/run --propagation=rslave --detach-netns /usr/bin/dockerd-rootless.sh ├─2655 /proc/self/exe --state-dir=/run/user/1000/dockerd-rootless --net=gvisor-tap-vsock --mtu=65520 --slirp4netns-sandbox=auto --slirp4netns-seccomp=auto --disable-host-loopback --port-driver=builtin --copy-up=/etc --copy-up=/run --propagation=rslave --detach-netns /usr/bin/dockerd-rootless.sh ├─2693 dockerd └─2713 /usr/bin/containerd --config /run/user/1000/docker/containerd/containerd.toml ... rootlesskit: Version: 3.1.0 ApiVersion: 1.1.2 NetworkDriver: gvisor-tap-vsock PortDriver: builtin StateDir: /run/user/1000/dockerd-rootless + systemctl --user enable docker.service Created symlink '/home/ubuntu/.config/systemd/user/default.target.wants/docker.service' → '/home/ubuntu/.config/systemd/user/docker.service'. [INFO] Installed docker.service successfully. [INFO] To control docker.service, run: `systemctl --user (start|stop|restart) docker.service` [INFO] To run docker.service on system startup, run: `sudo loginctl enable-linger ubuntu` [INFO] Creating CLI context "rootless" Successfully created context "rootless" [INFO] Using CLI context "rootless" Current context is now "rootless" [INFO] Make sure the following environment variable(s) are set (or add them to ~/.bashrc): export PATH=/usr/bin:$PATH [INFO] Some applications may require the following environment variable too: export DOCKER_HOST=unix:///run/user/1000/docker.sock

The tool wrote a user unit at ~/.config/systemd/user/docker.service, started it, and enabled it for default.target. The process tree shows the chain: rootlesskit creates the user, mount and network namespaces, and dockerd and a private containerd run inside them. The rootlesskit section of docker version names the network driver, gvisor-tap-vsock, and the port driver, builtin. The tool also created a CLI context and switched to it:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker context ls
NAME DESCRIPTION DOCKER ENDPOINT ERROR default Current DOCKER_HOST based configuration unix:///var/run/docker.sock rootless * Rootless mode unix:///run/user/1000/docker.sock

A context is a named daemon endpoint. default still points at /var/run/docker.sock; rootless points at your socket in /run/user/1000. docker context use default switches back, and scripts that ignore contexts can set DOCKER_HOST=unix:///run/user/1000/docker.sock instead, as the installer's last lines suggest. Manage the daemon with systemctl --user start|stop|restart docker and read its logs with journalctl --user -u docker. Running rootless Docker as a system-wide unit with User= is not supported.

What docker info reports

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker info
... Server: Containers: 0 Running: 0 Paused: 0 Stopped: 0 Images: 0 Server Version: 29.8.2 Storage Driver: overlayfs driver-type: io.containerd.snapshotter.v1 Logging Driver: json-file Cgroup Driver: systemd Cgroup Version: 2 ... Security Options: seccomp Profile: builtin rootless cgroupns ... Docker Root Dir: /home/ubuntu/.local/share/docker ... WARNING: No cpuset support WARNING: No io.weight support WARNING: No io.weight (per device) support WARNING: No io.max (rbps) support WARNING: No io.max (wbps) support WARNING: No io.max (riops) support WARNING: No io.max (wiops) support

Four lines matter. Security Options lists rootless and no apparmor: AppArmor is not available to rootless containers. Docker Root Dir is ~/.local/share/docker, a separate store, so images you pulled with the system daemon are not here. Storage Driver: overlayfs with driver-type: io.containerd.snapshotter.v1 means the rootless daemon uses the containerd image store with native kernel overlayfs, as the rootful daemon does on Docker 29. The docs' list of rootless storage drivers (overlay2 on kernel 5.11 or later, fuse-overlayfs, btrfs, vfs) applies to the older graph drivers; on this kernel 7.0 host nothing goes through FUSE. The WARNING lines say that the cpuset and io controllers are missing; the cgroup section below explains why.

The processes confirm who owns the daemon:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ ps -ww -o user,pid,args -C rootlesskit,dockerd,containerd
USER PID COMMAND root 1121 /usr/bin/containerd ubuntu 2647 rootlesskit --state-dir=/run/user/1000/dockerd-rootless --net=gvisor-tap-vsock --mtu=65520 --slirp4netns-sandbox=auto --slirp4netns-seccomp=auto --disable-host-loopback --port-driver=builtin --copy-up=/etc --copy-up=/run --propagation=rslave --detach-netns /usr/bin/dockerd-rootless.sh ubuntu 2693 dockerd ubuntu 2713 /usr/bin/containerd --config /run/user/1000/docker/containerd/containerd.toml

Everything rootless belongs to ubuntu. The root containerd is the system containerd.service from the containerd.io package, which disabling docker.service does not stop. It is idle now; stop it too if you want no root container service on the host at all.

How container UIDs map to the host

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker pull -q alpine:3.22 docker pull -q nginx:1.30-alpine docker pull -q python:3.14-slim
docker.io/library/alpine:3.22 docker.io/library/nginx:1.30-alpine docker.io/library/python:3.14-slim
$ docker run -d --name lab-root alpine:3.22 sleep 300 >/dev/null docker run -d --name lab-u101 --user 101 alpine:3.22 sleep 300 >/dev/null ps -o user,uid,pid,args -C sleep docker exec lab-root cat /proc/self/uid_map
USER UID PID COMMAND ubuntu 1000 3298 sleep 300 100100 100100 3363 sleep 300 0 1000 1 1 100000 65536

Container root shows up as ubuntu (UID 1000), and container UID 101 as host UID 100100. The uid_map explains both lines. Container UID 0 maps to host 1000 for a range of one, and container UIDs from 1 upward map to 100000 and onwards. So container UID n is host 100000 + (n - 1). This is different from userns-remap, where container UID 0 maps to the first subordinate UID itself (the lesson "User-namespace remapping" shows it). Bind mounts follow the same rule:

ubuntu@secopslog-docker-sec:~/lab/rootless · Docker 29.8.2
$ mkdir -p data docker run --rm -v "$PWD/data:/data" alpine:3.22 sh -c "touch /data/by-root && touch /data/by-101 && chown 101:101 /data/by-101" ls -ln data
total 0 -rw-r--r-- 1 100100 100100 0 Oct 8 03:09 by-101 -rw-r--r-- 1 1000 1000 0 Oct 8 03:09 by-root

Files that container root writes in your home directory belong to you, which is convenient. Files written as another container user get a subordinate UID, and you can only remove them through the namespace (rootlesskit rm -rf ..., used in the cleanup). Host paths that only root may write stay closed, because "root" in the container is just UID 1000:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run --rm -v /usr/local:/host-usr alpine:3.22 touch /host-usr/lab-rootless-test
touch: /host-usr/lab-rootless-test: Permission denied
$ docker run --rm -v /etc:/host-etc alpine:3.22 sh -c "touch /host-etc/lab-rootless-test && ls /host-etc/lab-rootless-test" ls -l /etc/lab-rootless-test
/host-etc/lab-rootless-test ls: cannot access '/etc/lab-rootless-test': No such file or directory

The second result catches people out. Writing through a bind mount of /etc appears to succeed, yet the file is not in the host's /etc. RootlessKit runs with --copy-up=/etc (visible in the process list above), so the daemon's mount namespace has its own writable copy of /etc and that copy is what got mounted. Do not use a bind mount of /etc under rootless to change host configuration; it changes nothing on the host. AppArmor confinement is gone as well:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker exec lab-root cat /proc/self/attr/current
runc (unconfined)

The rootful container ran under docker-default (enforce); this one is labelled runc (unconfined). Seccomp, dropped capabilities and the user namespace still apply. AppArmor no longer does.

Resource limits and cgroup v2 delegation

Limits for rootless containers use the cgroup subtree that systemd delegates to your user manager. The controllers you get are listed in its cgroup.controllers:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ cat /sys/fs/cgroup/user.slice/user-$(id -u).slice/user@$(id -u).service/cgroup.controllers
cpu memory pids
$ docker run --rm --memory 64m --cpus 0.5 --pids-limit 50 alpine:3.22 \ cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/cpu.max /sys/fs/cgroup/pids.max
67108864 50000 100000 50
$ docker run --rm --cpuset-cpus 0 --device-read-bps /dev/vda:1mb alpine:3.22 cat /sys/fs/cgroup/cgroup.controllers
WARNING: Your kernel does not support cpuset or the cgroup is not mounted. Cpuset discarded. WARNING: Your kernel does not support BPS Block I/O read limit or the cgroup is not mounted. Block I/O BPS read limit discarded. cpu memory pids

Ubuntu 26.04 delegates cpu, memory and pids by default (Docker's docs say many distributions delegate only memory and pids). So --memory, --cpus and --pids-limit reach the container's cgroup files, while --cpuset-cpus and the block I/O flags are dropped with a warning and the container still starts. That is the risky part: a limit you asked for silently does not exist. The fix is a systemd drop-in for every user manager, from the lesson files:

/etc/systemd/system/user@.service.d/delegate.conf
[Service]
Delegate=cpu cpuset io memory pids
ubuntu@secopslog-docker-sec:~/lab/rootless · Docker 29.8.2
$ cat delegate.conf sudo mkdir -p /etc/systemd/system/user@.service.d sudo cp delegate.conf /etc/systemd/system/user@.service.d/delegate.conf sudo systemctl daemon-reload cat /sys/fs/cgroup/user.slice/user-$(id -u).slice/user@$(id -u).service/cgroup.controllers systemctl --user restart docker
[Service] Delegate=cpu cpuset io memory pids cpuset cpu io memory pids
ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run --rm --cpuset-cpus 0 --device-read-bps /dev/vda:1mb alpine:3.22 \ cat /sys/fs/cgroup/cpuset.cpus.effective /sys/fs/cgroup/io.max
0 253:0 rbps=1048576 wbps=max riops=max wiops=max

After daemon-reload the user manager has all five controllers, and after the daemon restart the cpuset and read-bandwidth limits land in the container's cgroup (253:0 is the major:minor number of /dev/vda). cpuset delegation needs systemd 244 or later; this VM runs 259. If docker info ever shows Cgroup Driver: none, no cgroup flag works at all. "cgroups v2 and resource containment" covers the files themselves.

Networking on Docker 29

A rootless daemon cannot create veth pairs or iptables rules in the host network namespace. It lives in its own network namespace, and a user-mode TCP/IP stack moves its traffic. Since 29.5, gvisor-tap-vsock is the default network driver when slirp4netns is not installed, and the prerequisite check above showed it is not on this VM; where slirp4netns is installed it stays the default. pasta is an experimental alternative. Inside, containers look normal:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run --rm alpine:3.22 sh -c "ip -o -4 addr show eth0; ip route; ping -c1 -W3 1.1.1.1 | tail -2"
2: eth0 inet 172.17.0.2/16 brd 172.17.255.255 scope global eth0\ valid_lft forever preferred_lft forever default via 172.17.0.1 dev eth0 172.17.0.0/16 dev eth0 scope link src 172.17.0.2 1 packets transmitted, 1 packets received, 0% packet loss round-trip min/avg/max = 1.086/1.086/1.086 ms
$ docker run -d --name lab-web -p 8080:80 nginx:1.30-alpine >/dev/null; sleep 2 curl -sI http://127.0.0.1:8080/ | head -1 ss -ltnp | grep :8080
HTTP/1.1 200 OK LISTEN 0 4096 0.0.0.0:8080 0.0.0.0:* users:(("rootlesskit",pid=4276,fd=16)) LISTEN 0 4096 [::]:8080 [::]:* users:(("rootlesskit",pid=4276,fd=17))
$ IP=$(docker inspect -f "{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}" lab-web) echo "$IP" curl -sS -m 3 -o /dev/null "http://$IP/"
172.17.0.2 curl: (28) Connection timed out after 3080 milliseconds

The container has the usual bridge address and reaches the internet. ping works because Ubuntu 26.04 already sets net.ipv4.ping_group_range to 0 2147483647; on hosts where it is 1 0, set that range as the docs describe. -p 8080:80 works too, and ss shows who listens on the host: rootlesskit, through its builtin port driver, not docker-proxy. The container's IP, 172.17.0.2, exists only inside RootlessKit's namespace, so curl to it times out. Publish ports instead of relying on container IPs. By default the published port also does not show clients' real source IPs; RootlessKit 3 can preserve them once you set "userland-proxy": false in ~/.config/docker/daemon.json.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run -d --name lab-web80 -p 80:80 nginx:1.30-alpine; echo "exit=$?" docker rm -f lab-web80 >/dev/null
9494d851e7f37cee57bd208ba69c159b2feeaa174465035c9a7b1944d96f73a4 docker: Error response from daemon: failed to set up container networking: driver failed programming external connectivity on endpoint lab-web80 (1ef66d56233bc29eb66d01de8c237f551a0054bf0f559cbfb60494b6202bdd6b): error while calling RootlessKit PortManager.AddPort(): cannot expose privileged port 80, you can add 'net.ipv4.ip_unprivileged_port_start=80' to /etc/sysctl.conf (currently 1024), or set CAP_NET_BIND_SERVICE on rootlesskit binary, or choose a larger port number (>= 1024): listen tcp4 0.0.0.0:80: bind: permission denied Run 'docker run --help' for more information exit=126
$ docker run -d --name lab-web80 --cap-add NET_BIND_SERVICE -p 80:80 nginx:1.30-alpine; echo "exit=$?" docker rm -f lab-web80 >/dev/null
4b7866679cbfcc63c686e5eead333f9467facbb3218ee96567b5ee46021b6809 docker: Error response from daemon: failed to set up container networking: driver failed programming external connectivity on endpoint lab-web80 (6ca5e538b449fe969a1385a79de3470ad7cb0a3908b58054ad011ec163c4eace): error while calling RootlessKit PortManager.AddPort(): cannot expose privileged port 80, you can add 'net.ipv4.ip_unprivileged_port_start=80' to /etc/sysctl.conf (currently 1024), or set CAP_NET_BIND_SERVICE on rootlesskit binary, or choose a larger port number (>= 1024): listen tcp4 0.0.0.0:80: bind: permission denied Run 'docker run --help' for more information exit=126

Both runs print the error and exit=126. Docker creates the container before the port fails, so each command removes it again. The process refused is RootlessKit, which tries to listen on host port 80 without CAP_NET_BIND_SERVICE, so adding the capability to the container changes nothing. The error lists the real options. In order of how much they change on the host: publish a high port and put a proxy or load balancer in front; give only rootlesskit the capability; or lower net.ipv4.ip_unprivileged_port_start, which lets every process of every user bind low ports.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo setcap cap_net_bind_service=ep /usr/bin/rootlesskit /usr/sbin/getcap /usr/bin/rootlesskit systemctl --user restart docker
/usr/bin/rootlesskit cap_net_bind_service=ep
$ docker run -d --name lab-web80 -p 80:80 nginx:1.30-alpine >/dev/null; sleep 2 curl -sI http://127.0.0.1/ | head -1
HTTP/1.1 200 OK
$ docker rm -f lab-web80 lab-web sudo setcap -r /usr/bin/rootlesskit /usr/sbin/getcap /usr/bin/rootlesskit
lab-web80 lab-web

With the file capability on /usr/bin/rootlesskit and a daemon restart, port 80 works. setcap -r takes it away again, and the empty getcap output confirms no capability is left; the running RootlessKit keeps it until the next restart. Only that binary gains the capability, but /usr/bin/rootlesskit is executable by every account, so on a shared host anyone can use it to bind low ports; restrict who may run it (for example a dedicated group and mode 750) before you rely on this. A package upgrade replaces the binary and drops the capability, so put the setcap in configuration management.

Until 29.5, --net=host meant RootlessKit's namespace, so a container's listening ports were invisible to the real host. Docker 29.5 added proper support, and 29.7 fixed the cgroup mount for such containers:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run -d --name lab-host --net=host python:3.14-slim python3 -m http.server 8081 >/dev/null; sleep 3 ss -ltnp | grep :8081 curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8081/
LISTEN 0 5 0.0.0.0:8081 0.0.0.0:* users:(("python3",pid=5918,fd=3)) 200

The listener sits in the host's namespace and belongs to python3, owned by your account, and the host reaches it directly without the user-mode stack. The docs suggest it as a workaround for slow rootless networking. The cost is the same as on a rootful host: the container shares the host's network stack (see "Host namespace sharing as attack surface"). Ports below 1024 still need the sysctl, because the container is not RootlessKit.

What still does not work, and what rootless does not protect

From the docs and this lab: no AppArmor, no docker checkpoint, no overlay networks (so no Swarm), no SCTP port publishing, container IPs that cannot be reached from the host, --cap-add that only works for resources owned by the container's user namespace, and a data root that cannot be on NFS. Each user has their own daemon, image store and socket, so two engineers on one build host pull every image twice. Host-level monitoring and log agents that read /var/lib/docker or the root socket see none of these containers: each user's data lives in ~/.local/share/docker and the daemon logs to journalctl --user, so point the agents at each user's socket and data root.

The socket at /run/user/1000/docker.sock controls your whole account. Whoever can talk to it can mount your home directory, read your SSH keys and cloud credentials, and run anything as you. Do not mount it into containers and do not expose it on TCP without mutual TLS, for the same reasons as the root socket in "The Docker socket and daemon hardening". Containers still share the host kernel, so a kernel bug can still matter. Keep non-root images, dropped capabilities and seccomp on top.

Clean up

Uninstall the user service, remove the rootless data through the namespace (plain rm fails on files owned by subordinate UIDs), then undo the host changes and bring the system daemon back:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker rm -f lab-root lab-u101 lab-host dockerd-rootless-setuptool.sh uninstall 2>&1 | sed "s/\x1b\[[0-9;]*m//g" rootlesskit rm -rf ~/.local/share/docker ~/lab/rootless/data
lab-root lab-u101 lab-host + systemctl --user stop docker.service + systemctl --user disable docker.service Removed '/home/ubuntu/.config/systemd/user/default.target.wants/docker.service'. [INFO] Uninstalled docker.service [INFO] Deleted CLI context "rootless" Current context is now "default" [INFO] Configured CLI to use the "default" context. [INFO] [INFO] Make sure to unset or update the environment PATH, DOCKER_HOST, and DOCKER_CONTEXT environment variables if you have added them to `~/.bashrc`. [INFO] This uninstallation tool does NOT remove Docker binaries and data. [INFO] To remove data, run: `/usr/bin/rootlesskit rm -rf /home/ubuntu/.local/share/docker`
$ sudo loginctl disable-linger ubuntu sudo rm /etc/systemd/system/user@.service.d/delegate.conf sudo systemctl daemon-reload sudo systemctl enable --now docker.service docker.socket 2>/dev/null sudo docker info --format "{{.ServerVersion}} {{.Driver}}"
29.8.2 overlayfs

The uninstaller also deleted the rootless context and switched the CLI back to default. If you added DOCKER_HOST to ~/.bashrc, remove it.

Quick check
01A team moves a service to rootless Docker. docker run -p 443:8443 app fails with "cannot expose privileged port 443", and --cap-add NET_BIND_SERVICE makes no difference. Which change fixes it while changing the least on the host?
Incorrect — Privileged mode widens what the container can do inside its own user namespace. The process refused is RootlessKit on the host, so this does not help.
Correct — RootlessKit is the process that listens on the host port, so it needs the capability. Only that binary gains it, though anyone who can run it can use it, so restrict its permissions on a shared host.
Incorrect — This works, but it lets every process of every user bind low ports, which changes much more than the setcap does.
Incorrect — That setting matters for source IP propagation. It does not let RootlessKit bind port 443.
02On a rootless host, /etc/subuid has ubuntu:100000:65536. A container runs as --user 1000. Which host UID owns its files on a bind mount?
Incorrect — Only container UID 0 maps to your own UID (1000 here). Every other container UID goes through the subordinate range.
Incorrect — That is the userns-remap rule. In rootless mode the range starts at container UID 1, not 0, so there is an offset of one.
Incorrect — No UID is passed through unchanged in rootless mode. Host UID 0 is not even in the map.
Correct — Container UID n (n >= 1) maps to subuid + (n - 1), so 1000 becomes 100999.
03docker run --cpuset-cpus 0 app on a fresh rootless install prints "Cpuset discarded" and the container starts. What is going on, and what fixes it?
Correct — The rootless daemon can only use the controllers delegated to the user manager. Ubuntu 26.04 delegates cpu, memory and pids; the drop-in adds cpuset and io.
Incorrect — The warning wording suggests that, but the rootful daemon on the same kernel uses cpuset. After the drop-in, the rootless container gets cpuset.cpus.effective = 0.
Incorrect — --cpus worked before any change (cpu.max showed 50000 100000), and cpuset works after delegation.
Incorrect — CPU pinning goes through the cgroup, not a capability inside the container.

Try this

Work through “Clean up” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from rootless docker, keep “Clean up”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related