Rootless Docker
The daemon as an ordinary user: setup, networking, storage and limits.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/rootless.tar.gz && tar -xzf rootless.tar.gz, which creates ~/lab/rootless/. SHA-256: 5759a7e73517cb68bad62a4f96f2e334bcdee2ab0e37edcc3f0214f3e8b1e05dsecopslog-docker-sec). The lab stops the system Docker daemon, sets a file capability on rootlesskit and adds a systemd drop-in. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima), and unpack the lesson files into ~/lab/rootless. To get a clean VM back, run ./setup/create-lab.sh --profile sec --recreate.On a normal Docker host, look at a container process from the host side and you will find it owned by root. Here is a container started through the system daemon on the sec VM (the learner ubuntu is not in the docker group there, so it uses sudo):
The sleep inside the container is host UID 0, and the only thing keeping it in check is the AppArmor profile docker-default plus seccomp and dropped capabilities. If anything gets past those, it lands as root on the host. Rootless mode changes the identity underneath. The daemon, containerd and every container run as an ordinary login account, inside a user namespace that RootlessKit sets up. Container root becomes your UID, and the other container UIDs become numbers from your subordinate range that own nothing on the host. This lesson installs it on the sec VM with Docker 29.8.2 and measures what works and what does not.
Prerequisites, and the Ubuntu 26.04 catch
Rootless needs newuidmap and newgidmap (package uidmap), at least 65,536 subordinate UIDs and GIDs for your account in /etc/subuid and /etc/subgid, cgroup v2 with systemd for resource limits, and a systemd user session (dbus-user-session). The Docker apt packages put the setup tool and RootlessKit in docker-ce-rootless-extras. The lab VM has all of them:
Ubuntu 24.04 and later add one more condition. The kernel setting kernel.apparmor_restrict_unprivileged_userns=1 means an unprivileged process that creates a user namespace gets no privileges inside it unless an AppArmor profile grants userns. Try it with unshare:
Creating the namespace is allowed, but the process is put under a restrictive profile and cannot write its own UID map. RootlessKit would hit the same wall without a profile of its own. The deb install path is covered, because Ubuntu's apparmor package ships a profile for /usr/bin/rootlesskit:
flags=(unconfined) and the userns, rule are the important parts: the binary gets a name and permission to create user namespaces, and no other restriction. If you install the static binaries into ~/bin with the get.docker.com/rootless script, no packaged profile matches that path. Docker's troubleshooting page prescribes a profile for it, written as root and loaded with an AppArmor restart:
abi <abi/4.0>,include <tunables/global>"/home/<user>/bin/rootlesskit" flags=(unconfined) {userns,include if exists <local/home.<user>.bin.rootlesskit>}# then: sudo systemctl restart apparmor.service
Install with dockerd-rootless-setuptool.sh
Docker's docs suggest disabling the system daemon first; otherwise the setup tool wants --force and you keep a root daemon running next to the rootless one. Then enable lingering, so systemd keeps your user manager (and with it the daemon) running when you are not logged in and starts it at boot:
RuntimePath is the directory systemd gives your session, /run/user/1000. A real login (multipass shell, SSH, the console) goes through pam_systemd, which exports it as XDG_RUNTIME_DIR and starts the user bus. Switching accounts with sudo -iu or su skips that, and systemctl --user then fails with "Failed to connect to bus"; log in as the account directly instead. Now run the installer as yourself, without sudo. The sed only strips the colour codes the script prints:
The tool wrote a user unit at ~/.config/systemd/user/docker.service, started it, and enabled it for default.target. The process tree shows the chain: rootlesskit creates the user, mount and network namespaces, and dockerd and a private containerd run inside them. The rootlesskit section of docker version names the network driver, gvisor-tap-vsock, and the port driver, builtin. The tool also created a CLI context and switched to it:
A context is a named daemon endpoint. default still points at /var/run/docker.sock; rootless points at your socket in /run/user/1000. docker context use default switches back, and scripts that ignore contexts can set DOCKER_HOST=unix:///run/user/1000/docker.sock instead, as the installer's last lines suggest. Manage the daemon with systemctl --user start|stop|restart docker and read its logs with journalctl --user -u docker. Running rootless Docker as a system-wide unit with User= is not supported.
What docker info reports
Four lines matter. Security Options lists rootless and no apparmor: AppArmor is not available to rootless containers. Docker Root Dir is ~/.local/share/docker, a separate store, so images you pulled with the system daemon are not here. Storage Driver: overlayfs with driver-type: io.containerd.snapshotter.v1 means the rootless daemon uses the containerd image store with native kernel overlayfs, as the rootful daemon does on Docker 29. The docs' list of rootless storage drivers (overlay2 on kernel 5.11 or later, fuse-overlayfs, btrfs, vfs) applies to the older graph drivers; on this kernel 7.0 host nothing goes through FUSE. The WARNING lines say that the cpuset and io controllers are missing; the cgroup section below explains why.
The processes confirm who owns the daemon:
Everything rootless belongs to ubuntu. The root containerd is the system containerd.service from the containerd.io package, which disabling docker.service does not stop. It is idle now; stop it too if you want no root container service on the host at all.
How container UIDs map to the host
Container root shows up as ubuntu (UID 1000), and container UID 101 as host UID 100100. The uid_map explains both lines. Container UID 0 maps to host 1000 for a range of one, and container UIDs from 1 upward map to 100000 and onwards. So container UID n is host 100000 + (n - 1). This is different from userns-remap, where container UID 0 maps to the first subordinate UID itself (the lesson "User-namespace remapping" shows it). Bind mounts follow the same rule:
Files that container root writes in your home directory belong to you, which is convenient. Files written as another container user get a subordinate UID, and you can only remove them through the namespace (rootlesskit rm -rf ..., used in the cleanup). Host paths that only root may write stay closed, because "root" in the container is just UID 1000:
The second result catches people out. Writing through a bind mount of /etc appears to succeed, yet the file is not in the host's /etc. RootlessKit runs with --copy-up=/etc (visible in the process list above), so the daemon's mount namespace has its own writable copy of /etc and that copy is what got mounted. Do not use a bind mount of /etc under rootless to change host configuration; it changes nothing on the host. AppArmor confinement is gone as well:
The rootful container ran under docker-default (enforce); this one is labelled runc (unconfined). Seccomp, dropped capabilities and the user namespace still apply. AppArmor no longer does.
Resource limits and cgroup v2 delegation
Limits for rootless containers use the cgroup subtree that systemd delegates to your user manager. The controllers you get are listed in its cgroup.controllers:
Ubuntu 26.04 delegates cpu, memory and pids by default (Docker's docs say many distributions delegate only memory and pids). So --memory, --cpus and --pids-limit reach the container's cgroup files, while --cpuset-cpus and the block I/O flags are dropped with a warning and the container still starts. That is the risky part: a limit you asked for silently does not exist. The fix is a systemd drop-in for every user manager, from the lesson files:
[Service]Delegate=cpu cpuset io memory pids
After daemon-reload the user manager has all five controllers, and after the daemon restart the cpuset and read-bandwidth limits land in the container's cgroup (253:0 is the major:minor number of /dev/vda). cpuset delegation needs systemd 244 or later; this VM runs 259. If docker info ever shows Cgroup Driver: none, no cgroup flag works at all. "cgroups v2 and resource containment" covers the files themselves.
Networking on Docker 29
A rootless daemon cannot create veth pairs or iptables rules in the host network namespace. It lives in its own network namespace, and a user-mode TCP/IP stack moves its traffic. Since 29.5, gvisor-tap-vsock is the default network driver when slirp4netns is not installed, and the prerequisite check above showed it is not on this VM; where slirp4netns is installed it stays the default. pasta is an experimental alternative. Inside, containers look normal:
The container has the usual bridge address and reaches the internet. ping works because Ubuntu 26.04 already sets net.ipv4.ping_group_range to 0 2147483647; on hosts where it is 1 0, set that range as the docs describe. -p 8080:80 works too, and ss shows who listens on the host: rootlesskit, through its builtin port driver, not docker-proxy. The container's IP, 172.17.0.2, exists only inside RootlessKit's namespace, so curl to it times out. Publish ports instead of relying on container IPs. By default the published port also does not show clients' real source IPs; RootlessKit 3 can preserve them once you set "userland-proxy": false in ~/.config/docker/daemon.json.
Both runs print the error and exit=126. Docker creates the container before the port fails, so each command removes it again. The process refused is RootlessKit, which tries to listen on host port 80 without CAP_NET_BIND_SERVICE, so adding the capability to the container changes nothing. The error lists the real options. In order of how much they change on the host: publish a high port and put a proxy or load balancer in front; give only rootlesskit the capability; or lower net.ipv4.ip_unprivileged_port_start, which lets every process of every user bind low ports.
With the file capability on /usr/bin/rootlesskit and a daemon restart, port 80 works. setcap -r takes it away again, and the empty getcap output confirms no capability is left; the running RootlessKit keeps it until the next restart. Only that binary gains the capability, but /usr/bin/rootlesskit is executable by every account, so on a shared host anyone can use it to bind low ports; restrict who may run it (for example a dedicated group and mode 750) before you rely on this. A package upgrade replaces the binary and drops the capability, so put the setcap in configuration management.
Until 29.5, --net=host meant RootlessKit's namespace, so a container's listening ports were invisible to the real host. Docker 29.5 added proper support, and 29.7 fixed the cgroup mount for such containers:
The listener sits in the host's namespace and belongs to python3, owned by your account, and the host reaches it directly without the user-mode stack. The docs suggest it as a workaround for slow rootless networking. The cost is the same as on a rootful host: the container shares the host's network stack (see "Host namespace sharing as attack surface"). Ports below 1024 still need the sysctl, because the container is not RootlessKit.
What still does not work, and what rootless does not protect
From the docs and this lab: no AppArmor, no docker checkpoint, no overlay networks (so no Swarm), no SCTP port publishing, container IPs that cannot be reached from the host, --cap-add that only works for resources owned by the container's user namespace, and a data root that cannot be on NFS. Each user has their own daemon, image store and socket, so two engineers on one build host pull every image twice. Host-level monitoring and log agents that read /var/lib/docker or the root socket see none of these containers: each user's data lives in ~/.local/share/docker and the daemon logs to journalctl --user, so point the agents at each user's socket and data root.
The socket at /run/user/1000/docker.sock controls your whole account. Whoever can talk to it can mount your home directory, read your SSH keys and cloud credentials, and run anything as you. Do not mount it into containers and do not expose it on TCP without mutual TLS, for the same reasons as the root socket in "The Docker socket and daemon hardening". Containers still share the host kernel, so a kernel bug can still matter. Keep non-root images, dropped capabilities and seccomp on top.
Clean up
Uninstall the user service, remove the rootless data through the namespace (plain rm fails on files owned by subordinate UIDs), then undo the host changes and bring the system daemon back:
The uninstaller also deleted the rootless context and switched the CLI back to default. If you added DOCKER_HOST to ~/.bashrc, remove it.
docker run -p 443:8443 app fails with "cannot expose privileged port 443", and --cap-add NET_BIND_SERVICE makes no difference. Which change fixes it while changing the least on the host?/etc/subuid has ubuntu:100000:65536. A container runs as --user 1000. Which host UID owns its files on a bind mount?docker run --cpuset-cpus 0 app on a fresh rootless install prints "Cpuset discarded" and the container starts. What is going on, and what fixes it?--cpus worked before any change (cpu.max showed 50000 100000), and cpuset works after delegation.Try this
Work through “Clean up” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from rootless docker, keep “Clean up”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.