Run as non-root
A numeric USER, code the app cannot change, and the run-time overrides that undo it.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/nonroot.tar.gz && tar -xzf nonroot.tar.gz, which creates ~/lab/nonroot/. SHA-256: 5d60386ef52fa27ee501535babd0069887fbabc282e8c33b37c4943cf2157c92docker run --user 10001 alpine:3.22 id prints uid=10001 gid=0(root) groups=0(root). The process is not root, but its group is 0, so every file it creates belongs to group root and it can write anything that is group-writable for root. It is a small example of how "runs as non-root" is really decided by a pair of numbers, set in the image or at run time, that the kernel compares with the numbers stored on files. "What root in a container really is" showed why UID 0 in a container is UID 0 on the host. This lesson builds the fix, an image that runs as a fixed UID and GID, keeps its own code out of its reach, and still writes where it has to.
Use the main lab VM. The lesson files go in ~/lab/nonroot: a small shell "app" that appends a line to /data/starts.log each time it starts, and two Dockerfiles for it.
#!/bin/sh# lab app: appends one line per start to /data/starts.log, then prints the logset -eecho "$(date -u +%FT%TZ) started as uid=$(id -u) gid=$(id -g)" >> /data/starts.logcat /data/starts.log
A numeric USER in the image
FROM alpine:3.22# a service account with fixed numeric IDs, and the one directory it may writeRUN addgroup -S -g 10001 app \&& adduser -S -D -H -u 10001 -G app app \&& mkdir /data \&& chown 10001:10001 /data# the code stays owned by root: the app can run it but not change itCOPY --chmod=0755 app.sh /app/app.shUSER 10001:10001CMD ["/app/app.sh"]
addgroup -S and adduser -S -D -H create a system account with no password and no home directory, with fixed IDs. Pick IDs that mean nothing on the hosts the image will run on. Without a user namespace the number crosses unchanged, so 1000 would be the first login account on most machines, and 100 to 999 belong to distribution service accounts. 10001 is a common choice; distroless images use 65532 for their nonroot user ("Minimal bases: distroless, scratch, static and Alpine"). mkdir /data and chown give the account exactly one directory it owns.
USER 10001:10001 uses numbers rather than app. A name is resolved through the image's /etc/passwd when the container starts, and a scratch or distroless image may not have one. A number also lets tools check the claim without opening the image. Kubernetes is the common example: with runAsNonRoot: true the kubelet refuses to start a container whose user is 0, and also one whose user is a name it cannot verify. Build the image and read back what Docker stored:
Keep Dockerfile comments on their own lines. A comment after an instruction is not a comment, and for USER it becomes part of the user string:
The build succeeds, Config.User holds the whole line, and the container cannot start because no group is called 10001 # the app user. Nothing warns at build time, so this kind of mistake reaches a registry. Now run the real image twice with a named volume on /data:
The first run printed one line, the second printed two, both as uid=10001 gid=10001. The write worked on a brand-new volume because Docker copied the owner and mode of the image's /data into the empty volume at its first mount, which is why the Dockerfile creates and chowns /data at all. What happens with a volume that root wrote first, and with bind mounts, where the host directory's owner applies unchanged, is covered in "Volumes, bind mounts and tmpfs in practice" in Docker in depth, along with the fixes. None of them involves running the app as root again.
Code the app cannot change
COPY without --chown leaves files owned by root. --chmod=0755 makes the script executable for everyone and writable only by root:
UID 10001 can run /app/app.sh but cannot change it. That is the intended state. A process that is compromised through a bug cannot rewrite the code it runs, and anything it does write lands only in /data, where you expect writes. A common shortcut gives the whole application directory to the service account:
FROM alpine:3.22RUN addgroup -S -g 10001 app \&& adduser -S -D -H -u 10001 -G app app \&& mkdir /data \&& chown 10001:10001 /data# the common shortcut: hand the whole app directory to the service accountCOPY --chown=10001:10001 --chmod=0755 app.sh /app/app.shUSER 10001:10001CMD ["/app/app.sh"]
Now the process owns its code and can modify it. Inside a running container, a change made this way lasts until the container is removed, and it survives restarts. Use --chown for directories the app must write, not for the code itself. If a framework insists on writing next to its code, give it that one subdirectory.
USER is a default, and so is the group
Whoever starts the container can replace the image's choice:
--user 0 turns the same image back into a root container with root's supplementary groups. Docker Engine has no daemon setting that refuses UID 0 containers, so on plain Docker the controls are review and checks: look for --user and Compose user: values in the files that start services, and read Config.User from images in CI. User-namespace remapping and rootless Docker change what UID 0 means instead ("User-namespace remapping", "Rootless Docker").
The second pair shows the opening example. With --user 10001 alone, Docker looks the UID up in the image's /etc/passwd to find a primary group. Plain alpine has no entry for 10001, so the group falls back to 0. Inside lab-nonroot:1, which has the entry, --user 10001 would get group 10001. Pass both numbers, --user 10001:10001 on the command line and user: "10001:10001" in Compose, and the result does not depend on what the image contains.
Images that start as root on purpose
An empty Config.User does not always mean the service runs as root. Several official images start as root so their entrypoint can fix the ownership of a data directory, and then switch user before starting the server:
The redis image has no USER, yet the host sees redis-server running as UID 999. Its entrypoint uses setpriv to switch to the redis user, which needs the SETUID and SETGID capabilities in the container's default set. During that first moment the process is root, so this pattern still depends on the rest of the hardening: dropped capabilities, no-new-privileges and the default security profiles. For an audit, the host's ps (or /proc/<pid>/status) answers "which UID is it running as", while Config.User answers only "what did the image ask for". "Production best practices" in Docker in depth turns both into a pre-ship check.
Remove everything:
docker run --user 10001 alpine-based-image. The image has no /etc/passwd entry for 10001. Files the service creates in a shared volume show up as owned by 10001:0. Why group 0?--user 10001:10001 the lab shows groups=10001 only.uid=10001 gid=0(root) for this case. Pass --user 10001:10001 to make the group explicit.id before any file is written.COPY --chown=10001:10001 . /app and USER 10001:10001. A reviewer asks for the --chown to be removed from the code copy. What risk is the reviewer pointing at?/app/app.sh as 10001 after --chown and is refused without it.USER 10001:10001 # app user. The build succeeds, but docker run fails with "unable to find group". What went wrong?docker run failing with "unable to find group".Try this
Work through “Images that start as root on purpose” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from run as non-root, keep “Images that start as root on purpose”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.