Read-only root filesystem

Find what an image writes, give back only those paths as tmpfs, and know what the flag does not cover.

Advanced12 min · lesson 11 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (3 files, 1 KB): readonly.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/readonly.tar.gz && tar -xzf readonly.tar.gz, which creates ~/lab/readonly/. SHA-256: b1d113bee53a3607789572861dccf88a11ac70d021bd1367f56d7d44cbcfe32c

sh: can't create /usr/local/bin/lab-dropped: Read-only file system is the error an intruder gets when the container they landed in was started with --read-only. After code execution, the next useful step for an attacker is usually a write: a downloaded tool, an edited entrypoint script, a modified config that survives a restart. A read-only root filesystem refuses all of them for every path that comes from the image. The application still needs a few writable places, a PID file, a cache, /tmp, and the work in this lesson is finding exactly those paths and giving them back as small in-memory mounts.

Use the main lab VM, in ~/lab/readonly. The lesson files hold an nginx configuration, a Dockerfile and a compose.yaml used in the second half.

What --read-only changes

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --read-only alpine:3.22 sh -c 'echo test > /usr/local/bin/lab-dropped'
sh: can't create /usr/local/bin/lab-dropped: Read-only file system
$ docker run --rm --read-only alpine:3.22 sh -c "grep ' / ' /proc/mounts | cut -c1-96"
overlay / overlay ro,relatime,lowerdir=/var/lib/containerd/io.containerd.snapshotter.v1.overlayf

A container's root filesystem is an overlay mount: the image layers below, read-only, and a writable layer for this container on top ("Where Docker keeps data" in Docker in depth covers the layout). --read-only mounts the whole overlay read-only, so writes to image paths fail with EROFS, which sh prints as Read-only file system. The /proc/mounts line shows ro on /, and the layer directories under /var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/, the containerd image store of a fresh Docker 29 install. An engine upgraded from an older version that still uses the overlay2 graph driver shows paths under /var/lib/docker instead; the ro is what matters. docker inspect reports the setting as HostConfig.ReadonlyRootfs, which the audit at the end uses.

Only the root filesystem is affected. Volumes, bind mounts and tmpfs mounts are separate mounts with their own options, and /proc and /sys keep the protections Docker always applies.

Find the paths an image writes

Guessing which directories need to stay writable is slow. docker diff lists every path a running container has added (A), changed (C) or deleted (D) in its writable layer, so run the image once normally, exercise it, and read the list:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-ng-rw nginx:1.30-alpine >/dev/null sleep 2 docker diff lab-ng-rw
C /run A /run/nginx.pid C /var C /var/cache C /var/cache/nginx A /var/cache/nginx/client_temp A /var/cache/nginx/fastcgi_temp A /var/cache/nginx/proxy_temp A /var/cache/nginx/scgi_temp A /var/cache/nginx/uwsgi_temp C /etc C /etc/nginx C /etc/nginx/conf.d C /etc/nginx/conf.d/default.conf

Stock nginx writes in three places. /run/nginx.pid is the PID file. /var/cache/nginx/*_temp are the directories nginx creates for request bodies and proxy buffers. /etc/nginx/conf.d/default.conf was changed by a script in the image's entrypoint that adds an IPv6 listen line. For a real application, send it the kind of traffic it gets in production before you read the diff, since some paths are written only on the first upload or the first cache miss. Now start the same image read-only:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --name lab-ng-ro --read-only nginx:1.30-alpine
/docker-entrypoint.sh: /docker-entrypoint.d/ is not empty, will attempt to perform configuration ... 10-listen-on-ipv6-by-default.sh: info: can not modify /etc/nginx/conf.d/default.conf (read-only file system?) ... 2026/10/07 19:21:32 [emerg] 1#1: mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system) nginx: [emerg] mkdir() "/var/cache/nginx/client_temp" failed (30: Read-only file system)

Two things happen. The entrypoint script notices it cannot edit default.conf, logs an info line and carries on, which is fine because the IPv6 change is optional. Then nginx itself stops at its first mkdir() with error 30, EROFS, and the container exits 1. Give back exactly the two directories from the diff as tmpfs mounts, which live in memory and disappear with the container:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-ng --read-only --tmpfs /var/cache/nginx --tmpfs /run nginx:1.30-alpine >/dev/null sleep 2 docker ps --filter name=^lab-ng$ --format "{{.Names}} {{.Status}}" docker exec lab-ng wget -qO- http://127.0.0.1/ | grep "<title>" docker diff lab-ng
lab-ng Up 2 seconds <title>Welcome to nginx!</title>
$ docker exec lab-ng grep -E ' (/run|/var/cache/nginx) ' /proc/mounts docker logs lab-ng 2>&1 | grep 'can not modify'
tmpfs /run tmpfs rw,nosuid,nodev,noexec,relatime,mode=755,inode64 0 0 tmpfs /var/cache/nginx tmpfs rw,nosuid,nodev,noexec,relatime,mode=755,inode64 0 0 10-listen-on-ipv6-by-default.sh: info: can not modify /etc/nginx/conf.d/default.conf (read-only file system?)

The container is Up, serves its page, and docker diff prints nothing: the writable layer stays empty because the only writable places are the two tmpfs mounts, and docker diff does not look inside mounts. Docker mounts each --tmpfs with nosuid,nodev,noexec, so nothing written there can be executed or used as a setuid binary. Set a size on anything that can grow; tmpfs pages count against the container's memory limit ("Volumes, bind mounts and tmpfs in practice" in Docker in depth covers sizes and modes). The info line about default.conf is in the log on every start, a reminder that entrypoint scripts which edit files at start-up are the usual reason an image needs more writable paths than its main process does.

A configuration that writes only to /tmp

Better still is an application configured so that everything it writes lands in one place. nginx takes the PID file and every temporary path as directives, so a short configuration moves them all under /tmp. The image switches to the nginx user (UID 101) that the base image already has, so there is no root master process and no entrypoint that edits files. Port 80 would work for UID 101 on a bridge network, where Docker makes every port unprivileged ("Capabilities, cap-drop and no-new-privileges"); listening on 8080 keeps the image working on host networking and on platforms that do not set that sysctl:

nginx.conf
# nginx for a read-only root filesystem: runs as the image's nginx user (UID 101),
# listens on 8080 and keeps everything it writes under /tmp
worker_processes auto;
pid /tmp/nginx.pid;
error_log /dev/stderr notice;
events {
worker_connections 1024;
}
http {
include /etc/nginx/mime.types;
access_log /dev/stdout;
client_body_temp_path /tmp/client_temp;
proxy_temp_path /tmp/proxy_temp;
fastcgi_temp_path /tmp/fastcgi_temp;
uwsgi_temp_path /tmp/uwsgi_temp;
scgi_temp_path /tmp/scgi_temp;
server {
listen 8080;
root /usr/share/nginx/html;
}
}
Dockerfile
FROM nginx:1.30-alpine
COPY nginx.conf /etc/nginx/nginx.conf
USER 101:101
EXPOSE 8080

In Compose, read_only: true is the --read-only flag and tmpfs: takes the same path:options strings as --tmpfs. This service also drops every capability and sets no-new-privileges, as in "Capabilities, cap-drop and no-new-privileges":

compose.yaml
services:
web:
build: .
image: lab-ro-web:1
read_only: true
tmpfs:
- /tmp:size=16m
cap_drop:
- ALL
security_opt:
- no-new-privileges:true
ubuntu@secopslog-docker:~/lab/readonly · Docker 29.8.2
$ docker compose --progress quiet -p lab-ro up -d --build sleep 2 docker compose -p lab-ro exec -T web sh -c "id; wget -qO- http://127.0.0.1:8080/ | grep \"<title>\"; ls /tmp"
uid=101(nginx) gid=101(nginx) groups=101(nginx) <title>Welcome to nginx!</title> client_temp fastcgi_temp nginx.pid proxy_temp scgi_temp uwsgi_temp
$ docker diff lab-ro-web-1 docker compose -p lab-ro exec -T web grep " /tmp " /proc/mounts
tmpfs /tmp tmpfs rw,nosuid,nodev,noexec,relatime,size=16384k,inode64 0 0

nginx runs as UID 101, serves the page, and everything it created is under /tmp: the PID file and the five temporary directories. docker diff prints nothing, and the only writable mount is a 16 MiB tmpfs with noexec. This container needs no capabilities at all. The nginx project publishes a ready-made image built the same way, nginxinc/nginx-unprivileged, which listens on 8080 and keeps its PID file and temporary paths in /tmp. The same idea applies to your own applications. Language runtimes and frameworks usually let you move caches, sockets, PID files and temporary files with a setting or an environment variable, and one tmpfs on /tmp covers them. One rollout failure is common enough to plan for. Some runtimes extract a native library to /tmp and load it from there (JVM libraries such as Netty, Snappy and JNA, and some Python packages), which noexec refuses. Point them at a dedicated directory, for example with -Djava.io.tmpdir or -Dio.netty.native.workdir, on its own small tmpfs with exec, rather than making /tmp executable for everything.

Mounts are not covered

--read-only protects the image's paths and nothing else. A named volume or a bind mount stays writable, and unlike tmpfs it is not mounted noexec:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm --read-only --tmpfs /tmp -v lab-ro-data:/data alpine:3.22 sh -c ' for d in /tmp /data; do printf "#!/bin/sh\necho ran from $d\n" > $d/s.sh && chmod +x $d/s.sh && $d/s.sh done'
sh: /tmp/s.sh: Permission denied ran from /data
$ docker run --rm --read-only -v lab-ro-data:/data:ro alpine:3.22 sh -c 'cat /data/s.sh; echo x > /data/s.sh'
#!/bin/sh echo ran from /data sh: can't create /data/s.sh: Read-only file system

The same script, written and marked executable in both places, is refused in /tmp and runs from /data. A writable volume is therefore the place an intruder stages tools in a read-only container, and a bind mount of a host directory is a write to the host. Mount everything the application only reads with :ro (or readonly with --mount), as the second command shows: the read works and the write fails with the same Read-only file system. Where a volume must stay writable, keep it as narrow as the data it holds and watch it, because that is where unexpected files will show up. The fields to audit across containers are the root filesystem flag and the tmpfs list:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ for c in lab-ng-rw lab-ng lab-ro-web-1; do docker inspect -f '{{.Name}} readonly={{.HostConfig.ReadonlyRootfs}} tmpfs={{.HostConfig.Tmpfs}}' $c done
/lab-ng-rw readonly=false tmpfs=map[] /lab-ng readonly=true tmpfs=map[/run: /var/cache/nginx:] /lab-ro-web-1 readonly=true tmpfs=map[/tmp:size=16m]

lab-ng-rw is the finding: readonly=false and no tmpfs mounts. In Kubernetes the same control is readOnlyRootFilesystem: true in the container's securityContext; the Restricted Pod Security Standard does not require it, so set it explicitly. Clean up:

ubuntu@secopslog-docker:~/lab/readonly · Docker 29.8.2
$ docker compose --progress quiet -p lab-ro down docker rm -f lab-ng-rw lab-ng-ro lab-ng docker volume rm lab-ro-data docker rmi lab-ro-web:1
lab-ng-rw lab-ng-ro lab-ng lab-ro-data Untagged: lab-ro-web:1 Deleted: sha256:83322b60f6b1e41439b8350951af321674a8e4bc6c7136585164d67c2c0050ff
Quick check
01A container runs with --read-only, --tmpfs /tmp and a named volume on /var/lib/app. An attacker writes the same script to /tmp and to /var/lib/app and marks both executable. Only the copy in the volume runs. Why?
Incorrect — The tmpfs is writable; the script was written there. The refusal comes from the noexec option.
Incorrect — A volume is its own filesystem mounted over the path and inherits no permissions from the layer under it.
Correct — The lab shows noexec on every --tmpfs mount and the same script running from a named volume.
Incorrect — There is no rule about root and memory-backed filesystems; noexec applies to every user.
02You want to run an image you did not write with --read-only. What is the quickest reliable way to find the paths it needs to write?
Correct — The lab's nginx diff names /run/nginx.pid and /var/cache/nginx, the two paths the read-only run then needed.
Incorrect — That field lists only the tmpfs mounts someone asked for; it knows nothing about what the image writes.
Incorrect — Some failures are logged and ignored, like the entrypoint's default.conf edit, and the process may stop at the first one, so this takes many rounds.
Incorrect — Mounting over / would hide the image's files; the container would not start.
03nginx runs with --read-only --tmpfs /var/cache/nginx --tmpfs /run. It serves requests and writes its PID file, yet docker diff on the container prints nothing. Why?
Incorrect — docker diff works for any container; the lab ran it on nginx from Docker Hub and got a full list.
Incorrect — The diff is empty, not disabled; there are simply no changes in the container's own layer.
Incorrect — nginx still writes files; they go to the tmpfs paths you mounted.
Correct — The writable layer stayed empty because every write went to a mount.

Try this

Work through “Mounts are not covered” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from read-only root filesystem, keep “Mounts are not covered”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related