Read-only root filesystem
Find what an image writes, give back only those paths as tmpfs, and know what the flag does not cover.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/readonly.tar.gz && tar -xzf readonly.tar.gz, which creates ~/lab/readonly/. SHA-256: b1d113bee53a3607789572861dccf88a11ac70d021bd1367f56d7d44cbcfe32csh: can't create /usr/local/bin/lab-dropped: Read-only file system is the error an intruder gets when the container they landed in was started with --read-only. After code execution, the next useful step for an attacker is usually a write: a downloaded tool, an edited entrypoint script, a modified config that survives a restart. A read-only root filesystem refuses all of them for every path that comes from the image. The application still needs a few writable places, a PID file, a cache, /tmp, and the work in this lesson is finding exactly those paths and giving them back as small in-memory mounts.
Use the main lab VM, in ~/lab/readonly. The lesson files hold an nginx configuration, a Dockerfile and a compose.yaml used in the second half.
What --read-only changes
A container's root filesystem is an overlay mount: the image layers below, read-only, and a writable layer for this container on top ("Where Docker keeps data" in Docker in depth covers the layout). --read-only mounts the whole overlay read-only, so writes to image paths fail with EROFS, which sh prints as Read-only file system. The /proc/mounts line shows ro on /, and the layer directories under /var/lib/containerd/io.containerd.snapshotter.v1.overlayfs/snapshots/, the containerd image store of a fresh Docker 29 install. An engine upgraded from an older version that still uses the overlay2 graph driver shows paths under /var/lib/docker instead; the ro is what matters. docker inspect reports the setting as HostConfig.ReadonlyRootfs, which the audit at the end uses.
Only the root filesystem is affected. Volumes, bind mounts and tmpfs mounts are separate mounts with their own options, and /proc and /sys keep the protections Docker always applies.
Find the paths an image writes
Guessing which directories need to stay writable is slow. docker diff lists every path a running container has added (A), changed (C) or deleted (D) in its writable layer, so run the image once normally, exercise it, and read the list:
Stock nginx writes in three places. /run/nginx.pid is the PID file. /var/cache/nginx/*_temp are the directories nginx creates for request bodies and proxy buffers. /etc/nginx/conf.d/default.conf was changed by a script in the image's entrypoint that adds an IPv6 listen line. For a real application, send it the kind of traffic it gets in production before you read the diff, since some paths are written only on the first upload or the first cache miss. Now start the same image read-only:
Two things happen. The entrypoint script notices it cannot edit default.conf, logs an info line and carries on, which is fine because the IPv6 change is optional. Then nginx itself stops at its first mkdir() with error 30, EROFS, and the container exits 1. Give back exactly the two directories from the diff as tmpfs mounts, which live in memory and disappear with the container:
The container is Up, serves its page, and docker diff prints nothing: the writable layer stays empty because the only writable places are the two tmpfs mounts, and docker diff does not look inside mounts. Docker mounts each --tmpfs with nosuid,nodev,noexec, so nothing written there can be executed or used as a setuid binary. Set a size on anything that can grow; tmpfs pages count against the container's memory limit ("Volumes, bind mounts and tmpfs in practice" in Docker in depth covers sizes and modes). The info line about default.conf is in the log on every start, a reminder that entrypoint scripts which edit files at start-up are the usual reason an image needs more writable paths than its main process does.
A configuration that writes only to /tmp
Better still is an application configured so that everything it writes lands in one place. nginx takes the PID file and every temporary path as directives, so a short configuration moves them all under /tmp. The image switches to the nginx user (UID 101) that the base image already has, so there is no root master process and no entrypoint that edits files. Port 80 would work for UID 101 on a bridge network, where Docker makes every port unprivileged ("Capabilities, cap-drop and no-new-privileges"); listening on 8080 keeps the image working on host networking and on platforms that do not set that sysctl:
# nginx for a read-only root filesystem: runs as the image's nginx user (UID 101),# listens on 8080 and keeps everything it writes under /tmpworker_processes auto;pid /tmp/nginx.pid;error_log /dev/stderr notice;events {worker_connections 1024;}http {include /etc/nginx/mime.types;access_log /dev/stdout;client_body_temp_path /tmp/client_temp;proxy_temp_path /tmp/proxy_temp;fastcgi_temp_path /tmp/fastcgi_temp;uwsgi_temp_path /tmp/uwsgi_temp;scgi_temp_path /tmp/scgi_temp;server {listen 8080;root /usr/share/nginx/html;}}
FROM nginx:1.30-alpineCOPY nginx.conf /etc/nginx/nginx.confUSER 101:101EXPOSE 8080
In Compose, read_only: true is the --read-only flag and tmpfs: takes the same path:options strings as --tmpfs. This service also drops every capability and sets no-new-privileges, as in "Capabilities, cap-drop and no-new-privileges":
services:web:build: .image: lab-ro-web:1read_only: truetmpfs:- /tmp:size=16mcap_drop:- ALLsecurity_opt:- no-new-privileges:true
nginx runs as UID 101, serves the page, and everything it created is under /tmp: the PID file and the five temporary directories. docker diff prints nothing, and the only writable mount is a 16 MiB tmpfs with noexec. This container needs no capabilities at all. The nginx project publishes a ready-made image built the same way, nginxinc/nginx-unprivileged, which listens on 8080 and keeps its PID file and temporary paths in /tmp. The same idea applies to your own applications. Language runtimes and frameworks usually let you move caches, sockets, PID files and temporary files with a setting or an environment variable, and one tmpfs on /tmp covers them. One rollout failure is common enough to plan for. Some runtimes extract a native library to /tmp and load it from there (JVM libraries such as Netty, Snappy and JNA, and some Python packages), which noexec refuses. Point them at a dedicated directory, for example with -Djava.io.tmpdir or -Dio.netty.native.workdir, on its own small tmpfs with exec, rather than making /tmp executable for everything.
Mounts are not covered
--read-only protects the image's paths and nothing else. A named volume or a bind mount stays writable, and unlike tmpfs it is not mounted noexec:
The same script, written and marked executable in both places, is refused in /tmp and runs from /data. A writable volume is therefore the place an intruder stages tools in a read-only container, and a bind mount of a host directory is a write to the host. Mount everything the application only reads with :ro (or readonly with --mount), as the second command shows: the read works and the write fails with the same Read-only file system. Where a volume must stay writable, keep it as narrow as the data it holds and watch it, because that is where unexpected files will show up. The fields to audit across containers are the root filesystem flag and the tmpfs list:
lab-ng-rw is the finding: readonly=false and no tmpfs mounts. In Kubernetes the same control is readOnlyRootFilesystem: true in the container's securityContext; the Restricted Pod Security Standard does not require it, so set it explicitly. Clean up:
--read-only, --tmpfs /tmp and a named volume on /var/lib/app. An attacker writes the same script to /tmp and to /var/lib/app and marks both executable. Only the copy in the volume runs. Why?noexec on every --tmpfs mount and the same script running from a named volume.--read-only. What is the quickest reliable way to find the paths it needs to write?--read-only --tmpfs /var/cache/nginx --tmpfs /run. It serves requests and writes its PID file, yet docker diff on the container prints nothing. Why?Try this
Work through “Mounts are not covered” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from read-only root filesystem, keep “Mounts are not covered”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.