Volumes, bind mounts and tmpfs in practice
Ownership, SELinux labels, read-only mounts, backup and restore.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/volumes.tar.gz && tar -xzf volumes.tar.gz, which creates ~/lab/volumes/. SHA-256: 98c3a45c214d3bf128496ec4e15fef6e7eda2e009ea6c1b4b4f56bf36ea16672A deploy goes out, the new image runs its service as UID 10001 instead of root, and the container crash-loops on Permission denied writing to /data. Nothing about the volume changed. That is the usual shape of volume trouble in production: the mount works, and the numbers on the files do not match the numbers of the process. This lesson starts where "Volumes and bind mounts" in Docker for beginners ends. It assumes you know the three mount types and the -v versus --mount basics, and it covers ownership with a non-root user, the options only --mount has, SELinux labels, tmpfs settings, a backup you can restore, and external volumes in Compose.
Everything runs on the main lab VM, secopslog-docker, as ubuntu in ~/lab/volumes. The lesson files download has the app/Dockerfile and compose.yaml used below.
Ownership: what the container sees is what the disk stores
Linux checks permissions with numbers. A container process with UID 10001 can write a directory only if the directory's numeric owner, group or mode allows UID 10001, whatever user names exist in the image or on the host. The lab image creates such a user and a /data directory it owns:
FROM alpine:3.22# a non-root service account with a fixed numeric ID, and a data directory it ownsRUN adduser -D -H -u 10001 app \&& mkdir -p /data \&& chown 10001:10001 /dataUSER 10001
The write works because of a volume feature: when an empty named volume is mounted over a directory that exists in the image, Docker copies that directory's contents and its ownership and mode into the volume first. /data in the image belongs to 10001, so the new volume does too, and ls -ln shows the file the process created under that number.
The copy happens only when the volume is empty at the first mount by a container. If something else touched the volume first, the image never gets a say. That is the crash loop from the opening, reproduced: a root job (alpine has no /data, so the volume root stays root:root 755) seeds the volume, then the app starts:
The fix is one chown of the volume to the service's numeric ID, run once from a throwaway root container. In Compose the same idea is an init service that runs the chown and that the app waits for with depends_on: condition: service_completed_successfully. Do not "fix" it by running the app as root again; "Run as non-root" in Advanced container security covers how to choose and pin the IDs in the image so this stays predictable.
Bind mounts never copy anything. The container sees the host directory with the host's numeric owners, here ubuntu (1000):
After chown 10001:10001 the write works, and on the host the file belongs to a UID that has no name there. The other options are to run the container as your own UID with --user "$(id -u):$(id -g)" (fine for development tools, as "Volumes and bind mounts" in Docker for beginners showed), or to give the directory a shared group, make it group-writable and pass --group-add with that GID. What changes this picture entirely is user-namespace remapping, where container UID 0 maps to an unprivileged host range; "User-namespace remapping" in Advanced container security covers it.
What only --mount can say
-v packs everything into source:target:options and has no room for several newer options. The first is a volume subpath: mount one directory of a volume instead of all of it, which lets one volume hold the config for several services without each seeing the others' files. Seed a volume with two directories and mount only app, read-only:
The container sees app.ini and nothing from proxy, and readonly makes the write fail with Read-only file system. The subpath must already exist; Docker does not create it, and the error names the path inside the volume's _data directory. The -v attempt shows the syntax limit: subpath= is not a valid -v mode, so it fails with invalid mode. In Compose the long volume syntax has volume: { subpath: app }.
The second is volume-nocopy, which switches off the copy-from-image behaviour from the previous section:
The first container got nginx's two default pages copied into the new lab-html volume. The second printed nothing: its volume stayed empty because volume-nocopy told Docker not to seed it. Use it when the image's content at that path should never be persisted. -v accepts the same switch as the nocopy mode.
Read-only works with both syntaxes (:ro or readonly). One detail matters for bind mounts of directories that contain other mounts, for example a host path with a disk or tmpfs mounted below it. On kernel 5.12 and newer, Docker makes those submounts read-only too (the default bind-recursive=enabled). On older kernels, and on older engines, they stay writable, so a "read-only" bind of /srv could still be written under /srv/data. --mount adds bind-recursive=writable to get the old behaviour on purpose and bind-recursive=disabled to not mount the submounts at all.
SELinux labels: :z and :Z
On hosts with SELinux enforcing (Fedora, RHEL, Rocky, AlmaLinux), the label on a file decides access as much as its owner does. Containers run as the container_t type and may only touch files labelled for containers, so a bind mount of /srv/app/config, labelled var_t or user_home_t, fails with Permission denied even when the UIDs line up. The audit log (ausearch -m avc) shows the denial; the mount options fix it:
Two consequences catch people. :Z on a directory makes it unreadable for every other container, including the next version of the same service if it runs alongside the old one during a rollout; use :z for anything shared. And the relabel changes the host files themselves, so never put :z or :Z on system or home directories such as /, /etc, /usr or /home, which can leave the host unbootable or users locked out. In Compose, short syntax takes :z, and long syntax has bind: { selinux: z }.
The lab VM is Ubuntu with AppArmor, not SELinux, so the labels cannot be demonstrated there. What you can see is that the option is harmless where SELinux is off, and that it belongs to -v:
-v ...:Z ran and ls -Z shows ?, no label, because there is no SELinux policy to apply. --mount rejects a bare Z field: the docs say it does not support the SELinux relabel options, so on SELinux hosts bind mounts that need relabelling use -v or the Compose syntax above. "AppArmor and SELinux" in Advanced container security covers the policies themselves.
tmpfs: size, mode and noexec
A tmpfs mount lives in memory and vanishes with the container. "Volumes and bind mounts" in Docker for beginners showed the flags Docker adds by default; two of them change what you can do with it:
--mount type=tmpfs takes tmpfs-size and tmpfs-mode (here 64 MiB and 1770) and always mounts noexec, so the script written to it cannot run (Permission denied). That is a sensible default for scratch space, because a dropped payload in /tmp is a common first step of an attack. When a tool really must execute from scratch space (some installers and JIT caches do), --tmpfs takes raw mount options and exec lifts the restriction; --mount has no field for it in 29.8.2. Always set a size. Without one the maximum is half the host's RAM, and the docs are explicit that tmpfs data counts toward the container's memory limit, so a growing tmpfs pushes the container toward an OOM kill instead of a clean "no space left" error. Under memory pressure tmpfs pages can also be swapped to disk, so do not treat it as a guarantee that a secret never touches storage. The default mode is 1777, world-writable like /tmp.
A backup you can restore
The common pattern is a throwaway container that mounts the volume read-only next to a host directory and runs tar. The part that is easy to get wrong is consistency: a database writes its files continuously, and a tar of a running data directory can capture them mid-write. Postgres may then refuse to start from the copy, or start and replay to a state that never existed. Stop the writer, copy, start it again. Start a Postgres 18 container and give it some data (note the mount point: from version 18 the official image keeps its data under /var/lib/postgresql/18/docker and expects the volume at /var/lib/postgresql):
The database was down for the seconds the tar took; that is the price of a consistent file-level copy. The listing shows numeric owners, 70/70, the postgres UID of the Alpine-based image. Alpine's tar stores numeric IDs and restores them when it runs as root, which is what you want: the restored files must belong to UID 70 again. The archive itself is owned by root and readable by every local user (-rw-r--r--); a database backup holds every row, so move it to storage with restricted access and encrypt it if it leaves the host. Restore into a new volume and start a second server on it:
PG_VERSION is owned by 70 and the second server returns both rows. A backup is only proven once a restore like this has worked, so script the restore check next to the backup. When stopping the database is not acceptable, use the database's own tools instead: pg_dump for a logical dump or pg_basebackup for a physical one, both consistent while the server runs. "Recipe: PostgreSQL, MySQL and MongoDB" covers them and the other two engines.
Where the bytes go, and performance
Reads and writes on a volume or bind mount go straight to the host filesystem. They skip the overlay mount that serves the container's root filesystem, so there is no copy-up on first write ("Where Docker keeps data" measures that cost). That is the main performance rule: anything a service writes heavily, databases, queues, caches, uploads, belongs on a volume, not in the writable layer.
A named volume lives under /var/lib/docker/volumes, on the same disk as everything else Docker keeps. To put one service's data on a faster or bigger disk without giving up the named-volume workflow, the local driver can bind a directory you choose:
The file written through the volume landed in fastdisk, and docker volume inspect shows both the nominal mountpoint and the options that redirect it. The directory must exist before the first mount, and docker volume rm removes only the volume record; the data in fastdisk stays. The same driver mounts network storage with --opt type=nfs, for example:
docker volume create --driver local \--opt type=nfs --opt o=addr=10.0.20.5,rw,nfsvers=4.2 \--opt device=:/exports/app-data app-data
NFS behaves very differently from a local disk: higher latency on every synchronous write, locking that some databases do not trust, and root squashing that turns container UID 0 into nobody. Run databases on local or block storage and keep NFS for shared files. On Docker Desktop for Mac and Windows, a bind mount of a project directory crosses from your machine into Docker's VM (VirtioFS on current versions), which is noticeably slower for trees with many small files; keep dependency directories such as node_modules in a named volume there.
External volumes in Compose
Compose creates the volumes a project declares and names them <project>_<volume>. When data must outlive the project, or several projects share one volume, declare it external and create it yourself. Compose then refuses to start instead of quietly creating an empty one:
name: lab-volsservices:reader:image: alpine:3.22command: ["cat", "/shared/hello.txt"]volumes:- shared:/shared:rovolumes:shared:external: truename: lab-shared
With lab-shared missing, docker compose run stops with external volume "lab-shared" not found. Once it exists, the service reads it. docker compose down -v removes the project's network and would remove volumes Compose created, but it leaves external ones alone, which is the point: a mistyped down -v cannot delete the production database volume if that volume is external. "Multi-container apps with Compose" in Docker for beginners covers the rest of the volume syntax.
Clean up. hostdata, labels and fastdisk contain files owned by other UIDs, so their removal needs sudo:
Permission denied on its data directory, although the image creates that directory owned by 10001. What happened?docker run --rm -v pgdata:/data:ro -v /backup:/backup alpine tar czf ... while the database keeps serving traffic. A restore test fails to start the server. Which change fixes the method?/srv/shared/assets. The first was started with :Z. Now the second gets Permission denied on every file there. What should the mounts use?:Z labels the content privately for one container; :z labels it so every container can use it.--mount does not accept the relabel options; the lab showed it rejecting a Z field.Try this
Work through “External volumes in Compose” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from volumes, bind mounts and tmpfs in practice, keep “External volumes in Compose”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.