Deploy a Compose file to a swarm, with private images, secrets and configs that rotate.
Intermediate25 min · lesson 16 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (7 files, 2 KB): stacks.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/stacks.tar.gz && tar -xzf stacks.tar.gz, which creates ~/lab/stacks/. SHA-256: ebf2c807bde42aca9e6bf783d423b806714db7902d2ac7f082c7f2a32ca00cb7
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab simulates a three-node swarm with privileged Docker-in-Docker containers and runs a private registry next to it. If the VM does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec and open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima); reset it with ./setup/create-lab.sh --profile sec --recreate. The lesson works in ~/lab/stacks, where the lesson files unpack. On this VM ubuntu is not in the docker group, so commands for the VM's own daemon use sudo.
This Compose file runs with docker compose up on a single host, the way "Multi-container apps with Compose" (Docker for beginners) built such files:
compose.yaml
services:
web:
build: ./web
image: lab-private:5000/shop-web:1
container_name: shop-web
restart: unless-stopped
ports:
- "8080:80"
depends_on:
db:
condition: service_healthy
db:
image: lab-reg:5000/postgres:18-alpine
environment:
POSTGRES_PASSWORD_FILE: /run/secrets/db_password
secrets:
- db_password
healthcheck:
test: ["CMD", "pg_isready", "-U", "postgres"]
tools:
image: lab-reg:5000/alpine:3.22
profiles: [debug]
secrets:
db_password:
file: ./db_password_v1.txt
The lab uses the simulated cluster from "Swarm clusters and services", rebuilt with swarm-lab.sh from the lesson files, with both workers labelled disk=ssd:
services.tools: additional property 'profiles' is not allowed
docker stack deploy and docker stack config read the file with the older Compose file format 3 schema, not the Compose Specification that docker compose uses. The long form of depends_on with a condition is rejected, and so is profiles. Remove both and the file passes, but deploying it shows what else a stack silently drops:
Ignoring unsupported options: build, restart
Ignoring deprecated options:
container_name: Setting the container name is not supported.
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Creating network try_default
Creating secret try_db_password
Creating service try_tools
Creating service try_web
Creating service try_db
$docker stack rm try
Removing service try_db
Removing service try_tools
Removing service try_web
Removing secret try_db_password
Removing network try_default
There is no startup ordering in a swarm. Services start in parallel and are restarted until they stay up, so an application must retry its database connection instead of relying on depends_on. docker stack config -c file prints the file as Swarm will read it, with variables substituted and paths made absolute, and is the cheapest check to run in CI before a deploy.
The stack file
A shop stack: two replicas of a web image from a private registry, with its nginx configuration delivered as a Swarm config, and Postgres 18 with its password delivered as a Swarm secret. Postgres 18 keeps its data under /var/lib/postgresql (the image sets PGDATA=/var/lib/postgresql/18/docker), so that is the path the volume mounts; "Recipe: PostgreSQL, MySQL and MongoDB" covers the image itself.
Secrets and configs get explicit name: values with a version suffix. Without name:, Swarm calls them <stack>_<key> (shop_db_password), and the immutability section below shows why a version in the name matters. The long secret syntax sets uid, gid and mode for the file inside the container: 70 is the postgres user in the Alpine-based image, and 0400 makes the file readable by that user only. The placement constraint is deliberately wrong, as the volume section shows.
Private images and --with-registry-auth
The private registry lab-private:5000 requires a login. Its password file comes with the lesson files; it was created with htpasswd from apache2-utils:
WARNING! Your credentials are stored unencrypted in '/home/ubuntu/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
Login Succeeded
The registry speaks plain HTTP, and the nodes accept that only because swarm-lab.sh lists the lab subnet under --insecure-registry. With a password on top, that sends the credentials in clear text on every pull; it is tolerable on a private lab bridge only. A real registry needs TLS, and a password-protected registry never belongs in insecure-registries. The login warning is accurate too: without a credential helper the CLI keeps the password base64-encoded in ~/.docker/config.json ("Registries, tags and sharing images" in Docker for beginners covers helpers). The images were built on mgr1 and pushed, then the local copies removed, so every node, including mgr1, has to pull them. First deploy without passing credentials:
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Creating network shop_backend
Creating secret shop_db_password_v1
Creating config shop_web_conf_v1
Creating service shop_web
Creating service shop_db
$docker service ps shop_web --no-trunc --format "{{.Name}} {{.Node}} {{.CurrentState}} {{.Error}}" | head -n 2
shop_web.1 mgr1 Rejected 1 second ago "failed to resolve reference "lab-private:5000/shop-web:1": pull access denied, repository does not exist or may require authorization: authorization failed: no basic auth credentials"
shop_web.1 mgr1 Rejected 6 seconds ago "failed to resolve reference "lab-private:5000/shop-web:1": pull access denied, repository does not exist or may require authorization: authorization failed: no basic auth credentials"
The deploy itself succeeds: it only writes the desired state. The tasks are rejected on the nodes with no basic auth credentials, because docker login stored the credentials in the CLI's config on the VM, and the nodes pull for themselves. --with-registry-auth makes the CLI send them along with the service spec:
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Updating service shop_web (id: kfxdapoi6qmz1quj5mrzrfbtg)
Updating service shop_db (id: jsxb1tnzhsz3rzz56b3pplj1w)
ID NAME MODE REPLICAS IMAGE PORTS
jsxb1tnzhsz3 shop_db replicated 1/1 lab-reg:5000/postgres:18-alpine
kfxdapoi6qmz shop_web replicated 2/2 lab-private:5000/shop-web:1 *:8080->80/tcp
shop_db.1 wrk2 Running 14 seconds ago
shop_web.1 mgr1 Running 8 seconds ago
shop_web.2 wrk1 Running less than a second ago
shop web v1, config v1, on f2da1d3a5cc2
All three tasks run, and the web service answers through the routing mesh. The credentials were not used once and forgotten. The manager stores them in the service spec in the Raft log and hands them to every node that runs a task of the service. To see that, log out, change the image tag in the file and deploy without the flag:
Removing login credentials for lab-private:5000
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Updating service shop_web (id: kfxdapoi6qmz1quj5mrzrfbtg)
image lab-private:5000/shop-web:2 could not be accessed on a registry to record
its digest. Each node will access lab-private:5000/shop-web:2 independently,
possibly leading to different nodes running different
versions of the image.
Updating service shop_db (id: jsxb1tnzhsz3rzz56b3pplj1w)
shop_web.1 lab-private:5000/shop-web:2 mgr1 Running 31 seconds ago
shop_web.2 lab-private:5000/shop-web:2 wrk1 Running 20 seconds ago
$docker service inspect shop_web | grep -i -e registryauth -e pulloptions || echo "no registry credentials in the API response"
no registry credentials in the API response
Both new shop-web:2 tasks were pulled from the private registry, on nodes that had never pulled that tag, with no credentials on the client: the update kept the stored ones. They are not part of what the Engine API returns: docker service inspect, which prints the API's answer, shows no registry credentials at all. They still travel to every node that runs a task of the service, and anyone who controls a manager (root on it, with its Raft data and keys) can recover them. The deploy also printed a warning it did not print with credentials: without them the CLI could not ask the registry for the digest of shop-web:2, so the spec holds the bare tag and each node resolves it on its own (the pinning from "Swarm clusters and services" is lost until a deploy with credentials). Use a dedicated, read-only, pull-only registry token for deployments, not a personal account, rotate it, and redeploy with --with-registry-auth after rotating so the spec carries the new one.
Inside the database container the secret is a file on a read-only tmpfs at /run/secrets/db_password, owned by postgres with mode 0400, as the long syntax asked; without those fields it would be root-owned with mode 0444, readable by every user in the container. Swarm applies these fields when it mounts the secret, which is the behaviour "Recipe: PostgreSQL, MySQL and MongoDB" relies on when it contrasts stacks with plain Compose. The environment holds only the path in POSTGRES_PASSWORD_FILE, which the official image reads at initialisation. docker secret inspect shows the metadata but never the value; there is no command that reads a secret back out of the cluster. Managers keep secrets encrypted in the Raft log (autolock, from "Swarm clusters and services", protects the key), and a worker receives a secret only while it runs a task that was granted it. What can still leak a secret once it is in a container (an entrypoint that exports it, logs, core dumps) is covered in "Runtime secrets, done right" (Advanced container security).
Configs work the same way for non-sensitive files, with two differences: they are mounted directly into the container filesystem rather than on a tmpfs, and anyone with access to a manager's API can read them back (docker config inspect --pretty prints the content). Both secrets and configs are limited to 500 KB.
A local volume lives on one node
pgdata is a local volume, and in a swarm that means local to whichever node runs the task. The constraint node.labels.disk == ssd matches two workers. Write a row, then drain the node the database runs on:
$docker run --rm --network shop_backend -e PGPASSWORD="$(cat db_password_v1.txt)" lab-reg:5000/postgres:18-alpine \
psql -h db -U app -d shop -c "CREATE TABLE orders (id int)" -c "INSERT INTO orders VALUES (42)"
CREATE TABLE
INSERT 0 1
$docker service ps shop_db -f desired-state=running --format "{{.Node}}" | tee db-node.txt
docker node update --availability drain "$(cat db-node.txt)"
wrk2
wrk2
$docker service ps shop_db --format "{{.Name}} {{.Node}} {{.DesiredState}} {{.CurrentState}}"
shop_db.1 wrk1 Running Running less than a second ago
shop_db.1 wrk2 Shutdown Shutdown 7 seconds ago
$docker run --rm --network shop_backend -e PGPASSWORD="$(cat db_password_v1.txt)" lab-reg:5000/postgres:18-alpine \
psql -h db -U app -d shop -c "SELECT * FROM orders"
ERROR: relation "orders" does not exist
LINE 1: SELECT * FROM orders
^
$for n in lab-wrk1 lab-wrk2; do echo "$n: $(docker --context $n volume ls -q)"; done
lab-wrk1: shop_pgdata
lab-wrk2: shop_pgdata
The scheduler did what the constraint allowed: it started the database on the other ssd worker. That node had no shop_pgdata volume, so Docker created an empty one, the Postgres entrypoint initialised a brand-new database in it, and the table is gone from the service's point of view. Both workers now hold a shop_pgdata volume with different contents. Nothing failed; the service is green and empty. The data is still on the first node, which is why pinning the service back to exactly that node brings it back:
wrk2
wrk2
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Updating service shop_db (id: jsxb1tnzhsz3rzz56b3pplj1w)
Updating service shop_web (id: kfxdapoi6qmz1quj5mrzrfbtg)
image lab-private:5000/shop-web:2 could not be accessed on a registry to record
its digest. Each node will access lab-private:5000/shop-web:2 independently,
possibly leading to different nodes running different
versions of the image.
$docker service ps shop_db -f desired-state=running --format "{{.Node}}"
docker run --rm --network shop_backend -e PGPASSWORD="$(cat db_password_v1.txt)" lab-reg:5000/postgres:18-alpine \
psql -h db -U app -d shop -c "SELECT * FROM orders"
wrk2
id
----
42
(1 row)
Pin a stateful service with local volumes to one node, by a label only that node carries (pgdata=true here) or node.hostname == <host>. If that node is down, the task stays Pending, which is the outcome you want for a database: an alert instead of an empty replacement. The alternatives are a volume driver backed by shared or replicated storage, or Swarm cluster volumes (CSI plugins, available since Engine 23), or running the database outside the swarm. Replication between databases is the database's job, not Swarm's.
Secrets and configs are immutable
A new password goes into the secret file and the stack is redeployed:
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
failed to update secret shop_db_password_v1: Error response from daemon: rpc error: code = InvalidArgument desc = only updates to Labels are allowed
The deploy stops with only updates to Labels are allowed. A secret's data cannot change after it is created; the CLI tried to update shop_db_password_v1 in place, the manager refused, and nothing in the stack changed. Configs behave the same way. The way to rotate is a new object under a new name, which is what the _v1 suffix is for. Here the new file becomes version 2, and the stack points the same target at it, so the path inside the container does not change:
$mv db_password_v1.txt db_password_v2.txt
sed -i "s/db_password_v1/db_password_v2/g" stack.yaml
grep -A 2 "^ db_password:" stack.yaml
docker stack deploy -c stack.yaml shop
db_password:
name: shop_db_password_v2
file: ./db_password_v2.txt
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Creating secret shop_db_password_v2
Updating service shop_db (id: jsxb1tnzhsz3rzz56b3pplj1w)
Updating service shop_web (id: kfxdapoi6qmz1quj5mrzrfbtg)
image lab-private:5000/shop-web:2 could not be accessed on a registry to record
its digest. Each node will access lab-private:5000/shop-web:2 independently,
possibly leading to different nodes running different
versions of the image.
Changing the secret in the service spec replaced the database task, and the new task has the new file. The database password, though, did not change:
$docker run --rm --network shop_backend -e PGPASSWORD="$(cat db_password_v2.txt)" lab-reg:5000/postgres:18-alpine \
psql -h db -U app -d shop -c "SELECT 1"
psql: error: connection to server at "db" (10.0.1.5), port 5432 failed: FATAL: password authentication failed for user "app"
$docker service logs shop_db 2>&1 | grep -i "skipping initialization"
shop_db.1.hsrzc1zrxmce@wrk2 | PostgreSQL Database directory appears to contain a database; Skipping initialization
shop_db.1.pwei6cw3st4y@wrk2 | PostgreSQL Database directory appears to contain a database; Skipping initialization
The Postgres entrypoint reads POSTGRES_PASSWORD_FILE only when the data directory is empty; on every later start it skips initialisation, and the log shows that line for both restarts on wrk1 (after pinning and after the secret change). The secret and the database now disagree. Setting the role's password from the new secret file, through the trusted local socket inside the container, brings them back together:
In production, do it in the order that avoids the outage: change the credential in the database first (or create a second role or password that both versions accept), then roll out the new secret to the clients, then remove the old one. Swarm replaces tasks when a secret reference changes, so an application that reads its secret file only at startup gets the new value on that restart. An application that keeps running across a change never sees it, because the mounted file belongs to the old object. The old secret stays in the cluster until you remove it, and Swarm refuses to remove a secret that a service still uses.
A config follows the same rules. Editing web.conf in place fails; a new name rolls the web tasks with the new file:
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
failed to update config shop_web_conf_v1: Error response from daemon: rpc error: code = InvalidArgument desc = only updates to Labels are allowed
Since --detach=false was not specified, tasks will be created in the background.
In a future release, --detach=false will become the default.
Creating config shop_web_conf_v2
Updating service shop_web (id: kfxdapoi6qmz1quj5mrzrfbtg)
image lab-private:5000/shop-web:2 could not be accessed on a registry to record
its digest. Each node will access lab-private:5000/shop-web:2 independently,
possibly leading to different nodes running different
versions of the image.
Updating service shop_db (id: jsxb1tnzhsz3rzz56b3pplj1w)
$curl -s http://10.77.0.11:8080/
docker config ls --format "{{.Name}}"
shop web v2, config v2, on 066658b98ffe
shop_web_conf_v1
shop_web_conf_v2
Updating and removing a stack
Running docker stack deploy again with the same file and stack name is the update; Swarm compares each service with its spec and changes only what differs, using each service's update_config. Removing a service from the file does not remove it from the cluster unless you add --prune. The deploy also returns as soon as the desired state is written (the "Since --detach=false was not specified" message); --detach=false makes it wait for convergence, which is what a CI job usually wants.
Removing service shop_db
Removing service shop_web
Removing secret shop_db_password_v2
Removing config shop_web_conf_v1
Removing config shop_web_conf_v2
Removing network shop_backend
$docker secret ls --format "{{.Name}}"; docker config ls --format "{{.Name}}"
for n in lab-wrk1 lab-wrk2; do echo "$n: $(docker --context $n volume ls -q)"; done
lab-wrk1: shop_pgdata
lab-wrk2: shop_pgdata
docker stack rm removed the services, the network, and the secrets and configs that carry the stack's label. Both shop_pgdata volumes are still on their workers: a stack never removes volumes, and on a multi-node swarm they are spread over every node that ever ran the task. Removing them is a per-node docker volume rm, after you have checked which one holds the data you care about. The lab removes everything with the cluster:
$docker context use default
sudo docker rm -f lab-private
./swarm-lab.sh down
default
Current context is now "default"
lab-private
swarm lab removed
Quick check
01A Postgres service in a stack uses a local named volume and the constraint node.labels.zone == eu-1, which three nodes carry. Its node is drained for patching. What happens?
Incorrect — Two other nodes satisfy the constraint, so the scheduler places the task on one of them at once.
Correct — Local volumes exist per node. The lab's replacement task initialised a fresh database and the orders table was gone until the service was pinned back.
Incorrect — Swarm never moves volume data; the lab showed a separate shop_pgdata on each worker.
Incorrect — Nothing validates that; the service starts and the data problem is silent.
02You put a new password into db_password_v1.txt and run docker stack deploy again. What happens?
Incorrect — The secret cannot be updated in place, so nothing is restarted.
Incorrect — Older guides claimed this; on current versions the CLI tries to update the secret and the manager rejects it.
Incorrect — Names are unique; the CLI updates the existing object instead, which fails.
Correct — Secret and config data is immutable; create a new name such as _v2 and point the service at it.
03A stack was deployed with --with-registry-auth. Weeks later the CI job deploys a new image tag without the flag, after docker logout. The private image pulls fine on every node. Why?
Correct — The manager stores registry auth in the service spec and keeps it on later updates, which is also why a deploy token should be read-only and rotated.
Incorrect — Nodes have no such cache; they use what the service spec gives them.
Incorrect — The lab registry rejected the first, credential-less deploy with "no basic auth credentials".
Incorrect — It does; the lab printed "Removing login credentials" and the deploy sent none.
Try this
Work through “Updating and removing a stack” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from stacks, secrets and configs, keep “Updating and removing a stack”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.