Swarm clusters and services
Managers, Raft quorum, join tokens, services, rolling updates, rollback and autolock.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/swarmintro.tar.gz && tar -xzf swarmintro.tar.gz, which creates ~/lab/swarmintro/. SHA-256: 3f9b48a5c61c59b96e210d8683daeca5008cdaedbc9fc061340c55bbd45c5d0a./setup/create-lab.sh --profile sec, then open a shell in it with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima). To get back to a clean state at any point, recreate it with ./setup/create-lab.sh --profile sec --recreate. On this VM the ubuntu account is not in the docker group, so commands for the VM's own daemon use sudo.This is where the lesson is going: a three-node cluster, as its only manager reports it. The cluster does not exist yet, so there is nothing to run here; the same command comes back once you have built it.
That is Swarm mode: several Docker engines joined into one cluster, with a manager (mgr1, the Leader) that accepts docker service commands and two workers that run whatever the manager assigns. The ENGINE VERSION column is the ordinary Docker Engine; Swarm is built into it, and there is nothing else to install. This lesson builds that cluster, runs services on it, updates and rolls them back, and takes nodes and managers away to see what breaks and what keeps running.
Where Swarm stands
Swarm mode ships in every Docker Engine release and is maintained; its core is the open-source SwarmKit project, maintained mainly by Mirantis, which bought Docker's enterprise business in 2019. It changes slowly. New orchestration work (autoscaling, operators, admission policy, most of the tooling) happens around Kubernetes, the usual choice for large fleets. Swarm suits a handful of hosts run by a team that knows Docker and Compose and needs scheduling, rolling updates, secrets and overlay networks without a separate control plane. For a single host, Docker's documentation points to Compose.
The lab: three nodes in one VM
A real swarm is several machines. The lab simulates them with Docker-in-Docker: each node is a privileged docker:29-dind container running its own dockerd on the sec VM, with its own IP address on a bridge network lab-swarm (10.77.0.0/24). The script below, in the lesson files, starts the three nodes and a plain registry (lab-reg:5000) they can all pull from, copies a few base images from the VM into the nodes so they never pull from Docker Hub themselves, and creates a Docker context per node from the TLS client certificates the dind image generates. The nodes run with --insecure-registry 10.77.0.0/24, which lets them talk plain HTTP to the lab registries; that is acceptable on a private lab bridge and nowhere else, since a real registry needs TLS.
Run it from ~/lab/swarmintro, where the lesson files unpack:
#!/bin/sh# Simulates three Docker hosts inside one lab VM: Docker-in-Docker containers lab-mgr1, lab-wrk1 and# lab-wrk2 on the bridge network lab-swarm (10.77.0.0/24), plus a registry lab-reg:5000 they all use.# Each node is a full dockerd (docker:29-dind, privileged); --tmpfs /run gives a restarted node an empty# /run, as after a reboot. The node's TLS client certificates become a Docker context of the same name,# so `docker --context lab-wrk1 info` talks to that node's daemon. Base images are pulled once on the VM# (only if missing) and copied into lab-mgr1, so the nodes never pull from Docker Hub; alpine and postgres# are pushed to lab-reg for the other nodes.# Usage: ./swarm-lab.sh up|down (runs the VM's docker through sudo; contexts belong to the caller)[ -f /etc/secopslog-lab ] || { echo "refusing: run this only in the SecOpsLog lab VM" >&2; exit 2; }set -euNODES="lab-mgr1:10.77.0.11 lab-wrk1:10.77.0.12 lab-wrk2:10.77.0.13"IMAGES="nginx:1.30-alpine alpine:3.22 postgres:18-alpine"case "${1:-}" inup)for img in $IMAGES; dosudo docker image inspect "$img" >/dev/null 2>&1 || { echo "pulling $img (not on this VM yet)"; sudo docker pull -q "$img" >/dev/null; }donesudo docker network create --subnet 10.77.0.0/24 lab-swarm >/dev/nullsudo docker run -d --name lab-reg --network lab-swarm --ip 10.77.0.5 registry:3 >/dev/nullfor n in $NODES; doname=${n%%:*}; ip=${n#*:}sudo docker run -d --privileged --name "$name" --hostname "${name#lab-}" --tmpfs /run \--network lab-swarm --ip "$ip" docker:29-dind --insecure-registry 10.77.0.0/24 >/dev/nulldonefor n in $NODES; doname=${n%%:*}; ip=${n#*:}i=0; until sudo docker exec "$name" docker info >/dev/null 2>&1 && sudo docker exec "$name" test -f /certs/client/key.pem; doi=$((i+1)); [ $i -lt 60 ] || { echo "$name did not start" >&2; exit 1; }; sleep 1; doned=$(mktemp -d)sudo docker cp "$name:/certs/client" - | tar -x -C "$d"docker context create "$name" --description "lab node ${name#lab-}" \--docker "host=tcp://$ip:2376,ca=$d/client/ca.pem,cert=$d/client/cert.pem,key=$d/client/key.pem" >/dev/nullrm -rf "$d"echo "$name $ip $(docker --context "$name" version -f '{{.Server.Version}}') ready"donefor img in $IMAGES; do sudo docker save "$img" | docker --context lab-mgr1 load -q >/dev/null; donefor img in alpine:3.22 postgres:18-alpine; dodocker --context lab-mgr1 tag "$img" "lab-reg:5000/$img"docker --context lab-mgr1 push -q "lab-reg:5000/$img" >/dev/nulldoneecho "lab-reg:5000 has: $(curl -s http://10.77.0.5:5000/v2/_catalog)";;down)docker context use default >/dev/null 2>&1 || truefor n in $NODES; do docker context rm -f "${n%%:*}" >/dev/null 2>&1 || true; donesudo docker rm -f -v lab-mgr1 lab-wrk1 lab-wrk2 lab-reg >/dev/null 2>&1 || truesudo docker network rm lab-swarm >/dev/null 2>&1 || trueecho "swarm lab removed";;*) echo "usage: $0 up|down" >&2; exit 2 ;;esac
This VM had no postgres:18-alpine yet, so the script pulled it first; on a VM that already has all three images that line does not appear. Each node then reports its address and engine version, and the last line lists what the lab registry holds.
A Docker context is a named daemon endpoint plus credentials. docker --context lab-wrk1 info talks to the worker's daemon over TLS on port 2376; docker context use lab-mgr1 makes the manager the default for every later command, so from here on a plain docker command goes to mgr1. The same mechanism (usually with host=ssh://user@host) is how you would manage a real remote swarm from a workstation.
What the simulation does not give you: the three nodes share one kernel, one clock and one physical network, so a kernel panic, a clock skew or a cable failure cannot hit one node alone, and there is no firewall between them. On real hosts, open these between the nodes only: 2377/tcp to the managers (cluster management and joins), 7946/tcp and 7946/udp between all nodes (node discovery and gossip), 4789/udp between all nodes (VXLAN overlay traffic), and IP protocol 50 (ESP) if you use encrypted overlays. Keep every one of them closed to everything outside the cluster: 2377 accepts joins from anyone holding a token, and VXLAN has no authentication, so anyone who can send packets to 4789 on a node can inject traffic into its overlays ("Overlay networks and the routing mesh" shows that traffic on the wire). The --tmpfs /run in the script makes a restarted node start with an empty /run, as a rebooted host would.
Creating the swarm
docker swarm init turns mgr1 into a one-node swarm and makes it the first manager. --advertise-addr is the address the other nodes will use to reach it; on a host with several interfaces, set it explicitly, or the daemon picks one or refuses to guess. The node ID (25 characters) and the join token are random per cluster, so yours differ.
A join token has three parts: a version (SWMTKN-1), a digest of the cluster's root CA certificate, which lets the joining node verify it is talking to the right cluster, and a secret. Worker and manager tokens differ only in the secret. With the worker token and access to port 2377, anyone can add a machine that receives tasks and the secrets granted to them; the manager token adds a machine with a full copy of the cluster state. Tokens do not expire. Rotate them when they may have leaked, or after you finish adding nodes:
After --rotate, the old worker token is rejected with A valid join token is necessary to join this cluster, and only the new one works. Nodes that already joined are not affected. Rotating the manager token works the same way with docker swarm join-token --rotate manager.
docker node ls at this point is the listing this lesson opened with. * marks the node the command ran on, and the AVAILABILITY and MANAGER STATUS columns come back later. A worker refuses cluster commands, because only managers hold the cluster state:
Managers, Raft and quorum
Managers keep the cluster state (nodes, services, networks, secrets, configs) in a replicated log maintained with the Raft consensus algorithm. One manager is the leader and makes changes; a change is committed only when a majority of managers, the quorum, has stored it. With N managers the cluster tolerates the loss of (N-1)/2 of them, rounded down:
Two managers are worse than one: the quorum is still two, so losing either one stops the cluster, and you now have two machines that can fail. Run three managers in production, five if you must survive two failures or one failure during maintenance; more managers make every write slower. Spread managers across failure domains (racks, availability zones), keep their clocks synchronised, and keep workloads off them if they are busy (drain them, below).
Services and tasks
You do not start containers on a swarm; you declare services. A service says which image to run, how many copies, with which ports, networks and update policy, and the managers keep creating tasks until reality matches. A task is one scheduled unit of work bound to a node; on that node it becomes one container. Images must come from a registry every node can reach, because each node pulls for itself. The lab builds three versions of a small nginx image on mgr1 (2 is a normal release; 3 listens on the wrong port, so its health check fails) and pushes them to lab-reg:5000:
FROM nginx:1.30-alpineARG VERSION=1ARG PORT=80ENV VERSION=$VERSION PORT=$PORTCOPY default.conf.template /etc/nginx/templates/default.conf.templateHEALTHCHECK --interval=2s --timeout=2s --start-period=2s --retries=2 \CMD wget -q -O /dev/null http://127.0.0.1/healthz || exit 1
docker service create prints the service ID, then waits and reports progress until the tasks are running and have stayed up for the monitor period (verify: Waiting 10 seconds...), and ends with verify: Service <ID> converged. In a terminal those lines redraw in place as progress bars; the lab captured them without a terminal, so each refresh is a new line (most are elided here). -d returns at once instead, and -q waits silently.
REPLICAS reads running over desired: 3/3 is converged. docker service ps lists the tasks, one per slot (web.1 to web.3), and the scheduler spread them over the three nodes. The service spec now holds the tag plus @sha256: and a digest: at create time the CLI resolved the tag to a digest through the registry and pinned it, so every node runs exactly the same image even if someone pushes a new web:1 later. When the registry cannot be reached, the CLI warns image ... could not be accessed on a registry to record its digest and each node resolves the tag on its own; treat that warning as an error in a pipeline.
The service publishes port 8080 on every node, including nodes without a task, and the client address the app sees is a 10.0.0.x address of Swarm's ingress network rather than the VM's. That is the routing mesh, which "Overlay networks and the routing mesh" takes apart.
A replicated service runs a number of copies wherever they fit. A global service runs exactly one task on every node that is available, including nodes that join later, which suits per-host agents such as log shippers and monitoring:
Global task names end in the node ID rather than a slot number, and there is nothing to scale.
Rolling updates and rollback
The create command above already set the update policy, because that is where it belongs: --update-parallelism 1 replaces one task at a time, --update-delay 5s waits between batches, --update-monitor 10s watches each new task for 10 seconds after it starts, and --update-failure-action rollback reverts the whole service to its previous spec if a task fails during that window. The defaults are parallelism 1, no delay, a 5-second monitor and pause. A failure means a task that exits or becomes unhealthy within the monitor window; --update-max-failure-ratio (default 0) sets how many such failures, as a fraction of the tasks updated, the update tolerates before the failure action fires.
Each slot shows its history: the running web:2 task and, indented with \_, the web:1 task it replaced on the same node. Version 3 is the broken release:
The update started with one slot, web.2 in this run (the order varies). Its web:3 task never became healthy (nginx listens on 8080, the health check probes port 80), so it was stopped, a second web:3 task fared no better, and within the monitor window the update turned into a rollback: the CLI printed rollback: update rolled back due to failure, UpdateStatus says rollback_completed, and all three slots run web:2 again. Only one slot was ever on the bad version, because parallelism was 1. With the default pause action the update would have stopped instead, leaving one slot on the failed version and the service marked paused, until you run docker service rollback web or another update.
Set the update policy when you create the service (or in the stack file, as "Stacks, secrets and configs" does). Flags given to docker service update are part of the new spec, so a rollback reverts them together with the image. Here a service created with the defaults gets the broken image and the rollback policy in one update:
The update rolled back, and with it the failure action went back to the default pause. The next bad release on that service pauses instead of rolling back.
When running stays below desired
docker service ls shows 0/1 while a service is starting and also when it never will. The reason is in the task list:
The task stays Pending with no suitable node (scheduling constraints not satisfied on 3 nodes): no node carries the label the constraint asks for. The scheduler picks up the label on its next pass: three seconds after --label-add, the task is running on wrk2. Other common ERROR values are insufficient resources on N nodes (reservations larger than any node's free CPU or memory), No such image or pull access denied (a registry or credentials problem, see "Stacks, secrets and configs"), and task: non-zero exit (1) for a crashing process. Use --no-trunc, because the default output cuts the message.
Node maintenance: pause, drain, demote
With wrk1 paused, scaling to six placed no new task on it but left its existing web.1 alone. Draining wrk2 moved its two web tasks to mgr1 (the only active node left) and shut down the global agent there: drain stops every task on the node, global ones included, and the scheduler places replicated tasks elsewhere. Drain before you patch or reboot a node, wait until docker node ps <node> shows nothing running, do the work, then set it back:
Setting active again does not move tasks back; the scheduler only uses the node for new tasks, so rebalance with docker service update --force <service> if you need to.
Losing quorum
Promote both workers so the cluster has three managers, then freeze two of them with docker pause, which stops every process in a node container at once, the lab's closest equivalent of two hosts dropping off the network:
With one of three managers left there is no majority: docker node ls and every change fail with The swarm does not have a leader, and nothing can be scheduled, scaled or updated. The containers already running on mgr1 are untouched: quorum protects the cluster's ability to make decisions, not the workloads. A real outage would look the same, with the added risk that tasks on the lost nodes are not replaced until the managers agree again.
Once the frozen managers answer again, they re-elect a leader and the cluster accepts changes. The second command shows the trap with drain on a manager: mgr1 is Drain and still Leader. Draining a manager only stops it from running tasks; it keeps its vote. Drain one of three managers for maintenance, lose another, and quorum is gone. To take a manager out of the vote, demote it, and keep the count odd. If quorum is lost for good (two of three managers destroyed), the remaining manager can rebuild a one-manager cluster with docker swarm init --force-new-cluster, keeping services and tasks; then promote new managers. Make sure the lost managers really are gone, or wipe them before they rejoin, because a returning old manager still believes in the old cluster.
Autolock and backups
Raft data on each manager, under /var/lib/docker/swarm, is encrypted, including secrets. By default the key that decrypts it sits on the same disk, so anyone who copies a manager's disk gets everything. Autolock encrypts that key with an unlock key that you hold, and a manager that restarts stays locked until someone provides it:
After the restart, mgr1 refuses every cluster command with Swarm is encrypted and needs to be unlocked. docker swarm unlock reads the key from the terminal or, as here, from standard input. The workers may show Unknown for a few seconds afterwards, until their next heartbeat; in this run they were already Ready. The price is operational: an unattended reboot of a manager needs a person or a secret store to unlock it, and if every manager is locked and the key is lost, the cluster cannot be recovered. Keep the key in a password manager or vault, not next to the cluster, and rotate it with docker swarm unlock-key --rotate.
Back up a manager by stopping Docker on it and copying /var/lib/docker/swarm (on Docker 29 the swarm state is still under the daemon's data root), then starting Docker again. Do it on a non-leader manager so the cluster keeps quorum meanwhile. With autolock on, the copy is useless without the unlock key, so record which key belongs to which backup (stored separately). Restoring means putting that directory on a fresh host and running docker swarm init --force-new-cluster. Do that only when the old managers are gone for good: if any of them comes back, you have two clusters claiming the same identity and the same workers. The restored state includes the old node list, so remove nodes that no longer exist (docker node rm) and promote new managers. With autolock on, unlock the restored manager with the key that belongs to the backup before anything else. Volume data is not part of it.
Clean up
./swarm-lab.sh down switches back to the VM's own daemon and removes the three nodes with everything inside them, the registry and the network. The next two lessons start again from ./swarm-lab.sh up.
docker service update --image app:2 --update-failure-action rollback app, the new tasks fail their health check, and the service rolls back. A week later app:3 is also broken. What happens to that update?docker service ls has shown report 0/1 for ten minutes and docker service ps report shows the task as Pending. What do you do next?Try this
Work through “Clean up” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from swarm clusters and services, keep “Clean up”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.