Overlay networks and the routing mesh
VXLAN overlays, VIP vs DNS round robin, the ingress mesh, host mode and encryption.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/swarmnet.tar.gz && tar -xzf swarmnet.tar.gz, which creates ~/lab/swarmnet/. SHA-256: d51e95d042f7a1ad3655c2be98cd80a20339488b6a446aa9d733f8d278d897db./setup/create-lab.sh --profile sec and open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima); reset it with ./setup/create-lab.sh --profile sec --recreate. The daemon.json change follows "Configuring the daemon safely": back up, merge, validate, restore. On this VM ubuntu is not in the docker group, so commands for the VM's own daemon use sudo.Every node in the swarm reports Ready, DNS resolves the service name, and a request from a container on one node to a task on another still times out. Managers and workers talk over TCP 2377 and gossip over 7946, so the cluster looks healthy. Service traffic between nodes takes a different path, VXLAN over UDP port 4789, and a single firewall rule on that port breaks every cross-node connection while leaving the control plane untouched. This lesson follows that path: overlay networks, service discovery, the routing mesh, encrypted overlays, and the firewall requirements that go with them.
The lab is the simulated cluster from "Swarm clusters and services", started from ~/lab/swarmnet where the lesson files unpack: three docker:29-dind nodes (mgr1 10.77.0.11, wrk1 10.77.0.12, wrk2 10.77.0.13) on the VM bridge lab-swarm, reached through Docker contexts and built with swarm-lab.sh, which pulls any base image the VM lacks before copying it into the nodes. Because the nodes share the VM's kernel and bridge, the VM can capture their traffic and filter it, which is how the lesson reproduces the timeout further down. Between real hosts the same packets cross a physical network and its firewalls.
The last command builds the small nginx test image from that lesson, which answers with its version, its container hostname and the client address it sees, and pushes it to the lab registry.
An overlay network
An overlay network is a layer-2 network for containers that spans hosts. Each node attached to it gets a Linux bridge for the network inside a dedicated network namespace, plus a VXLAN interface; a frame from a container on one node is wrapped in a UDP packet, sent to the node that hosts the destination container, and unwrapped there. You create overlays on a manager. --attachable additionally allows standalone containers (docker run) to join, which is handy for debugging.
Before any task used appnet, wrk2 did not know about it; managers create overlays in the cluster state and extend them to a worker only when a task on that worker needs the network. ingress is the overlay Swarm creates for the routing mesh. appnet got the subnet 10.0.1.0/24 from Swarm's default pool (10.0.0.0/8 carved into /24s; change it with docker swarm init --default-addr-pool if it collides with your networks) and VXLAN network identifier 4097.
Names: a VIP, tasks., or DNS round robin
Docker's embedded DNS server (127.0.0.11 in every container) answers for services on networks the container shares with them. A debug container on appnet asks for api and for tasks.api:
api resolves to one address, 10.0.1.2, the service's virtual IP (VIP). tasks.api returns the two task addresses. Connections to the VIP are balanced across healthy tasks by IPVS in the kernel of the calling node, which is why four requests alternated between the two containers. The VIP stays the same while tasks are replaced, rescheduled or scaled, so clients that cache DNS answers keep working. Balancing is per connection, not per request: a client that holds one keep-alive connection open talks to one task.
With --endpoint-mode dnsrr there is no VIP (VirtualIPs is null) and the service name itself returns every task address. Clients then do their own balancing, and DNS caching becomes your problem: a client that resolved once and kept the answer keeps calling a task that may have moved. Use dnsrr for software that wants to see its peers (clustered databases, some proxies with their own health checking), and the default VIP mode otherwise. A dnsrr service cannot publish ports through the routing mesh; it can publish in host mode, shown below.
VXLAN on the wire
Capturing on the VM bridge while the debug container on mgr1 opens a connection to an api task:
Each capture line is two packets: the outer UDP packet from node to node (10.77.0.11 to 10.77.0.13, destination port 4789, VXLAN network identifier 4097, the ID appnet showed above) and, inside it, the container-to-container TCP SYN from the debug container (10.0.1.8) to the api task at 10.0.1.3, port 80, then the SYN-ACK coming back. The connection went to the VIP, but IPVS on mgr1 had already rewritten the destination to a task address before encapsulation. The inner traffic is not encrypted. Anyone who can capture on the network between the hosts reads it in plain text, which matters as soon as that network is shared or crosses a data centre you do not control.
Encapsulation costs 50 bytes per packet. The inner MSS of 1410 above comes from an inner MTU of 1450 on a 1500-byte network. Where the underlying MTU is smaller (some VPNs and cloud networks), set the overlay MTU explicitly with docker network create -d overlay --opt com.docker.network.driver.mtu=<value>, or small requests work and large responses hang.
A firewall that drops 4789/udp
Now the opening symptom. A rule in the VM's DOCKER-USER chain drops UDP port 4789 on the bridge between the nodes, which is what a host firewall or a security group without that port does between real hosts. Node status and DNS still work, and the TCP connection to the VIP times out:
Delete the rule and the same request goes through:
The same picture appears when a cloud security group allows 2377 but not 4789/udp, when 7946 is blocked (node discovery then flaps and overlay peers are not learned), or when a host firewall on one node filters forwarded traffic. Between all nodes, allow 7946/tcp, 7946/udp and 4789/udp; towards managers, 2377/tcp; and IP protocol 50 (ESP) when you use encrypted overlays. Keep all of them closed to everything outside the cluster: VXLAN has no authentication, so anyone who can send packets to 4789 on a node can inject traffic into its overlays.
The routing mesh
A port published with -p goes on the ingress overlay. Every node in the swarm listens on it, whether or not it runs a task, and forwards the connection to a healthy task on any node:
The single web task runs on wrk1, yet all three node addresses answer, and the reply comes from the same container each time. The client address the application sees is 10.0.0.2, 10.0.0.3 or 10.0.0.4, the ingress-network address of the node that received the connection, not the VM's address. The mesh applies source NAT so that replies return through the node that accepted the connection. That is convenient behind an external load balancer that knows nothing about task placement, and it means published ports are open on every node and the application loses the real client address, which breaks IP allow-lists, rate limits and access logs.
Host-mode publishing binds the port only on nodes that run a task, directly to the task, without the mesh and without NAT:
mgr1 runs no edge task and refuses the connection (curl exit 7). The two workers answer with the client address 10.77.0.1, the VM's own address on the bridge, which is the real source. docker service ls lists only mesh ports in its PORTS column, so host-mode ports do not show up there; check the service spec instead. Host mode pairs naturally with global services: one task per node, each node publishes the port, and a load balancer that health-checks the nodes, or a proxy such as Traefik or HAProxy running as that global service, sends traffic to them. With a replicated service in host mode, two tasks of the same service cannot land on one node (the port is taken), and the load balancer has to follow task placement.
Encrypted overlays
--opt encrypted makes Swarm set up IPsec (ESP in transport mode, AES-GCM) between every pair of nodes that share the network. The managers generate the keys and rotate them every 12 hours.
The same kind of request now shows ESP packets between the nodes, with no VXLAN header and no inner addresses visible. Encryption is per network and off by default, and it costs CPU and throughput on every packet; measure it on your hardware before turning it on for heavy east-west traffic. Firewalls between nodes must pass IP protocol 50, which many security group templates do not include, and the failure looks exactly like the 4789 drop above. Encryption covers overlay traffic only: the routing mesh's ingress network is not encrypted by this option, and traffic from clients to published ports is whatever the application makes it, so terminate TLS in the application or the proxy as well. On Windows nodes, encrypted overlays are not supported.
Swarm and the nftables firewall backend
Docker 29 added an experimental nftables firewall backend ("firewall-backend": "nftables" in daemon.json), which "Publishing ports and the packet path" covers. The overlay driver's rules have not been migrated, so the daemon refuses to enter swarm mode with it. The first command backs up daemon.json (an empty {} when there is none, as here), merges the key into the backup with jq, validates the result and restarts; the last one restores the backup:
If you moved a host to the nftables backend, it cannot be a swarm node until you switch back to iptables, the default; the reverse also holds, since a daemon in swarm mode keeps iptables. Plan the backend per host role.
-p 80:80 logs every request as coming from 10.0.0.x, so its IP allow-list blocks real customers. What change gives nginx the real client address?Try this
Work through “Swarm and the nftables firewall backend” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from overlay networks and the routing mesh, keep “Swarm and the nftables firewall backend”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.