Overlay networks and the routing mesh

VXLAN overlays, VIP vs DNS round robin, the ingress mesh, host mode and encryption.

Intermediate18 min · lesson 17 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (3 files, 2 KB): swarmnet.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/swarmnet.tar.gz && tar -xzf swarmnet.tar.gz, which creates ~/lab/swarmnet/. SHA-256: d51e95d042f7a1ad3655c2be98cd80a20339488b6a446aa9d733f8d278d897db
Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec). The lab simulates a three-node swarm with privileged Docker-in-Docker containers, adds and removes an iptables rule on the VM, and at the end changes /etc/docker/daemon.json and restarts dockerd. If the VM does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec and open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima); reset it with ./setup/create-lab.sh --profile sec --recreate. The daemon.json change follows "Configuring the daemon safely": back up, merge, validate, restore. On this VM ubuntu is not in the docker group, so commands for the VM's own daemon use sudo.

Every node in the swarm reports Ready, DNS resolves the service name, and a request from a container on one node to a task on another still times out. Managers and workers talk over TCP 2377 and gossip over 7946, so the cluster looks healthy. Service traffic between nodes takes a different path, VXLAN over UDP port 4789, and a single firewall rule on that port breaks every cross-node connection while leaving the control plane untouched. This lesson follows that path: overlay networks, service discovery, the routing mesh, encrypted overlays, and the firewall requirements that go with them.

The lab is the simulated cluster from "Swarm clusters and services", started from ~/lab/swarmnet where the lesson files unpack: three docker:29-dind nodes (mgr1 10.77.0.11, wrk1 10.77.0.12, wrk2 10.77.0.13) on the VM bridge lab-swarm, reached through Docker contexts and built with swarm-lab.sh, which pulls any base image the VM lacks before copying it into the nodes. Because the nodes share the VM's kernel and bridge, the VM can capture their traffic and filter it, which is how the lesson reproduces the timeout further down. Between real hosts the same packets cross a physical network and its firewalls.

ubuntu@secopslog-docker-sec:~/lab/swarmnet · Docker 29.8.2
$ ./swarm-lab.sh up
Successfully created context "lab-mgr1" lab-mgr1 10.77.0.11 29.8.2 ready Successfully created context "lab-wrk1" lab-wrk1 10.77.0.12 29.8.2 ready Successfully created context "lab-wrk2" lab-wrk2 10.77.0.13 29.8.2 ready lab-reg:5000 has: {"repositories":["alpine","postgres"]}
$ docker context use lab-mgr1 docker swarm init --advertise-addr 10.77.0.11 >/dev/null for n in lab-wrk1 lab-wrk2; do docker --context $n swarm join --token "$(docker swarm join-token -q worker)" 10.77.0.11:2377; done docker build -q -t lab-reg:5000/web:1 web/ >/dev/null && docker push -q lab-reg:5000/web:1
lab-mgr1 Current context is now "lab-mgr1" This node joined a swarm as a worker. This node joined a swarm as a worker. lab-reg:5000/web:1

The last command builds the small nginx test image from that lesson, which answers with its version, its container hostname and the client address it sees, and pushes it to the lab registry.

An overlay network

An overlay network is a layer-2 network for containers that spans hosts. Each node attached to it gets a Linux bridge for the network inside a dedicated network namespace, plus a VXLAN interface; a frame from a container on one node is wrapped in a UDP packet, sent to the node that hosts the destination container, and unwrapped there. You create overlays on a manager. --attachable additionally allows standalone containers (docker run) to join, which is handy for debugging.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker network create -d overlay --attachable appnet docker network ls --filter driver=overlay
ccsft7ofh60aowk2sg5n4wt9k NETWORK ID NAME DRIVER SCOPE ccsft7ofh60a appnet overlay swarm 5r5apmm6foh0 ingress overlay swarm
$ docker --context lab-wrk2 network ls --filter driver=overlay
NETWORK ID NAME DRIVER SCOPE 5r5apmm6foh0 ingress overlay swarm
$ docker service create -q --name api --network appnet --replicas 2 \ --constraint node.role==worker lab-reg:5000/web:1 docker service ps api --format "{{.Name}} {{.Node}} {{.CurrentState}}" docker --context lab-wrk2 network ls --filter driver=overlay
dg0og2p8c37wx21zhnka75ex3 api.1 wrk1 Running 5 seconds ago api.2 wrk2 Running 5 seconds ago NETWORK ID NAME DRIVER SCOPE ccsft7ofh60a appnet overlay swarm 5r5apmm6foh0 ingress overlay swarm
$ docker network inspect appnet --format "{{.Driver}} {{json .IPAM.Config}} {{json .Options}}"
overlay [{"Subnet":"10.0.1.0/24","Gateway":"10.0.1.1"}] {"com.docker.network.driver.overlay.vxlanid_list":"4097"}

Before any task used appnet, wrk2 did not know about it; managers create overlays in the cluster state and extend them to a worker only when a task on that worker needs the network. ingress is the overlay Swarm creates for the routing mesh. appnet got the subnet 10.0.1.0/24 from Swarm's default pool (10.0.0.0/8 carved into /24s; change it with docker swarm init --default-addr-pool if it collides with your networks) and VXLAN network identifier 4097.

Names: a VIP, tasks., or DNS round robin

Docker's embedded DNS server (127.0.0.11 in every container) answers for services on networks the container shares with them. A debug container on appnet asks for api and for tasks.api:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker run -d --name debug --network appnet lab-reg:5000/alpine:3.22 sleep 1d
a85f83525397621dd021f01a24b6bbaa888048bf8811edb81f3d1ed2690b3209
$ docker exec debug nslookup -type=a api docker exec debug nslookup -type=a tasks.api
Server: 127.0.0.11 Address: 127.0.0.11:53 Non-authoritative answer: Name: api Address: 10.0.1.2 Server: 127.0.0.11 Address: 127.0.0.11:53 Non-authoritative answer: Name: tasks.api Address: 10.0.1.3 Name: tasks.api Address: 10.0.1.4
$ for i in 1 2 3 4; do docker exec debug wget -qO- http://api/; done
web v1 on a928836a6806, client 10.0.1.8 web v1 on 610c3dcf8d8e, client 10.0.1.8 web v1 on a928836a6806, client 10.0.1.8 web v1 on 610c3dcf8d8e, client 10.0.1.8

api resolves to one address, 10.0.1.2, the service's virtual IP (VIP). tasks.api returns the two task addresses. Connections to the VIP are balanced across healthy tasks by IPVS in the kernel of the calling node, which is why four requests alternated between the two containers. The VIP stays the same while tasks are replaced, rescheduled or scaled, so clients that cache DNS answers keep working. Balancing is per connection, not per request: a client that holds one keep-alive connection open talks to one task.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker service create -q --name cache --network appnet --endpoint-mode dnsrr --replicas 2 lab-reg:5000/alpine:3.22 sleep 1d docker exec debug nslookup -type=a cache
jf043uceqrmlb24mz9dj4bzwk Server: 127.0.0.11 Address: 127.0.0.11:53 Non-authoritative answer: Name: cache Address: 10.0.1.9 Name: cache Address: 10.0.1.10
$ docker service inspect api --format "{{json .Endpoint.VirtualIPs}}" docker service inspect cache --format "{{json .Endpoint.VirtualIPs}} {{.Spec.EndpointSpec.Mode}}"
[{"NetworkID":"ccsft7ofh60aowk2sg5n4wt9k","Addr":"10.0.1.2/24"}] null dnsrr

With --endpoint-mode dnsrr there is no VIP (VirtualIPs is null) and the service name itself returns every task address. Clients then do their own balancing, and DNS caching becomes your problem: a client that resolved once and kept the answer keeps calling a task that may have moved. Use dnsrr for software that wants to see its peers (clustered databases, some proxies with their own health checking), and the default VIP mode otherwise. A dnsrr service cannot publish ports through the routing mesh; it can publish in host mode, shown below.

VXLAN on the wire

Capturing on the VM bridge while the debug container on mgr1 opens a connection to an api task:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ BR=br-$(sudo docker network inspect -f "{{.Id}}" lab-swarm | cut -c1-12) sudo timeout 10 tcpdump -ni $BR -c 2 "udp port 4789" 2>/dev/null & sleep 1; docker exec debug wget -qO- http://api/ >/dev/null; wait
02:57:52.834183 IP 10.77.0.11.56943 > 10.77.0.12.4789: VXLAN, flags [I] (0x08), vni 4097 IP 10.0.1.8.54330 > 10.0.1.3.80: Flags [S], seq 2256122492, win 64860, options [mss 1410,sackOK,TS val 319566991 ecr 0,nop,wscale 9], length 0 02:57:52.834342 IP 10.77.0.12.55107 > 10.77.0.11.4789: VXLAN, flags [I] (0x08), vni 4097 IP 10.0.1.3.80 > 10.0.1.8.54330: Flags [S.], seq 2748574456, ack 2256122493, win 64308, options [mss 1410,sackOK,TS val 978162796 ecr 319566991,nop,wscale 9], length 0

Each capture line is two packets: the outer UDP packet from node to node (10.77.0.11 to 10.77.0.13, destination port 4789, VXLAN network identifier 4097, the ID appnet showed above) and, inside it, the container-to-container TCP SYN from the debug container (10.0.1.8) to the api task at 10.0.1.3, port 80, then the SYN-ACK coming back. The connection went to the VIP, but IPVS on mgr1 had already rewritten the destination to a task address before encapsulation. The inner traffic is not encrypted. Anyone who can capture on the network between the hosts reads it in plain text, which matters as soon as that network is shared or crosses a data centre you do not control.

Encapsulation costs 50 bytes per packet. The inner MSS of 1410 above comes from an inner MTU of 1450 on a 1500-byte network. Where the underlying MTU is smaller (some VPNs and cloud networks), set the overlay MTU explicitly with docker network create -d overlay --opt com.docker.network.driver.mtu=<value>, or small requests work and large responses hang.

A firewall that drops 4789/udp

Now the opening symptom. A rule in the VM's DOCKER-USER chain drops UDP port 4789 on the bridge between the nodes, which is what a host firewall or a security group without that port does between real hosts. Node status and DNS still work, and the TCP connection to the VIP times out:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -I DOCKER-USER -p udp --dport 4789 -j DROP docker node ls --format "{{.Hostname}} {{.Status}}" docker exec debug nslookup -type=a api | tail -n 2 docker exec debug wget -T 3 -qO- http://api/ || echo "wget exit $?"
mgr1 Ready wrk1 Ready wrk2 Ready Address: 10.0.1.2 wget: download timed out wget exit 1

Delete the rule and the same request goes through:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -D DOCKER-USER -p udp --dport 4789 -j DROP docker exec debug wget -T 3 -qO- http://api/
web v1 on a928836a6806, client 10.0.1.8

The same picture appears when a cloud security group allows 2377 but not 4789/udp, when 7946 is blocked (node discovery then flaps and overlay peers are not learned), or when a host firewall on one node filters forwarded traffic. Between all nodes, allow 7946/tcp, 7946/udp and 4789/udp; towards managers, 2377/tcp; and IP protocol 50 (ESP) when you use encrypted overlays. Keep all of them closed to everything outside the cluster: VXLAN has no authentication, so anyone who can send packets to 4789 on a node can inject traffic into its overlays.

The routing mesh

A port published with -p goes on the ingress overlay. Every node in the swarm listens on it, whether or not it runs a task, and forwards the connection to a healthy task on any node:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker service create -q --name web --replicas 1 --constraint node.hostname==wrk1 -p 8080:80 lab-reg:5000/web:1 for ip in 10.77.0.11 10.77.0.12 10.77.0.13; do curl -s http://$ip:8080/; done
nt6677emvy7cggaepcpwnr1o0 web v1 on 5a25191889d3, client 10.0.0.2 web v1 on 5a25191889d3, client 10.0.0.3 web v1 on 5a25191889d3, client 10.0.0.4
$ docker network ls --filter name=ingress docker network inspect ingress --format "{{json .IPAM.Config}}"
NETWORK ID NAME DRIVER SCOPE 5r5apmm6foh0 ingress overlay swarm [{"Subnet":"10.0.0.0/24","Gateway":"10.0.0.1"}]

The single web task runs on wrk1, yet all three node addresses answer, and the reply comes from the same container each time. The client address the application sees is 10.0.0.2, 10.0.0.3 or 10.0.0.4, the ingress-network address of the node that received the connection, not the VM's address. The mesh applies source NAT so that replies return through the node that accepted the connection. That is convenient behind an external load balancer that knows nothing about task placement, and it means published ports are open on every node and the application loses the real client address, which breaks IP allow-lists, rate limits and access logs.

Host-mode publishing binds the port only on nodes that run a task, directly to the task, without the mesh and without NAT:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker service create -q --name edge --mode global --constraint node.role==worker \ --publish mode=host,published=8081,target=80 lab-reg:5000/web:1 for ip in 10.77.0.11 10.77.0.12 10.77.0.13; do curl -s -m 3 http://$ip:8081/ || echo "$ip: curl exit $?"; done
vzz2fnvs7d8y1zibcvyv28339 10.77.0.11: curl exit 7 web v1 on e034f36f15ca, client 10.77.0.1 web v1 on d04b8c5a0de2, client 10.77.0.1
$ docker service ls --format "{{.Name}} {{.Ports}}"
api cache edge web *:8080->80/tcp

mgr1 runs no edge task and refuses the connection (curl exit 7). The two workers answer with the client address 10.77.0.1, the VM's own address on the bridge, which is the real source. docker service ls lists only mesh ports in its PORTS column, so host-mode ports do not show up there; check the service spec instead. Host mode pairs naturally with global services: one task per node, each node publishes the port, and a load balancer that health-checks the nodes, or a proxy such as Traefik or HAProxy running as that global service, sends traffic to them. With a replicated service in host mode, two tasks of the same service cannot land on one node (the port is taken), and the load balancer has to follow task placement.

Encrypted overlays

--opt encrypted makes Swarm set up IPsec (ESP in transport mode, AES-GCM) between every pair of nodes that share the network. The managers generate the keys and rotate them every 12 hours.

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ docker network create -d overlay --opt encrypted --attachable secnet docker service create -q --name sapi --network secnet --replicas 1 --constraint node.hostname==wrk2 lab-reg:5000/web:1 docker run -d --name sdebug --network secnet lab-reg:5000/alpine:3.22 sleep 1d >/dev/null docker exec sdebug wget -qO- http://sapi/
fge32wajlwdsq5wwqm221yi7i m09vnn4ryk3gdwzrcnz8bmxa9 web v1 on 273c270df967, client 10.0.2.6
$ BR=br-$(sudo docker network inspect -f "{{.Id}}" lab-swarm | cut -c1-12) sudo timeout 10 tcpdump -ni $BR -c 2 "esp or udp port 4789" 2>/dev/null & sleep 1; docker exec sdebug wget -qO- http://sapi/ >/dev/null; wait
02:58:22.406478 IP 10.77.0.11 > 10.77.0.13: ESP(spi=0xeec1ea64,seq=0x6), length 116 02:58:22.406547 IP 10.77.0.13 > 10.77.0.11: ESP(spi=0x82b241cc,seq=0x6), length 116

The same kind of request now shows ESP packets between the nodes, with no VXLAN header and no inner addresses visible. Encryption is per network and off by default, and it costs CPU and throughput on every packet; measure it on your hardware before turning it on for heavy east-west traffic. Firewalls between nodes must pass IP protocol 50, which many security group templates do not include, and the failure looks exactly like the 4789 drop above. Encryption covers overlay traffic only: the routing mesh's ingress network is not encrypted by this option, and traffic from clients to published ports is whatever the application makes it, so terminate TLS in the application or the proxy as well. On Windows nodes, encrypted overlays are not supported.

ubuntu@secopslog-docker-sec:~/lab/swarmnet · Docker 29.8.2
$ docker rm -f debug sdebug >/dev/null docker context use default ./swarm-lab.sh down
default Current context is now "default" swarm lab removed

Swarm and the nftables firewall backend

Docker 29 added an experimental nftables firewall backend ("firewall-backend": "nftables" in daemon.json), which "Publishing ports and the packet path" covers. The overlay driver's rules have not been migrated, so the daemon refuses to enter swarm mode with it. The first command backs up daemon.json (an empty {} when there is none, as here), merges the key into the backup with jq, validates the result and restarts; the last one restores the backup:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo cp -a /etc/docker/daemon.json /etc/docker/daemon.json.bak 2>/dev/null || echo '{}' | sudo tee /etc/docker/daemon.json.bak >/dev/null jq '. + {"firewall-backend": "nftables"}' /etc/docker/daemon.json.bak | sudo tee /etc/docker/daemon.json sudo dockerd --validate --config-file /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format '{{json .FirewallBackend}}'
{ "firewall-backend": "nftables" } configuration OK {"Driver":"nftables","Info":[["EnableUserlandProxy","true"],["UserlandProxyPath","/usr/bin/docker-proxy"]]}
$ sudo docker swarm init --advertise-addr 127.0.0.1
Error response from daemon: --firewall-backend=nftables is incompatible with swarm mode
$ sudo cp /etc/docker/daemon.json.bak /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format "{{.FirewallBackend.Driver}}"
iptables

If you moved a host to the nftables backend, it cannot be a swarm node until you switch back to iptables, the default; the reverse also holds, since a daemon in swarm mode keeps iptables. Plan the backend per host role.

Quick check
01A new three-node swarm forms cleanly and every node is Ready. A frontend task on one node cannot reach a backend task on another, although DNS returns the backend's VIP. The security group allows 2377/tcp between the nodes. What is missing?
Correct — 2377 carries cluster management only. Gossip uses 7946 and the VXLAN data plane 4789/udp; the lab dropped 4789 and got exactly this symptom.
Incorrect — Services on a shared overlay reach each other without publishing; publishing is for clients outside the swarm.
Incorrect — --attachable only lets standalone containers join; service tasks are attached regardless.
Incorrect — VIPs work on every node; the lab's debug container on mgr1 reached tasks on both workers through the VIP.
02An nginx service published with -p 80:80 logs every request as coming from 10.0.0.x, so its IP allow-list blocks real customers. What change gives nginx the real client address?
Incorrect — Endpoint mode changes how other services resolve the name; it does not change how external clients reach a published port.
Incorrect — Encryption protects traffic between nodes; the mesh still applies source NAT.
Incorrect — With the mesh, a request can still be forwarded and NATed, whether or not the receiving node has a task.
Correct — Host mode binds the port on the node running the task, without the mesh and its source NAT; the lab showed the real address 10.77.0.1.
03Clients of a dnsrr service keep sending traffic to a task address for minutes after that task was rescheduled to another node. Why?
Incorrect — The embedded DNS answers with the current task addresses; the stale address comes from the client side.
Incorrect — dnsrr services do not use a VIP or the mesh at all.
Correct — In VIP mode the address never changes; with dnsrr the name maps straight to task addresses, so client-side DNS caching serves stale ones.
Incorrect — The problem is the cached name-to-address answer, not ARP.

Try this

Work through “Swarm and the nftables firewall backend” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from overlay networks and the routing mesh, keep “Swarm and the nftables firewall backend”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related