Publishing ports and the packet path
From a client to a container and back: NAT, docker-proxy, IPv6, firewalls and the nftables backend.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/ports.tar.gz && tar -xzf ports.tar.gz, which creates ~/lab/ports/. SHA-256: 5d18c8d6d8754ccb41f7747fb6c14601d6047fc3c64afc1f75afd97cca8f30edsecopslog-docker) and only reads firewall rules. The second half changes the host firewall, daemon.json and the firewall backend: run it only in the SecOpsLog disposable lab VM (secopslog-docker-sec), never on a host anyone depends on. If it does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec with Lima), and get back to a clean state at any point with ./setup/create-lab.sh --profile sec --recreate. Each terminal below names the VM it runs on. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo.One -p flag, two listeners:
docker port reports the mapping twice, once for 0.0.0.0 and once for [::]: a published port without a host address is published on every IPv4 and every IPv6 address of the host, and ss shows a docker-proxy process listening on each. "Container networking basics" (Docker for beginners) showed how to bind 127.0.0.1 instead. This lesson follows a packet from a client to the container and back, shows which part of Docker handles which kind of client, and covers what that means for host firewalls, IPv6 and the new nftables backend. Reading the rules needs root, so the commands use sudo; they are filtered to this lab's network because every other network on a host adds rules of its own.
The packet path
Here is the nat table behind that diagram:
PREROUTING sends every packet addressed to one of the host's own addresses into Docker's DOCKER chain, and OUTPUT does the same for connections the host itself opens, except to 127.0.0.0/8. In DOCKER, the DNAT rule rewrites TCP port 8088 to 10.89.10.2:80 for packets that did not come from the container's own bridge. POSTROUTING masquerades anything leaving 10.89.10.0/24 through another interface: that is how containers reach the internet with the host's address. Now the filter table:
FORWARD jumps first to DOCKER-USER, the chain Docker creates for your rules and never fills (since Docker 28.2.2 it has no built-in RETURN rule, so you can append as well as insert), then to DOCKER-FORWARD. For this bridge, DOCKER-CT lets replies of established connections back in, DOCKER-FORWARD lets anything the containers send out, and DOCKER accepts new connections from outside to exactly 10.89.10.2 port 80 and drops every other new connection into the bridge. The FORWARD policy is DROP on this host, so a forwarded packet nothing accepts is dropped. Docker sets that policy when it turns IP forwarding on itself, so that enabling forwarding does not turn the host into a router for its other interfaces. Two more tables complete the picture:
The raw-table rule drops packets addressed to the container's own IP unless they come from its bridge. It runs before NAT, so it only catches packets sent to 10.89.10.2 directly, which is what a neighbour that routes to the container subnet would do. The IPv6 DOCKER chain is empty because lab-pub is IPv4-only. The [::]:8088 listener still works, as the next section shows, but no IPv6 NAT rule is involved.
Three ways in
All three requests worked, and nginx's log shows two different client addresses. The request to the host's own address, 192.168.2.4 here (yours will differ), went through the OUTPUT DNAT rule and kept its source address. The requests to 127.0.0.1 and [::1] came from 10.89.10.1, the bridge address: they were accepted by docker-proxy, which opened a new connection to the container. Loopback is excluded from the DNAT rules, and there is no IPv6 rule at all on an IPv4-only network, so for those clients the userland proxy is the path. conntrack shows the difference:
Each conntrack entry shows the original direction first and the expected reply second. The connection to 192.168.2.4:8088 expects its reply from 10.89.10.2 port 80: that is the DNAT. The 127.0.0.1 connection stops at 127.0.0.1:8088, where docker-proxy accepted it, and the second command shows the proxy's own connections from 10.89.10.1 to the container. conntrack -L lists IPv4 unless you add -f ipv6, which is why the [::1] client does not appear in the first list. TIME_WAIT counters and ports differ on every run. Traffic in the other direction is masqueraded:
The container connected from 10.89.10.2; the reply is expected at 192.168.2.4, the host's address. The remote server never learns the container's address.
Choosing the address
Still on the main VM, publish a second container on 127.0.0.1 only, and a third with -P:
A port published on 127.0.0.1 gets one mapping, refuses connections to the host's LAN address (curl exit 7), and a raw rule drops packets for 127.0.0.1:8089 that arrive on any interface other than lo. Before Docker 28.0, neighbours on the same layer-2 segment could reach ports published on 127.0.0.1 by sending packets for 127.0.0.1 to the host's MAC address; the release notes list that rule as a security fix. -P publishes every EXPOSEd port on a free port from the ephemeral range, which starts at 32768, again on both address families. To change the default for every container on a host, set "ip": "127.0.0.1" in daemon.json ("Configuring the daemon safely" covers the procedure); explicit -p addresses still win.
That is all for the main VM. Remove its containers and network:
What a neighbour can reach
Everything from here on runs on the sec VM, secopslog-docker-sec, as ubuntu with sudo for docker. Get the lesson files there too; they unpack to ~/lab/ports, where these steps run. On a single VM there is no second machine to test from, so the lesson builds one: a network namespace joined to the host by a veth pair, with an address on a separate subnet and the host as its default gateway. To the host it is a LAN neighbour on interface lab-lan0.
#!/bin/sh# Simulates a second machine on the Docker host's LAN: a network namespace "lab-lan" joined to the host# by a veth pair. Host side: lab-lan0, 192.0.2.1/24. Neighbour: 192.0.2.10/24, default route via the host.# Usage: sudo ./lan-neighbour.sh up|down Then: sudo ip netns exec lab-lan <command>[ -f /etc/secopslog-lab ] || { echo "refusing: run this only in the SecOpsLog lab VM" >&2; exit 2; }set -eucase "${1:-}" inup)ip netns add lab-lanip link add lab-lan0 type veth peer name eth0 netns lab-lanip addr add 192.0.2.1/24 dev lab-lan0ip link set lab-lan0 upip -n lab-lan link set lo upip -n lab-lan addr add 192.0.2.10/24 dev eth0ip -n lab-lan link set eth0 upip -n lab-lan route add default via 192.0.2.1;;down)ip netns del lab-lan 2>/dev/null || trueip link del lab-lan0 2>/dev/null || true;;*)echo "usage: $0 up|down" >&2exit 2;;esac
The neighbour reached the published port, and nginx logged its real address 192.0.2.10: inbound DNAT does not hide the client, so access logs and allow-lists in the application keep working. Going straight to the container's address timed out, although the neighbour routes 10.89.10.0/24 through the host. That is the raw-table rule. Before Docker 28.0, a neighbour that routed to the container subnet could reach published ports directly on the container IP; the 28.0 release notes list that change among the security fixes. lan-neighbour.sh refuses to run outside the lab VM, because it adds interfaces and routes to the host.
Gateway modes
NAT is the default way a bridge network meets the outside, not the only one. The bridge driver option com.docker.network.bridge.gateway_mode_ipv4 (and _ipv6) selects nat (the default), nat-unprotected (NAT, but any container port is reachable by direct routing), routed (no NAT: the container addresses are meant to be routed to) or isolated (no gateway address on the host at all). routed is the mode to use when your network routes a container subnet to the host:
The container listens on ports 80 and 81 and publishes only 80. The neighbour reached 10.89.11.2:80 directly, port 81 timed out, and the host port 8081 refused the connection. There are no NAT rules for the subnet and no raw drop; the filter rule that accepts port 80 is still there, so "published" now means "allowed through", not "mapped to a host port". docker port shows the IPv4 mapping with an empty host port for the same reason. The [::]:8081 line is docker-proxy listening on the host's IPv6 addresses for this IPv4-only network; the routed mode applies to IPv4 only. The daemon option allow-direct-routing removes the direct-routing protection for every network; prefer per-network modes.
Host firewalls: ufw and firewalld
Docker's install documentation warns that if you use ufw or firewalld, ports you publish bypass your firewall rules. Watch it happen:
ufw is active with incoming traffic denied and only SSH allowed, and the neighbour still gets a 200 from port 8080. The diagram explains it: ufw's "incoming" rules live in the INPUT chain, and a DNATed packet is forwarded, never input. In FORWARD, Docker's jumps come before ufw's chains, and DOCKER-FORWARD accepts the packet before ufw's routed policy is consulted. firewalld behaves similarly from the other side: Docker puts its bridges in a firewalld zone called docker with target ACCEPT. Disabling ufw at the end also reset the built-in chain policies: FORWARD now reads ACCEPT, where it was DROP above, which matters again in the nftables section. The fixes are to not publish what should not be public (bind 127.0.0.1, or put the service behind a reverse proxy), to filter in front of the host (cloud security groups), or to put your rules where forwarded packets go: DOCKER-USER.
A rule for --dport 8080 matched nothing: its packet counter is 0. By the time a packet reaches DOCKER-USER, DNAT has already rewritten it to the container address and port 80. Match the original destination through conntrack instead:
With -m conntrack --ctorigdstport 8080 the rule matches (3 packets: the SYN and its retransmissions), the neighbour times out, and the host itself still gets a 200 because its loopback connection goes through docker-proxy and never crosses FORWARD. --ctorigdst matches the original destination address the same way. Scope rules with -i (the external interface) so they do not catch traffic between containers, and put an ACCEPT for RELATED,ESTABLISHED first when you build an allow-list. Rules added with iptables are lost on reboot; persist them with your configuration management or a unit that runs after Docker starts. Persist only your DOCKER-USER rules: do not iptables-save the whole ruleset into iptables-persistent, because restoring Docker's own chains at boot conflicts with the rules dockerd writes when it starts. The policy side of this, which ports to allow from where and how to control egress, belongs to "Container network hardening" (Advanced container security).
userland-proxy false, and IPv6 networks
docker-proxy costs a process per published port and address family and hides the client address for loopback and IPv6 clients. Setting "userland-proxy": false replaces it with kernel rules where it can. It is a daemon setting, so this uses the procedure from "Configuring the daemon safely": back up the current file (on this VM there is none, so the backup is an empty {}), merge the new key into the backup with jq, validate, restart, and at the end restore the backup. If your VM was created with --mtu, the backup holds the kit's MTU settings and the merge keeps them. This part restarts dockerd three times; systemd allows only a few restarts of docker.service in a short window, so if a restart fails with start request repeated too quickly, wait a minute or run sudo systemctl reset-failed docker docker.socket (the recorded run resets the counter between restarts).
With the proxy off, the IPv4-only network got only an IPv4 mapping: no process exists to carry IPv6 clients to an IPv4 container, so the host's IPv6 address no longer reaches it (000). Loopback IPv4 still works, now without a proxy process. The --ipv6 network got a ULA subnet (an fd00::/8 prefix Docker picks when you give none), mappings on both families, and an ip6tables DNAT rule, so IPv6 clients reach lab-web6 without a proxy. The rule's ! -s fe80::/10 keeps link-local traffic out of it. If IPv6 clients must reach a service, give its network IPv6; publishing alone only covers them through docker-proxy.
The nftables firewall backend
Docker 29.0 added "firewall-backend": "nftables", experimental in 29.x: Docker writes native nftables rules instead of iptables rules. It changes how you add your own rules, so try it before you plan on it:
Docker now owns two nftables tables, ip docker-bridges and ip6 docker-bridges, and its iptables chains are gone except the old jump from FORWARD to DOCKER-USER, which stays until it is removed or the host reboots. Two things are now yours. Docker does not enable IP forwarding with this backend; it reports an error when a network needs forwarding and it is off. It is 1 here only because the earlier iptables-backend start enabled it, so on a migrated host set net.ipv4.ip_forward=1 and net.ipv6.conf.all.forwarding=1 in a sysctl.d file. And an iptables FORWARD chain with a DROP policy still drops packets that Docker's nftables rules accepted. The iptables backend sets that policy when it enables forwarding itself, as the main VM showed; here the policy already reads ACCEPT because disabling ufw earlier reset it. On a host that switched backends without a reboot, run sudo iptables -P FORWARD ACCEPT and sudo ip6tables -P FORWARD ACCEPT.
Forwarding on and nothing in iptables dropping it means the host now forwards between all of its interfaces, not only Docker's bridges. On a host with more than one network, a neighbour that uses this host as its gateway can reach the other networks through it. Docker's documentation asks you to add firewall rules that block unwanted forwarding between non-Docker interfaces before you enable forwarding, and to keep such blocking on any multi-homed host that is not meant to be a router. A table of your own does it; for a host with interfaces eth0 and eth1:
table inet no-ext-forwarding {chain forward {type filter hook forward priority filter; policy accept;iifname "eth0" oifname "eth1" dropiifname "eth1" oifname "eth0" drop}}
Docker's own chains still accept what goes to and from its bridges, and your drops between the external interfaces stop everything else. Publish a port and look at the rules:
The same logic in a different shape: a DNAT rule for port 8080, a raw-priority drop for direct access, and a per-bridge chain that accepts established traffic, traffic between containers on the bridge (ICC), the published port, and drops the rest (UNPUBLISHED PORT DROP). There is no DOCKER-USER. Your rules go in a table of your own, with a base chain on the hook you need:
# Your own nftables rules live in your own table. Docker never touches it.table inet lab-filter {chain forward {type filter hook forward priority filter - 10; policy accept;iifname "lab-lan0" ct original proto-dst 8080 counter drop}}
ct original proto-dst 8080 plays the role of --ctorigdstport. In nftables a drop in any base chain is final, whatever the priority, so your drop wins over Docker's accept; the priority (filter - 10, before Docker's chain) only decides the order. The reverse is not true: an accept in your table does not override a drop in Docker's. For that, the daemon option bridge-accept-fwmark lets packets carrying a firewall mark you set be accepted by Docker's chains. One limit decides the backend for some hosts:
Swarm mode is incompatible with the nftables backend, because the overlay network rules have not been migrated. Roll back and confirm the iptables backend is active again:
Rootless Docker publishes ports through RootlessKit's port driver rather than through the host rules shown here; "Rootless Docker" (Advanced container security) covers how its ports and client addresses behave.
-p 5432:5432 on a server where ufw denies all incoming traffic except SSH. A scanner on the LAN still connects to 5432. Where would a blocking rule actually take effect?iptables -I DOCKER-USER -p tcp --dport 8080 -j DROP has a packet counter of 0, and the published port 8080 (container port 80) is still reachable. Why?"userland-proxy": false. A service on an IPv4-only network published with -p 8080:80 stops answering on the host's IPv6 address but still answers on IPv4. What change brings IPv6 clients back without the proxy?Try this
Work through “The nftables firewall backend” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from publishing ports and the packet path, keep “The nftables firewall backend”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.