Container network hardening

Internal networks, an egress proxy, DOCKER-USER policy and the nftables backend.

Advanced18 min · lesson 17 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (4 files, 3 KB): nethard.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/nethard.tar.gz && tar -xzf nethard.tar.gz, which creates ~/lab/nethard/. SHA-256: 92e9f24f7e04884464ef0e2e41a5916c889110a2dee5521ed99c02effa3b2ac9

In 2019 an attacker read the credentials of an AWS instance role out of the metadata service at 169.254.169.254 and used them to copy data on about 100 million people from Capital One. The entry point was a server-side request forgery in a web application, not a container escape. The lesson for container operators is the same whichever layer the bug is in: a hardened image on a flat network is still one request away from a cloud role or a neighbouring database. The image hardening in this course decides what a process can do; the network decides what it can reach. This lesson is about reach. The first part runs on the main lab VM (secopslog-docker) in ~/lab/nethard, where the lesson files go; the host-firewall part runs on the sec VM at the end.

The default bridge connects everything

Put two containers on the default bridge and one reaches the other with nothing configured. One change to note while reading the address: Docker 29 removed the top-level NetworkSettings address fields, so the old docker inspect -f '{{.NetworkSettings.IPAddress}}' template errors, and you read the per-network address under .NetworkSettings.Networks instead.

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-a alpine:3.22 sleep 600 docker run -d --name lab-b alpine:3.22 sleep 600
c73a146b0f4fc2020a5adcc754b70d60c8e2e47711ddb4add551e30ff9bfa3fb ec7d4e1f6f34f3b88c469da11c8b55b1a2e56867b4adf8dceba95c3c7b9c26de
$ B=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' lab-b) echo "lab-b is $B" docker exec lab-a ping -c1 -W2 "$B" | grep -E 'bytes from|loss'
lab-b is 172.17.0.3 64 bytes from 172.17.0.3: seq=0 ttl=64 time=1.141 ms 1 packets transmitted, 1 packets received, 0% packet loss
$ docker network inspect bridge -f '{{index .Options "com.docker.network.bridge.enable_icc"}}'
true

lab-a pinged lab-b because inter-container communication (ICC) is on for the default bridge, which the last command confirms. "Networks, drivers and DNS" (Docker in depth) covers the driver model and why you create your own networks; the point here is that the default bridge is a single flat segment, and a process that lands on it can scan its neighbours. Two settings change that: turn ICC off, and keep unrelated workloads on separate networks so there is no shared segment to scan in the first place.

Turning inter-container traffic off

com.docker.network.bridge.enable_icc=false tells Docker to drop traffic between containers on that bridge while still allowing each container its gateway and the outside. Create a network with it, and a container on it can resolve a neighbour by name (embedded DNS is unaffected) but cannot connect to it:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker network create --subnet 10.89.40.0/24 -o com.docker.network.bridge.enable_icc=false lab-jobs docker run -d --name lab-j1 --network lab-jobs nginx:1.30-alpine docker run -d --name lab-j2 --network lab-jobs alpine:3.22 sleep 600
b6e876f45d16b147584e6bce530b704cb38cde232be744d81e8c1df503223b15 c953de54f5d2677cb9f469392a4d8acc0417523a68d32c78db8118a6981e3273 5749a84e25845f160efafe93208f751246a1f5ab8446b60cf886533734fc034e
$ docker exec lab-j2 nslookup lab-j1 | grep Address: | tail -1 docker exec lab-j2 wget -q -T 3 -O /dev/null http://lab-j1/ || echo "lab-j2 -> lab-j1: wget exit $?" docker exec lab-j2 wget -q -T 5 -O /dev/null http://example.com/ && echo 'lab-j2 -> example.com: ok'
Address: 10.89.40.2 wget: download timed out lab-j2 -> lab-j1: wget exit 1 lab-j2 -> example.com: ok

The name resolved to 10.89.40.2, the HTTP request to the neighbour timed out, and the request to the internet succeeded: ICC governs container-to-container traffic on the bridge, not egress. The rule Docker writes for it is visible in the per-bridge forward chain:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ BR=br-$(docker network inspect -f '{{.Id}}' lab-jobs | cut -c1-12) sudo iptables -S DOCKER-FORWARD | grep "$BR"
-A DOCKER-FORWARD -i br-b6e876f45d16 -o br-b6e876f45d16 -j DROP -A DOCKER-FORWARD -i br-b6e876f45d16 ! -o br-b6e876f45d16 -j ACCEPT

The first rule drops packets that both enter and leave this bridge (one container to another); the second allows packets that leave it for elsewhere. ICC off is a blunt instrument for one network. For a real two-tier application you reach for separate networks and an internal network, which is next.

An internal network, and the only way out

An internal network (--internal, or internal: true in Compose) gets no gateway route and no NAT, so its members reach each other and nothing else; "Networks, drivers and DNS" shows the routing. That is the right default for a data tier. But a tier that must call out, to one API or one package registry, still should not get the open internet. The pattern is a dual-homed egress proxy: the application sits on an internal network with no route out, the proxy sits on both that network and a second network that does have egress, and the application is configured to send everything through the proxy. The proxy is the one place egress is allowed, so it is the one place egress is logged and filtered.

compose.yaml
name: lab-nethard
services:
app:
image: nginx:1.30-alpine
networks: [apps]
environment:
http_proxy: http://egress:3128
https_proxy: http://egress:3128
no_proxy: localhost,127.0.0.1
egress:
image: python:3.14-slim
command: ["python3", "-u", "/proxy/egress-proxy.py"]
environment:
EGRESS_ALLOW: example.com
volumes:
- ./egress-proxy.py:/proxy/egress-proxy.py:ro
read_only: true
user: "65534:65534"
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
networks: [apps, outside]
networks:
apps:
internal: true
outside: {}

apps is internal; outside is not. Only egress joins both, and it runs a small allow-list forward proxy (standard library, no real credentials; use a maintained proxy such as Squid or an Envoy egress gateway in production). The application points http_proxy/https_proxy at it. Bring it up and confirm the shape:

ubuntu@secopslog-docker:~/lab/nethard · Docker 29.8.2
$ docker compose --progress quiet up -d docker compose ps --format "{{.Service}} {{.State}}"
app running egress running
$ docker network inspect lab-nethard_apps lab-nethard_outside -f '{{.Name}} internal={{.Internal}}' docker inspect lab-nethard-egress-1 -f '{{range $n, $_ := .NetworkSettings.Networks}}{{$n}} {{end}}'
lab-nethard_apps internal=true lab-nethard_outside internal=false lab-nethard_apps lab-nethard_outside

With the proxy not yet involved, the application cannot reach the internet at all, by name or by address, because its network has no route out:

ubuntu@secopslog-docker:~/lab/nethard · Docker 29.8.2
$ docker compose exec -T app curl -s --noproxy '*' -m 5 -o /dev/null http://example.com/ || echo "direct: curl exit $?" IP=$(getent ahostsv4 example.com | awk 'NR==1{print $1}') docker compose exec -T app curl -s --noproxy '*' -m 5 -o /dev/null http://$IP/ || echo "direct to $IP: curl exit $?"
direct: curl exit 6 direct to 172.66.147.243: curl exit 7

curl exits 6 (cannot resolve: the internal network's DNS has no upstream) and 7 (cannot connect) for the raw address. Now through the proxy, which allows only example.com:

ubuntu@secopslog-docker:~/lab/nethard · Docker 29.8.2
$ docker compose exec -T app curl -s -m 10 -o /dev/null -w 'example.com via proxy: %{http_code}\n' http://example.com/ docker compose exec -T app curl -s -m 10 -o /dev/null -w 'example.net via proxy: %{http_code}\n' http://example.net/ docker compose exec -T app curl -s -m 10 -p -o /dev/null -w 'example.com via CONNECT: %{http_code}\n' http://example.com/
example.com via proxy: 200 example.net via proxy: 403 example.com via CONNECT: 200

The allowed host returns 200 over plain HTTP and over a CONNECT tunnel (how a client proxies HTTPS); the host that is not on the list is refused with 403 before any connection leaves. The proxy's log is the egress record:

ubuntu@secopslog-docker:~/lab/nethard · Docker 29.8.2
$ docker compose logs --no-log-prefix egress
egress proxy on :3128, allow=['example.com'] ALLOW 172.19.0.3 GET example.com:80 DENY 172.19.0.3 GET example.net:80 ALLOW 172.19.0.3 CONNECT example.com:80

Every outbound attempt is one line with a verdict. A destination nobody approved shows up as a DENY the moment a workload tries it, which is the signal you want long before data leaves. This is the policy "Publishing ports and the packet path" (Docker in depth) points to: that lesson owns how published ports and DOCKER-USER work; this one decides what the rules should say.

Publish only what must be reached, and bind it narrowly

A workload on an internal network cannot be published at all, because there is no host route to NAT to it; Docker silently maps nothing:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-iw --network lab-nethard_apps -p 8099:80 nginx:1.30-alpine docker port lab-iw; echo "docker port printed $(docker port lab-iw | wc -l) mappings" curl -s -m 3 -o /dev/null http://127.0.0.1:8099/ || echo "127.0.0.1:8099: curl exit $?"
15dd6c28cae1515c96cdce8c60bc072eff7c7eefa3791fe6d8950339f5c76847 docker port printed 0 mappings 127.0.0.1:8099: curl exit 7

docker port prints zero mappings and the host cannot reach it. For a service that must be published, bind it to the address that actually needs it. -p 8091:80 binds every host interface; -p 127.0.0.1:8092:80 binds only loopback, so only the host itself connects:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run -d --name lab-pub -p 8091:80 nginx:1.30-alpine docker run -d --name lab-loc -p 127.0.0.1:8092:80 nginx:1.30-alpine IP=$(hostname -I | awk '{print $1}') for p in 8091 8092; do curl -s -m 3 -o /dev/null -w "$IP:$p %{http_code}\n" http://$IP:$p/ || echo "$IP:$p curl exit $?"; done
092bc49bf587aa5c7fca5efe07fa94f34a8857e7c6e67fc00b244622349bafdd 7a0bdf56c6a78b0dba49159ebd30792203b85048b8afd19a11b89b42cd059e53 192.168.2.4:8091 200 192.168.2.4:8092 000 192.168.2.4:8092 curl exit 7

The all-interfaces port answered from the host's LAN address; the loopback-bound one refused (curl exit 7). A database or an admin endpoint that only a local process uses belongs on 127.0.0.1. "Publishing ports and the packet path" shows setting this default for every container with the daemon's "ip": "127.0.0.1".

Verify, because this drifts

None of these settings shows up in application logs, and all of them drift: someone adds a container to the default bridge to save a minute, a debugging session leaves a port on 0.0.0.0 over a weekend. The checks are two commands. List what every container publishes, and list each network's isolation settings:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker ps --filter name=lab- --format '{{.Names}}\t{{.Ports}}' | sort
lab-a lab-b lab-iw 80/tcp lab-j1 80/tcp lab-j2 lab-loc 127.0.0.1:8092->80/tcp lab-nethard-app-1 80/tcp lab-nethard-egress-1 lab-pub 0.0.0.0:8091->80/tcp, [::]:8091->80/tcp
$ for n in $(docker network ls --filter driver=bridge --format '{{.Name}}' | grep -E '^(bridge|lab-)'); do docker network inspect "$n" -f '{{.Name}} internal={{.Internal}} icc={{or (index .Options "com.docker.network.bridge.enable_icc") "true"}} containers={{len .Containers}}' done
bridge internal=false icc=true containers=4 lab-jobs internal=false icc=false containers=2 lab-nethard_apps internal=true icc=true containers=3 lab-nethard_outside internal=false icc=true containers=1

The port sweep shows which containers expose a host port and on which address; lab-pub on 0.0.0.0 is the one to question. The network sweep shows internal, icc and the container count per network, so an internal tier that quietly gained a second, non-internal network stands out. Run both after any change window and keep the output with the ticket.

On a cloud host, add one destination to every egress policy by reflex: 169.254.169.254, the instance metadata service. A container rides the host's route to it, so an SSRF bug or a compromised process can read the node's role credentials over plain HTTP, exactly the Capital One path. Drop it for workloads that have no business minting cloud credentials. On AWS, also require IMDSv2 (HttpTokens=required) and set its PUT response hop limit to 1. The hop limit applies only to the IMDSv2 token request, so the extra hop a bridged container adds stops the token from reaching it; while IMDSv1 is still allowed, a plain GET from the container works whatever the hop limit, which is the Capital One path. Two gaps remain for the firewall rule shown next: it sits on the forwarding path, so it does not cover --network host containers or the host's own processes, and IPv6-enabled instances also serve the metadata service at fd00:ec2::254, which needs a rule of its own.

Clean up the main VM before moving to the sec VM:

ubuntu@secopslog-docker:~/lab/nethard · Docker 29.8.2
$ docker rm -f lab-a lab-b lab-j1 lab-j2 lab-iw lab-pub lab-loc >/dev/null docker compose --progress quiet down docker network rm lab-jobs
lab-jobs

Policy in the host firewall (sec VM)

Watch out
Run this only in the SecOpsLog disposable lab VM (secopslog-docker-sec): it changes the host firewall, adds a route and switches the daemon's firewall backend. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, then open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima). Reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo.

DOCKER-USER is the iptables chain Docker jumps to first and never fills, so rules you put there are evaluated before Docker's own NAT and filter rules. Unpack the lesson files into ~/lab/nethard on the sec VM too. The first thing they provide is a stand-in for a cloud metadata service: a network namespace on 169.254.169.254 serving an obviously fake document, which also acts as a LAN neighbour at 192.0.2.10. The script adds a host route, so it refuses to run anywhere but a lab VM; read it before you run it:

metadata-stub.sh
#!/bin/sh
# Stands in for a cloud metadata service in the lab: a network namespace "lab-meta" joined to the host by a
# veth pair, answering HTTP on 169.254.169.254 with a fake, harmless credential document. Host side:
# lab-meta0, 192.0.2.1/24. Namespace: 192.0.2.10/24 (it doubles as a LAN neighbour) and 169.254.169.254/32.
# It replaces the host's route to 169.254.169.254, so it refuses to run outside the SecOpsLog lab VM.
# Usage: sudo ./metadata-stub.sh up|down
[ -f /etc/secopslog-lab ] || { echo "refusing: run this only in the SecOpsLog lab VM" >&2; exit 2; }
set -eu
DOC=/run/lab-meta
case "${1:-}" in
up)
ip netns add lab-meta
ip link add lab-meta0 type veth peer name eth0 netns lab-meta
ip addr add 192.0.2.1/24 dev lab-meta0
ip link set lab-meta0 up
ip -n lab-meta link set lo up
ip -n lab-meta addr add 192.0.2.10/24 dev eth0
ip -n lab-meta addr add 169.254.169.254/32 dev eth0
ip -n lab-meta link set eth0 up
ip -n lab-meta route add default via 192.0.2.1
ip route add 169.254.169.254/32 via 192.0.2.10 dev lab-meta0
mkdir -p $DOC/latest/meta-data/iam/security-credentials
echo "lab-role: FAKE-LAB-CREDENTIAL (SecOpsLog metadata stub, not a real key)" \
> $DOC/latest/meta-data/iam/security-credentials/lab-role
setsid ip netns exec lab-meta python3 -m http.server 80 --bind 169.254.169.254 --directory $DOC \
>/dev/null 2>&1 < /dev/null &
sleep 1
;;
down)
ip netns pids lab-meta 2>/dev/null | xargs -r kill
ip route del 169.254.169.254/32 2>/dev/null || true
ip netns del lab-meta 2>/dev/null || true
ip link del lab-meta0 2>/dev/null || true
rm -rf $DOC
;;
*)
echo "usage: $0 up|down" >&2
exit 2
;;
esac
ubuntu@secopslog-docker-sec:~/lab/nethard · Docker 29.8.2
$ sudo ./metadata-stub.sh up ip route get 169.254.169.254 | head -1
169.254.169.254 via 192.0.2.10 dev lab-meta0 src 192.0.2.1 uid 1000

The host now routes 169.254.169.254 to the stand-in. Next, a published web service and two containers: lab-tool on the web service's bridge (named lab-br-web so the rules can refer to it) and lab-other on the default bridge:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ for i in nginx:1.30-alpine alpine:3.22; do sudo docker pull -q $i; done sudo docker network create --subnet 10.89.30.0/24 -o com.docker.network.bridge.name=lab-br-web lab-web-net sudo docker run -d --name lab-web --network lab-web-net -p 8080:80 nginx:1.30-alpine sudo docker run -d --name lab-tool --network lab-web-net alpine:3.22 sleep 3600 sudo docker run -d --name lab-other alpine:3.22 sleep 3600
docker.io/library/nginx:1.30-alpine docker.io/library/alpine:3.22 807205941c2a1826e562d3037bb90ebadbf72dec02740b92ed6a9f0e42aaa2bc d57051a58a55cc954a71f10a8045586910e5a53e60cf4b38bc8c6a44705b5adb 2dee2d4db0fb7aa19265dd3ea46f7e41fdd6107b4eb79e896b9a8778573c2c73 4ba65e04ada6775100242ac98488a73b3c42ce54cc96625b4eb00103651185fc

Before any policy, every container reaches the metadata address and the internet, and the neighbour reaches the published port:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -S DOCKER-USER for c in lab-tool lab-other; do sudo docker exec $c wget -q -T 3 -O- http://169.254.169.254/latest/meta-data/iam/security-credentials/lab-role 2>&1 | sed "s/^/$c metadata: /" { sudo docker exec $c wget -q -T 5 -O /dev/null http://example.com/ && echo ok; } 2>&1 | sed "s/^/$c example.com: /" done sudo ip netns exec lab-meta curl -s -m 3 -o /dev/null -w 'neighbour -> 192.0.2.1:8080: %{http_code}\n' http://192.0.2.1:8080/
-N DOCKER-USER lab-tool metadata: lab-role: FAKE-LAB-CREDENTIAL (SecOpsLog metadata stub, not a real key) lab-tool example.com: ok lab-other metadata: lab-role: FAKE-LAB-CREDENTIAL (SecOpsLog metadata stub, not a real key) lab-other example.com: ok neighbour -> 192.0.2.1:8080: 200

Now the policy: drop the metadata address for all containers, and default-deny new outbound connections from the web bridge while letting replies back. Since Docker 28.2.2 the chain is created empty with no built-in RETURN, so you can append rules and they run (the old advice to insert above a terminating RETURN no longer applies):

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -A DOCKER-USER -d 169.254.169.254/32 -j DROP sudo iptables -A DOCKER-USER -i lab-br-web -m conntrack --ctstate RELATED,ESTABLISHED -j RETURN sudo iptables -A DOCKER-USER -i lab-br-web ! -o lab-br-web -j DROP sudo iptables -S DOCKER-USER
-N DOCKER-USER -A DOCKER-USER -d 169.254.169.254/32 -j DROP -A DOCKER-USER -i lab-br-web -m conntrack --ctstate RELATED,ESTABLISHED -j RETURN -A DOCKER-USER -i lab-br-web ! -o lab-br-web -j DROP
$ for c in lab-tool lab-other; do sudo docker exec $c wget -q -T 3 -O- http://169.254.169.254/latest/meta-data/iam/security-credentials/lab-role 2>&1 | sed "s/^/$c metadata: /" { sudo docker exec $c wget -q -T 5 -O /dev/null http://example.com/ && echo ok; } 2>&1 | sed "s/^/$c example.com: /" done sudo docker exec lab-tool wget -q -T 3 -O /dev/null http://lab-web/ && echo 'lab-tool -> lab-web: ok' sudo ip netns exec lab-meta curl -s -m 3 -o /dev/null -w 'neighbour -> 192.0.2.1:8080: %{http_code}\n' http://192.0.2.1:8080/
lab-tool metadata: wget: download timed out lab-tool example.com: wget: download timed out lab-other metadata: wget: download timed out lab-other example.com: ok lab-tool -> lab-web: ok neighbour -> 192.0.2.1:8080: 200

The metadata read times out for both containers, egress from the web bridge is gone, but lab-tool still reaches lab-web on that bridge (established and intra-bridge traffic is allowed), and the neighbour still reaches the published port. The packet counters show each rule matching traffic:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -L DOCKER-USER -v -n
Chain DOCKER-USER (1 references) pkts bytes target prot opt in out source destination 6 360 DROP all -- * * 0.0.0.0/0 169.254.169.254 5 1402 RETURN all -- lab-br-web * 0.0.0.0/0 0.0.0.0/0 ctstate RELATED,ESTABLISHED 5 300 DROP all -- lab-br-web !lab-br-web 0.0.0.0/0 0.0.0.0/0

Rules added with iptables are lost on reboot but survive a daemon restart, because Docker recreates its own chains around yours without flushing DOCKER-USER:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo systemctl restart docker sudo iptables -S DOCKER-USER sudo iptables -S FORWARD | head -3
-N DOCKER-USER -A DOCKER-USER -d 169.254.169.254/32 -j DROP -A DOCKER-USER -i lab-br-web -m conntrack --ctstate RELATED,ESTABLISHED -j RETURN -A DOCKER-USER -i lab-br-web ! -o lab-br-web -j DROP -P FORWARD DROP -A FORWARD -j DOCKER-USER -A FORWARD -j DOCKER-FORWARD

Persist them the same way you persist any host firewall rule (your configuration management, or a unit ordered after docker.service). One caveat this course states wherever it matters: DOCKER-USER is an iptables-backend feature.

The nftables backend writes its own tables

Docker 29.0 added an experimental firewall backend that writes native nftables rules instead of iptables rules. Under it there is no DOCKER-USER chain; Docker owns its own nftables tables and you own yours. Switching it is a daemon.json change, so follow the procedure from "Configuring the daemon safely" (Docker in depth): back up the file (or record that there was none as {}), merge the new key with jq instead of overwriting whatever is there (the lab kit may have written MTU settings on VPN hosts), validate, then restart:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -F DOCKER-USER sudo cp -a /etc/docker/daemon.json /etc/docker/daemon.json.bak 2>/dev/null || echo '{}' | sudo tee /etc/docker/daemon.json.bak >/dev/null jq '. + {"firewall-backend": "nftables"}' /etc/docker/daemon.json.bak | sudo tee /etc/docker/daemon.json sudo dockerd --validate --config-file /etc/docker/daemon.json sudo systemctl restart docker sudo docker info --format '{{.FirewallBackend.Driver}}' sudo docker start lab-web lab-tool lab-other >/dev/null
{ "firewall-backend": "nftables" } configuration OK nftables

One leftover needs a decision. While the iptables backend ran, Docker set the iptables FORWARD policy to DROP, and nothing resets it after the switch, so it would still drop container traffic that the nftables rules accept:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo iptables -S FORWARD | head -1 sudo sysctl net.ipv4.ip_forward sudo iptables -P FORWARD ACCEPT sudo iptables -S FORWARD | head -1
-P FORWARD DROP net.ipv4.ip_forward = 1 -P FORWARD ACCEPT

The lab sets the policy back to ACCEPT. Understand what that does on a real host: with ip_forward=1, that DROP policy was the host's only default-deny for forwarded packets, and without it the host routes between any of its interfaces (a VPN, a second NIC, other bridges). Docker's nftables documentation says the same: when forwarding is on, you need your own rules that block unwanted forwarding between non-Docker interfaces. On a real host, put that drop in an nftables forward chain of your own before you open the policy, and trial the experimental backend on a disposable host first. With the policy open, the metadata address is reachable again, because the old DOCKER-USER rules went with the iptables backend:

ubuntu@secopslog-docker-sec:~ · Docker 29.8.2
$ sudo docker exec lab-tool wget -q -T 3 -O- http://169.254.169.254/latest/meta-data/iam/security-credentials/lab-role
lab-role: FAKE-LAB-CREDENTIAL (SecOpsLog metadata stub, not a real key)

Put the same policy in a table of your own. A drop in any nftables base chain is final regardless of priority, so a drop in your table wins over Docker's accept; the priority only decides ordering:

lab-policy.nft
# The same policy for the nftables firewall backend. Docker never touches a table it did not create.
table inet lab-policy {
chain forward {
type filter hook forward priority filter - 10; policy accept;
ip daddr 169.254.169.254 counter drop
iifname "lab-br-web" oifname != "lab-br-web" ct state new counter drop
}
}
ubuntu@secopslog-docker-sec:~/lab/nethard · Docker 29.8.2
$ sudo nft -f lab-policy.nft for c in lab-tool lab-other; do sudo docker exec $c wget -q -T 3 -O- http://169.254.169.254/latest/meta-data/iam/security-credentials/lab-role 2>&1 | sed "s/^/$c metadata: /" done { sudo docker exec lab-tool wget -q -T 5 -O /dev/null http://example.com/ && echo ok; } 2>&1 | sed "s/^/lab-tool example.com: /" { sudo docker exec lab-other wget -q -T 5 -O /dev/null http://example.com/ && echo ok; } 2>&1 | sed "s/^/lab-other example.com: /" sudo ip netns exec lab-meta curl -s -m 3 -o /dev/null -w "neighbour -> 192.0.2.1:8080: %{http_code}\n" http://192.0.2.1:8080/ sudo nft list table inet lab-policy
lab-tool metadata: wget: download timed out lab-other metadata: wget: download timed out lab-tool example.com: wget: download timed out lab-other example.com: ok neighbour -> 192.0.2.1:8080: 200 table inet lab-policy { chain forward { type filter hook forward priority filter - 10; policy accept; ip daddr 169.254.169.254 counter packets 6 bytes 360 drop iifname "lab-br-web" oifname != "lab-br-web" ct state new counter packets 6 bytes 360 drop } }

The metadata address is dropped for every container and new egress from the web bridge is dropped, while the neighbour still reaches the published port and lab-other still reaches the internet, all from a table Docker never touches. The counters on the two rules show they matched. Roll the backend back to iptables by restoring the backup, and clean up:

ubuntu@secopslog-docker-sec:~/lab/nethard · Docker 29.8.2
$ sudo nft delete table inet lab-policy sudo docker rm -f lab-web lab-tool lab-other >/dev/null sudo docker network rm lab-web-net sudo cp /etc/docker/daemon.json.bak /etc/docker/daemon.json sudo systemctl reset-failed docker sudo systemctl restart docker sudo docker info --format '{{.FirewallBackend.Driver}}' sudo iptables -S FORWARD | head -1 sudo iptables -P FORWARD DROP sudo iptables -S FORWARD | head -1 sudo ./metadata-stub.sh down
lab-web-net iptables -P FORWARD ACCEPT -P FORWARD DROP

The daemon is back on iptables, but the FORWARD policy is still ACCEPT: Docker sets DROP only when it turns IP forwarding on itself, and forwarding was already on. Put the policy back by hand, as the last iptables lines do; on a real host this is the step people forget. The stand-in metadata service and its route are gone.

Quick check
01A data-tier container runs on a network created with --internal. An attacker gets code execution in an application container that is attached to that same internal network. What can they still reach?
Incorrect — Internal only removes the route to the outside. Members of the network still reach each other unless ICC is off or they are on separate networks.
Incorrect — An internal network has no upstream resolver and no default route; names do not resolve and there is no path out.
Correct — --internal removes the gateway route and NAT only. Blocking lateral traffic needs ICC off or separate networks.
Incorrect — The internal flag does no such redirection; it only withholds the outside route.
02You want a container that must call one external API but nothing else. Which design gives you both the restriction and a record of what it tried to reach?
Incorrect — Publishing is about inbound reach and does not constrain or log the container's outbound destinations at all.
Correct — The internal network removes any direct route out, and the dual-homed proxy is the single point that both allows and logs each destination.
Incorrect — Host networking removes the container's isolation entirely and grants it every destination the host has, the opposite of restriction.
Incorrect — ICC off stops container-to-container traffic on that bridge; it neither restricts egress to the internet nor records it.
03On a Docker 29 host using the default iptables firewall backend, a teammate says your default-deny egress rule must be inserted above DOCKER-USER's built-in RETURN or it will never run. Is that correct?
Incorrect — That was true on older Docker; since 28.2.2 the chain has no built-in RETURN.
Incorrect — The target is irrelevant; the premise about a built-in RETURN is what is wrong.
Incorrect — Egress filtering is exactly what DOCKER-USER is for; it sees forwarded traffic before Docker's own rules.
Correct — The chain ships with no RETURN now, so you can append or insert; if you switch to the nftables backend there is no DOCKER-USER and you write your own table.

Try this

Work through “The nftables backend writes its own tables” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from container network hardening, keep “The nftables backend writes its own tables”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related