Container network hardening
Internal networks, an egress proxy, DOCKER-USER policy and the nftables backend.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/nethard.tar.gz && tar -xzf nethard.tar.gz, which creates ~/lab/nethard/. SHA-256: 92e9f24f7e04884464ef0e2e41a5916c889110a2dee5521ed99c02effa3b2ac9In 2019 an attacker read the credentials of an AWS instance role out of the metadata service at 169.254.169.254 and used them to copy data on about 100 million people from Capital One. The entry point was a server-side request forgery in a web application, not a container escape. The lesson for container operators is the same whichever layer the bug is in: a hardened image on a flat network is still one request away from a cloud role or a neighbouring database. The image hardening in this course decides what a process can do; the network decides what it can reach. This lesson is about reach. The first part runs on the main lab VM (secopslog-docker) in ~/lab/nethard, where the lesson files go; the host-firewall part runs on the sec VM at the end.
The default bridge connects everything
Put two containers on the default bridge and one reaches the other with nothing configured. One change to note while reading the address: Docker 29 removed the top-level NetworkSettings address fields, so the old docker inspect -f '{{.NetworkSettings.IPAddress}}' template errors, and you read the per-network address under .NetworkSettings.Networks instead.
lab-a pinged lab-b because inter-container communication (ICC) is on for the default bridge, which the last command confirms. "Networks, drivers and DNS" (Docker in depth) covers the driver model and why you create your own networks; the point here is that the default bridge is a single flat segment, and a process that lands on it can scan its neighbours. Two settings change that: turn ICC off, and keep unrelated workloads on separate networks so there is no shared segment to scan in the first place.
Turning inter-container traffic off
com.docker.network.bridge.enable_icc=false tells Docker to drop traffic between containers on that bridge while still allowing each container its gateway and the outside. Create a network with it, and a container on it can resolve a neighbour by name (embedded DNS is unaffected) but cannot connect to it:
The name resolved to 10.89.40.2, the HTTP request to the neighbour timed out, and the request to the internet succeeded: ICC governs container-to-container traffic on the bridge, not egress. The rule Docker writes for it is visible in the per-bridge forward chain:
The first rule drops packets that both enter and leave this bridge (one container to another); the second allows packets that leave it for elsewhere. ICC off is a blunt instrument for one network. For a real two-tier application you reach for separate networks and an internal network, which is next.
An internal network, and the only way out
An internal network (--internal, or internal: true in Compose) gets no gateway route and no NAT, so its members reach each other and nothing else; "Networks, drivers and DNS" shows the routing. That is the right default for a data tier. But a tier that must call out, to one API or one package registry, still should not get the open internet. The pattern is a dual-homed egress proxy: the application sits on an internal network with no route out, the proxy sits on both that network and a second network that does have egress, and the application is configured to send everything through the proxy. The proxy is the one place egress is allowed, so it is the one place egress is logged and filtered.
name: lab-nethardservices:app:image: nginx:1.30-alpinenetworks: [apps]environment:http_proxy: http://egress:3128https_proxy: http://egress:3128no_proxy: localhost,127.0.0.1egress:image: python:3.14-slimcommand: ["python3", "-u", "/proxy/egress-proxy.py"]environment:EGRESS_ALLOW: example.comvolumes:- ./egress-proxy.py:/proxy/egress-proxy.py:roread_only: trueuser: "65534:65534"cap_drop: [ALL]security_opt: [no-new-privileges:true]networks: [apps, outside]networks:apps:internal: trueoutside: {}
apps is internal; outside is not. Only egress joins both, and it runs a small allow-list forward proxy (standard library, no real credentials; use a maintained proxy such as Squid or an Envoy egress gateway in production). The application points http_proxy/https_proxy at it. Bring it up and confirm the shape:
With the proxy not yet involved, the application cannot reach the internet at all, by name or by address, because its network has no route out:
curl exits 6 (cannot resolve: the internal network's DNS has no upstream) and 7 (cannot connect) for the raw address. Now through the proxy, which allows only example.com:
The allowed host returns 200 over plain HTTP and over a CONNECT tunnel (how a client proxies HTTPS); the host that is not on the list is refused with 403 before any connection leaves. The proxy's log is the egress record:
Every outbound attempt is one line with a verdict. A destination nobody approved shows up as a DENY the moment a workload tries it, which is the signal you want long before data leaves. This is the policy "Publishing ports and the packet path" (Docker in depth) points to: that lesson owns how published ports and DOCKER-USER work; this one decides what the rules should say.
Publish only what must be reached, and bind it narrowly
A workload on an internal network cannot be published at all, because there is no host route to NAT to it; Docker silently maps nothing:
docker port prints zero mappings and the host cannot reach it. For a service that must be published, bind it to the address that actually needs it. -p 8091:80 binds every host interface; -p 127.0.0.1:8092:80 binds only loopback, so only the host itself connects:
The all-interfaces port answered from the host's LAN address; the loopback-bound one refused (curl exit 7). A database or an admin endpoint that only a local process uses belongs on 127.0.0.1. "Publishing ports and the packet path" shows setting this default for every container with the daemon's "ip": "127.0.0.1".
Verify, because this drifts
None of these settings shows up in application logs, and all of them drift: someone adds a container to the default bridge to save a minute, a debugging session leaves a port on 0.0.0.0 over a weekend. The checks are two commands. List what every container publishes, and list each network's isolation settings:
The port sweep shows which containers expose a host port and on which address; lab-pub on 0.0.0.0 is the one to question. The network sweep shows internal, icc and the container count per network, so an internal tier that quietly gained a second, non-internal network stands out. Run both after any change window and keep the output with the ticket.
On a cloud host, add one destination to every egress policy by reflex: 169.254.169.254, the instance metadata service. A container rides the host's route to it, so an SSRF bug or a compromised process can read the node's role credentials over plain HTTP, exactly the Capital One path. Drop it for workloads that have no business minting cloud credentials. On AWS, also require IMDSv2 (HttpTokens=required) and set its PUT response hop limit to 1. The hop limit applies only to the IMDSv2 token request, so the extra hop a bridged container adds stops the token from reaching it; while IMDSv1 is still allowed, a plain GET from the container works whatever the hop limit, which is the Capital One path. Two gaps remain for the firewall rule shown next: it sits on the forwarding path, so it does not cover --network host containers or the host's own processes, and IPv6-enabled instances also serve the metadata service at fd00:ec2::254, which needs a rule of its own.
Clean up the main VM before moving to the sec VM:
Policy in the host firewall (sec VM)
secopslog-docker-sec): it changes the host firewall, adds a route and switches the daemon's firewall backend. If the VM does not exist, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, then open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima). Reset it at any point with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu account is not in the docker group, so docker runs through sudo.DOCKER-USER is the iptables chain Docker jumps to first and never fills, so rules you put there are evaluated before Docker's own NAT and filter rules. Unpack the lesson files into ~/lab/nethard on the sec VM too. The first thing they provide is a stand-in for a cloud metadata service: a network namespace on 169.254.169.254 serving an obviously fake document, which also acts as a LAN neighbour at 192.0.2.10. The script adds a host route, so it refuses to run anywhere but a lab VM; read it before you run it:
#!/bin/sh# Stands in for a cloud metadata service in the lab: a network namespace "lab-meta" joined to the host by a# veth pair, answering HTTP on 169.254.169.254 with a fake, harmless credential document. Host side:# lab-meta0, 192.0.2.1/24. Namespace: 192.0.2.10/24 (it doubles as a LAN neighbour) and 169.254.169.254/32.# It replaces the host's route to 169.254.169.254, so it refuses to run outside the SecOpsLog lab VM.# Usage: sudo ./metadata-stub.sh up|down[ -f /etc/secopslog-lab ] || { echo "refusing: run this only in the SecOpsLog lab VM" >&2; exit 2; }set -euDOC=/run/lab-metacase "${1:-}" inup)ip netns add lab-metaip link add lab-meta0 type veth peer name eth0 netns lab-metaip addr add 192.0.2.1/24 dev lab-meta0ip link set lab-meta0 upip -n lab-meta link set lo upip -n lab-meta addr add 192.0.2.10/24 dev eth0ip -n lab-meta addr add 169.254.169.254/32 dev eth0ip -n lab-meta link set eth0 upip -n lab-meta route add default via 192.0.2.1ip route add 169.254.169.254/32 via 192.0.2.10 dev lab-meta0mkdir -p $DOC/latest/meta-data/iam/security-credentialsecho "lab-role: FAKE-LAB-CREDENTIAL (SecOpsLog metadata stub, not a real key)" \> $DOC/latest/meta-data/iam/security-credentials/lab-rolesetsid ip netns exec lab-meta python3 -m http.server 80 --bind 169.254.169.254 --directory $DOC \>/dev/null 2>&1 < /dev/null &sleep 1;;down)ip netns pids lab-meta 2>/dev/null | xargs -r killip route del 169.254.169.254/32 2>/dev/null || trueip netns del lab-meta 2>/dev/null || trueip link del lab-meta0 2>/dev/null || truerm -rf $DOC;;*)echo "usage: $0 up|down" >&2exit 2;;esac
The host now routes 169.254.169.254 to the stand-in. Next, a published web service and two containers: lab-tool on the web service's bridge (named lab-br-web so the rules can refer to it) and lab-other on the default bridge:
Before any policy, every container reaches the metadata address and the internet, and the neighbour reaches the published port:
Now the policy: drop the metadata address for all containers, and default-deny new outbound connections from the web bridge while letting replies back. Since Docker 28.2.2 the chain is created empty with no built-in RETURN, so you can append rules and they run (the old advice to insert above a terminating RETURN no longer applies):
The metadata read times out for both containers, egress from the web bridge is gone, but lab-tool still reaches lab-web on that bridge (established and intra-bridge traffic is allowed), and the neighbour still reaches the published port. The packet counters show each rule matching traffic:
Rules added with iptables are lost on reboot but survive a daemon restart, because Docker recreates its own chains around yours without flushing DOCKER-USER:
Persist them the same way you persist any host firewall rule (your configuration management, or a unit ordered after docker.service). One caveat this course states wherever it matters: DOCKER-USER is an iptables-backend feature.
The nftables backend writes its own tables
Docker 29.0 added an experimental firewall backend that writes native nftables rules instead of iptables rules. Under it there is no DOCKER-USER chain; Docker owns its own nftables tables and you own yours. Switching it is a daemon.json change, so follow the procedure from "Configuring the daemon safely" (Docker in depth): back up the file (or record that there was none as {}), merge the new key with jq instead of overwriting whatever is there (the lab kit may have written MTU settings on VPN hosts), validate, then restart:
One leftover needs a decision. While the iptables backend ran, Docker set the iptables FORWARD policy to DROP, and nothing resets it after the switch, so it would still drop container traffic that the nftables rules accept:
The lab sets the policy back to ACCEPT. Understand what that does on a real host: with ip_forward=1, that DROP policy was the host's only default-deny for forwarded packets, and without it the host routes between any of its interfaces (a VPN, a second NIC, other bridges). Docker's nftables documentation says the same: when forwarding is on, you need your own rules that block unwanted forwarding between non-Docker interfaces. On a real host, put that drop in an nftables forward chain of your own before you open the policy, and trial the experimental backend on a disposable host first. With the policy open, the metadata address is reachable again, because the old DOCKER-USER rules went with the iptables backend:
Put the same policy in a table of your own. A drop in any nftables base chain is final regardless of priority, so a drop in your table wins over Docker's accept; the priority only decides ordering:
# The same policy for the nftables firewall backend. Docker never touches a table it did not create.table inet lab-policy {chain forward {type filter hook forward priority filter - 10; policy accept;ip daddr 169.254.169.254 counter dropiifname "lab-br-web" oifname != "lab-br-web" ct state new counter drop}}
The metadata address is dropped for every container and new egress from the web bridge is dropped, while the neighbour still reaches the published port and lab-other still reaches the internet, all from a table Docker never touches. The counters on the two rules show they matched. Roll the backend back to iptables by restoring the backup, and clean up:
The daemon is back on iptables, but the FORWARD policy is still ACCEPT: Docker sets DROP only when it turns IP forwarding on itself, and forwarding was already on. Put the policy back by hand, as the last iptables lines do; on a real host this is the step people forget. The stand-in metadata service and its route are gone.
--internal. An attacker gets code execution in an application container that is attached to that same internal network. What can they still reach?--internal removes the gateway route and NAT only. Blocking lateral traffic needs ICC off or separate networks.DOCKER-USER's built-in RETURN or it will never run. Is that correct?RETURN.RETURN is what is wrong.DOCKER-USER is for; it sees forwarded traffic before Docker's own rules.RETURN now, so you can append or insert; if you switch to the nftables backend there is no DOCKER-USER and you write your own table.Try this
Work through “The nftables backend writes its own tables” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from container network hardening, keep “The nftables backend writes its own tables”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.