Network troubleshooting
A method for 'it cannot connect', with nsenter, tcpdump, rule counters and conntrack.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/nettrouble.tar.gz && tar -xzf nettrouble.tar.gz, which creates ~/lab/nettrouble/. SHA-256: 61831bb57cce19744d90e66f0ebd3e2a144babbc865f56ed641bcf017398f545Use the main lab VM (secopslog-docker). The two files below are in this lesson's download; they describe a small app with nginx in front of a Python API, and three faults are built in. Start it and try it:
# A deliberately broken two-tier app: three faults, found and fixed in the lesson.name: lab-troubleservices:web:image: nginx:1.30-alpineports:- "8090:8090"volumes:- ./default.conf:/etc/nginx/conf.d/default.conf:ronetworks: [front]api:image: python:3.14-slimcommand: ["python3", "-m", "http.server", "8000", "--bind", "127.0.0.1"]networks: [back]networks:front:ipam:config:- subnet: 10.89.20.0/24back:ipam:config:- subnet: 10.89.21.0/24
server {listen 80;# Resolve "api" per request through Docker's embedded DNS server, so nginx# starts (and logs a clear error) even while the name does not resolve.resolver 127.0.0.11 valid=10s;set $api_upstream http://api:8000;location / {proxy_pass $api_upstream;}}
Both containers are up, and the same port fails in two different ways. Those two messages already carry information. The connection to 127.0.0.1 was accepted and then reset: that is docker-proxy, which accepted on the host side and closed the connection when its own connection to the container failed ("Publishing ports and the packet path" showed that loopback clients go through docker-proxy). The connection to the host's address was refused outright: DNAT forwarded the SYN to the container, and the container's kernel answered with a reset because nothing listens on that port. Learn to read the failure before reaching for tools:
Then work through the same four questions every time, in order, and stop at the first one that fails. Fixing that layer often reveals the next fault, which is exactly what happens here:
1. What listens, and where
The mapping sends host port 8090 to container port 8090, while nginx listens on port 80 (0.0.0.0:80). The second LISTEN line is dockerd's embedded DNS listener, present in every container on a user-defined network. ss is not in the nginx image, and many images have no tools at all, so the command runs the host's ss inside the container's network namespace with nsenter, as in "Networks, drivers and DNS". Fix the mapping (with sed here, or in an editor) and apply it; Compose recreates only the changed service:
A 502 is progress: the request now reaches nginx, and nginx cannot reach its upstream.
2. Which network, and does the name resolve
nginx resolves api per request through 127.0.0.11 (that is what the resolver line in default.conf does; without it nginx would refuse to start while the name is missing, which is a harder failure to read). The lookup fails from the web container too: the server is 127.0.0.11, so the container is on a user-defined network and DNS works, it just does not know api. docker inspect shows why: web is only on front, api only on back, and names resolve only between containers that share a network. A Server line other than 127.0.0.11 would have told you something different: the client is on the default bridge or uses --network host. Add back to web's networks:
Still 502, now with a different error in nginx's log, connect() failed (111: Connection refused): the name resolved to the api's address and the TCP connection was refused. The sleep 2 gives the recreated web container a moment to start; a request sent while it is still starting gets no HTTP answer at all, and curl prints 000. One side effect of the fix is worth knowing:
With two networks, web has one default route, and Docker chose the back network's gateway. Published ports follow the default gateway, so the DNAT rule now targets web's back address. Docker picks the gateway and may change it when connections change; when it matters (for example, a service that must publish only on its front network), set gw_priority on the network in Compose or --network name=...,gw-priority=1 on docker run.
3. Down to the packets
A refused connection already says a host answered. Capturing on the back bridge while sending one request shows who answered and how:
The first packet is web's SYN to 10.89.21.2:8000; the reply comes straight back from the api with flags R., a reset. The packet reached the right container, so the network, the bridge and the rules are fine. The api's kernel refused because no socket listens on that address and port. Timestamps, ports and sequence numbers differ on every run. If the capture had shown SYNs with no answer, the next stop would be rule counters; if it had shown nothing, the client was sending somewhere else. Capture on the bridge to see a whole network, or on one container's host-side veth (find it as in "Networks, drivers and DNS") to see one container. Now look at the api's sockets, in two ways:
python3 listens on 127.0.0.1:8000, a loopback address that exists only inside the api's own namespace, so connections arriving on its network interface are refused. The second command gets the same answer without root on the host: a throwaway container started with --network container:lab-trouble-api-1 shares the api's network namespace and brings its own tools, here BusyBox netstat from alpine. Any image with the tools you need works as that sidecar. "Container networking basics" (Docker for beginners) showed this loopback mistake from the outside; the fix is always in the application's configuration:
4. The host's rules
When packets go out and nothing answers, look at what the host's firewall did with them. Every iptables rule counts the packets it matched:
The DNAT rule counted 3 packets for 3 requests to the host's address: one per connection, because only the first packet of a connection goes through the nat table and conntrack handles the rest. The request to 127.0.0.1 did not count at all, because it went through docker-proxy, not DNAT. That is how counters answer "did my traffic even get here": run the request, list again, and see which rule moved. A DROP rule whose counter rises is your answer; a DNAT rule that does not move means the packets never reached this host on that port, or reached it on a different address or protocol. curl localhost tries ::1 before 127.0.0.1, for example, which takes a different path. On the nftables backend, sudo nft list ruleset shows the same kind of counters on Docker's rules.
conntrack shows each translated connection: the original direction to 192.168.2.4:8090 and the reply expected from 10.89.21.3:80. Entries in SYN_SENT that never progress are the conntrack view of a timeout. DOCKER-USER is empty on this host, so no local rule of yours is filtering; on a host where someone added rules there, it is the first place to look when a published port times out from outside but works from the host itself, as the DOCKER-USER part of "Publishing ports and the packet path" demonstrated.
Two causes this lab does not reproduce. MTU problems show up as small requests working while large responses hang, typically on overlay networks or hosts behind a VPN: compare ip link MTUs inside the container and on the host path, and set com.docker.network.driver.mtu on the network to match the smallest one. And when the client is another machine, add the layers between the two hosts (cloud security groups, network ACLs) to question 4 before blaming Docker. Clean up:
curl http://app-host:8080 times out. On app-host, curl http://127.0.0.1:8080 works. The DNAT rule's counter for 8080 rises with every remote attempt. What does that tell you?10.89.21.2:8000 followed immediately by a packet with flags R. from 10.89.21.2. Which conclusion is justified?Try this
Work through “4. The host's rules” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from network troubleshooting, keep “4. The host's rules”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.