Network troubleshooting

A method for 'it cannot connect', with nsenter, tcpdump, rule counters and conntrack.

Intermediate16 min · lesson 14 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (2 files, 1 KB): nettrouble.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-hard/nettrouble.tar.gz && tar -xzf nettrouble.tar.gz, which creates ~/lab/nettrouble/. SHA-256: 61831bb57cce19744d90e66f0ebd3e2a144babbc865f56ed641bcf017398f545

Use the main lab VM (secopslog-docker). The two files below are in this lesson's download; they describe a small app with nginx in front of a Python API, and three faults are built in. Start it and try it:

compose.yaml
# A deliberately broken two-tier app: three faults, found and fixed in the lesson.
name: lab-trouble
services:
web:
image: nginx:1.30-alpine
ports:
- "8090:8090"
volumes:
- ./default.conf:/etc/nginx/conf.d/default.conf:ro
networks: [front]
api:
image: python:3.14-slim
command: ["python3", "-m", "http.server", "8000", "--bind", "127.0.0.1"]
networks: [back]
networks:
front:
ipam:
config:
- subnet: 10.89.20.0/24
back:
ipam:
config:
- subnet: 10.89.21.0/24
default.conf
server {
listen 80;
# Resolve "api" per request through Docker's embedded DNS server, so nginx
# starts (and logs a clear error) even while the name does not resolve.
resolver 127.0.0.11 valid=10s;
set $api_upstream http://api:8000;
location / {
proxy_pass $api_upstream;
}
}
ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ docker compose up -d docker compose ps --format "{{.Name}}: {{.Status}}"
Network lab-trouble_back Creating Network lab-trouble_front Creating Network lab-trouble_back Creating Network lab-trouble_front Creating Network lab-trouble_front Created Network lab-trouble_front Created Container lab-trouble-web-1 Creating Network lab-trouble_back Created Network lab-trouble_back Created Container lab-trouble-api-1 Creating Container lab-trouble-web-1 Created Container lab-trouble-api-1 Created Container lab-trouble-api-1 Starting Container lab-trouble-web-1 Starting Container lab-trouble-api-1 Started Container lab-trouble-web-1 Started lab-trouble-api-1: Up Less than a second lab-trouble-web-1: Up Less than a second
$ curl -sS http://127.0.0.1:8090/
curl: (56) Recv failure: Connection reset by peer
$ curl -sS http://$(hostname -I | awk "{print \$1}"):8090/
curl: (7) Failed to connect to 192.168.2.4 port 8090 after 0 ms: Could not connect to server

Both containers are up, and the same port fails in two different ways. Those two messages already carry information. The connection to 127.0.0.1 was accepted and then reset: that is docker-proxy, which accepted on the host side and closed the connection when its own connection to the container failed ("Publishing ports and the packet path" showed that loopback clients go through docker-proxy). The connection to the host's address was refused outright: DNAT forwarded the SYN to the container, and the container's kernel answered with a reset because nothing listens on that port. Learn to read the failure before reaching for tools:

Then work through the same four questions every time, in order, and stop at the first one that fails. Fixing that layer often reveals the next fault, which is exactly what happens here:

Four questions for "it cannot connect"
11. What listens, on which address and port?
docker port; nsenter -n ... ss -tlnp; a sidecar on --network container:NAME
22. Which network, and does the name resolve?
docker inspect .NetworkSettings.Networks; nslookup from the client's side; the client's logs
33. What do the packets do?
tcpdump on the bridge or veth: SYN and SYN-ACK, SYN and RST, or SYN and silence
44. What do the host's rules do?
iptables -L -v -n counters (nft list ruleset on the nftables backend), conntrack, DOCKER-USER
Run each check from where the failing client sits: the host, another container, or another machine.

1. What listens, and where

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ docker compose port web 8090
0.0.0.0:8090
$ sudo nsenter -n -t $(docker inspect -f '{{.State.Pid}}' lab-trouble-web-1) ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("nginx",pid=602107,fd=6),("nginx",pid=602106,fd=6),("nginx",pid=602105,fd=6),("nginx",pid=602104,fd=6),("nginx",pid=602018,fd=6)) LISTEN 0 4096 127.0.0.11:41203 0.0.0.0:* users:(("dockerd",pid=17027,fd=44))

The mapping sends host port 8090 to container port 8090, while nginx listens on port 80 (0.0.0.0:80). The second LISTEN line is dockerd's embedded DNS listener, present in every container on a user-defined network. ss is not in the nginx image, and many images have no tools at all, so the command runs the host's ss inside the container's network namespace with nsenter, as in "Networks, drivers and DNS". Fix the mapping (with sed here, or in an editor) and apply it; Compose recreates only the changed service:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ sed -i 's/"8090:8090"/"8090:80"/' compose.yaml docker compose up -d curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8090/
Container lab-trouble-api-1 Running Container lab-trouble-web-1 Recreate Container lab-trouble-web-1 Recreated Container lab-trouble-web-1 Starting Container lab-trouble-web-1 Started 502

A 502 is progress: the request now reaches nginx, and nginx cannot reach its upstream.

2. Which network, and does the name resolve

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ docker compose logs web 2>&1 | grep -o 'api could not be resolved.*' | sed 's/, client.*//' | tail -n 1
api could not be resolved (2: Server failure)
$ docker compose exec -T web nslookup api
Server: 127.0.0.11 Address: 127.0.0.11:53 ** server can't find api: SERVFAIL ** server can't find api: SERVFAIL
$ docker inspect -f '{{.Name}}: {{range $k, $v := .NetworkSettings.Networks}}{{$k}}={{$v.IPAddress}} {{end}}' lab-trouble-web-1 lab-trouble-api-1
/lab-trouble-web-1: lab-trouble_front=10.89.20.2 /lab-trouble-api-1: lab-trouble_back=10.89.21.2

nginx resolves api per request through 127.0.0.11 (that is what the resolver line in default.conf does; without it nginx would refuse to start while the name is missing, which is a harder failure to read). The lookup fails from the web container too: the server is 127.0.0.11, so the container is on a user-defined network and DNS works, it just does not know api. docker inspect shows why: web is only on front, api only on back, and names resolve only between containers that share a network. A Server line other than 127.0.0.11 would have told you something different: the client is on the default bridge or uses --network host. Add back to web's networks:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ sed -i 's/networks: \[front\]/networks: [front, back]/' compose.yaml docker compose up -d sleep 2 curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8090/ docker compose logs web 2>&1 | grep -o 'connect() failed.*upstream' | tail -n 1
Container lab-trouble-api-1 Running Container lab-trouble-web-1 Recreate Container lab-trouble-web-1 Recreated Container lab-trouble-web-1 Starting Container lab-trouble-web-1 Started 502 connect() failed (111: Connection refused) while connecting to upstream, client: 10.89.21.1, server: , request: "GET / HTTP/1.1", upstream

Still 502, now with a different error in nginx's log, connect() failed (111: Connection refused): the name resolved to the api's address and the TCP connection was refused. The sleep 2 gives the recreated web container a moment to start; a request sent while it is still starting gets no HTTP answer at all, and curl prints 000. One side effect of the fix is worth knowing:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ docker compose exec -T web ip route sudo iptables -t nat -S DOCKER | grep -e '--dport 8090'
default via 10.89.21.1 dev eth0 10.89.20.0/24 dev eth1 scope link src 10.89.20.2 10.89.21.0/24 dev eth0 scope link src 10.89.21.3 -A DOCKER ! -i br-716c10204831 -p tcp -m tcp --dport 8090 -j DNAT --to-destination 10.89.21.3:80

With two networks, web has one default route, and Docker chose the back network's gateway. Published ports follow the default gateway, so the DNAT rule now targets web's back address. Docker picks the gateway and may change it when connections change; when it matters (for example, a service that must publish only on its front network), set gw_priority on the network in Compose or --network name=...,gw-priority=1 on docker run.

3. Down to the packets

A refused connection already says a host answered. Capturing on the back bridge while sending one request shows who answered and how:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ BR=br-$(docker network inspect -f '{{.Id}}' lab-trouble_back | cut -c1-12) sudo timeout 10 tcpdump -lnni $BR -c 2 'tcp port 8000' 2>/dev/null & sleep 2 curl -s -o /dev/null http://127.0.0.1:8090/ wait
02:49:17.861541 IP 10.89.21.3.53848 > 10.89.21.2.8000: Flags [S], seq 1029520932, win 64240, options [mss 1460,sackOK,TS val 507770958 ecr 0,nop,wscale 10], length 0 02:49:17.861631 IP 10.89.21.2.8000 > 10.89.21.3.53848: Flags [R.], seq 0, ack 1029520933, win 0, length 0

The first packet is web's SYN to 10.89.21.2:8000; the reply comes straight back from the api with flags R., a reset. The packet reached the right container, so the network, the bridge and the rules are fine. The api's kernel refused because no socket listens on that address and port. Timestamps, ports and sequence numbers differ on every run. If the capture had shown SYNs with no answer, the next stop would be rule counters; if it had shown nothing, the client was sending somewhere else. Capture on the bridge to see a whole network, or on one container's host-side veth (find it as in "Networks, drivers and DNS") to see one container. Now look at the api's sockets, in two ways:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ sudo nsenter -n -t $(docker inspect -f '{{.State.Pid}}' lab-trouble-api-1) ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess LISTEN 0 4096 127.0.0.11:40877 0.0.0.0:* users:(("dockerd",pid=17027,fd=35)) LISTEN 0 5 127.0.0.1:8000 0.0.0.0:* users:(("python3",pid=601953,fd=3))
$ docker run --rm --network container:lab-trouble-api-1 alpine:3.22 netstat -tln
Active Internet connections (only servers) Proto Recv-Q Send-Q Local Address Foreign Address State tcp 0 0 127.0.0.11:40877 0.0.0.0:* LISTEN tcp 0 0 127.0.0.1:8000 0.0.0.0:* LISTEN

python3 listens on 127.0.0.1:8000, a loopback address that exists only inside the api's own namespace, so connections arriving on its network interface are refused. The second command gets the same answer without root on the host: a throwaway container started with --network container:lab-trouble-api-1 shares the api's network namespace and brings its own tools, here BusyBox netstat from alpine. Any image with the tools you need works as that sidecar. "Container networking basics" (Docker for beginners) showed this loopback mistake from the outside; the fix is always in the application's configuration:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ sed -i 's/"--bind", "127.0.0.1"/"--bind", "0.0.0.0"/' compose.yaml docker compose up -d sleep 2 curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8090/
Container lab-trouble-web-1 Running Container lab-trouble-api-1 Recreate Container lab-trouble-api-1 Recreated Container lab-trouble-api-1 Starting Container lab-trouble-api-1 Started 200

4. The host's rules

When packets go out and nothing answers, look at what the host's firewall did with them. Every iptables rule counts the packets it matched:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ IP=$(hostname -I | awk '{print $1}') sudo iptables -t nat -L DOCKER -v -n | grep -e pkts -e 'dpt:8090' for i in 1 2 3; do curl -s -o /dev/null http://$IP:8090/; done curl -s -o /dev/null http://127.0.0.1:8090/ sudo iptables -t nat -L DOCKER -v -n | grep -e 'dpt:8090'
pkts bytes target prot opt in out source destination 0 0 DNAT tcp -- !br-716c10204831 * 0.0.0.0/0 0.0.0.0/0 tcp dpt:8090 to:10.89.21.3:80 3 180 DNAT tcp -- !br-716c10204831 * 0.0.0.0/0 0.0.0.0/0 tcp dpt:8090 to:10.89.21.3:80

The DNAT rule counted 3 packets for 3 requests to the host's address: one per connection, because only the first packet of a connection goes through the nat table and conntrack handles the rest. The request to 127.0.0.1 did not count at all, because it went through docker-proxy, not DNAT. That is how counters answer "did my traffic even get here": run the request, list again, and see which rule moved. A DROP rule whose counter rises is your answer; a DNAT rule that does not move means the packets never reached this host on that port, or reached it on a different address or protocol. curl localhost tries ::1 before 127.0.0.1, for example, which takes a different path. On the nftables backend, sudo nft list ruleset shows the same kind of counters on Docker's rules.

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ IP=$(hostname -I | awk '{print $1}') sudo conntrack -L -p tcp --orig-port-dst 8090 --orig-src $IP 2>/dev/null | sed 's/ \[ASSURED\]//; s/ mark=0 use=1//' sudo conntrack -L -p tcp --orig-port-dst 8090 --orig-src 127.0.0.1 2>/dev/null | sed 's/ \[ASSURED\]//; s/ mark=0 use=1//' | tail -n 1
tcp 6 119 TIME_WAIT src=192.168.2.4 dst=192.168.2.4 sport=36146 dport=8090 src=10.89.21.3 dst=192.168.2.4 sport=80 dport=36146 tcp 6 119 TIME_WAIT src=192.168.2.4 dst=192.168.2.4 sport=36128 dport=8090 src=10.89.21.3 dst=192.168.2.4 sport=80 dport=36128 tcp 6 119 TIME_WAIT src=192.168.2.4 dst=192.168.2.4 sport=36140 dport=8090 src=10.89.21.3 dst=192.168.2.4 sport=80 dport=36140 tcp 6 106 TIME_WAIT src=127.0.0.1 dst=127.0.0.1 sport=55262 dport=8090 src=127.0.0.1 dst=127.0.0.1 sport=8090 dport=55262
$ sudo iptables -S DOCKER-USER
-N DOCKER-USER

conntrack shows each translated connection: the original direction to 192.168.2.4:8090 and the reply expected from 10.89.21.3:80. Entries in SYN_SENT that never progress are the conntrack view of a timeout. DOCKER-USER is empty on this host, so no local rule of yours is filtering; on a host where someone added rules there, it is the first place to look when a published port times out from outside but works from the host itself, as the DOCKER-USER part of "Publishing ports and the packet path" demonstrated.

Two causes this lab does not reproduce. MTU problems show up as small requests working while large responses hang, typically on overlay networks or hosts behind a VPN: compare ip link MTUs inside the container and on the host path, and set com.docker.network.driver.mtu on the network to match the smallest one. And when the client is another machine, add the layers between the two hosts (cloud security groups, network ACLs) to question 4 before blaming Docker. Clean up:

ubuntu@secopslog-docker:~/lab/nettrouble · Docker 29.8.2
$ docker compose down; docker compose ps -a --format "{{.Name}}" | wc -l
Container lab-trouble-api-1 Stopping Container lab-trouble-web-1 Stopping Container lab-trouble-web-1 Stopped Container lab-trouble-web-1 Removing Container lab-trouble-web-1 Removed Container lab-trouble-api-1 Stopped Container lab-trouble-api-1 Removing Container lab-trouble-api-1 Removed Network lab-trouble_front Removing Network lab-trouble_back Removing Network lab-trouble_front Removed Network lab-trouble_back Removed 0
Quick check
01From another server, curl http://app-host:8080 times out. On app-host, curl http://127.0.0.1:8080 works. The DNAT rule's counter for 8080 rises with every remote attempt. What does that tell you?
Incorrect — Then the local request would fail too, since it also has to reach the container's interface.
Correct — A rising DNAT counter proves arrival; a timeout after it points at a drop in forwarding, while local requests take the proxy path.
Incorrect — The IPv4 DNAT counter rising shows the remote traffic arrives over IPv4.
Incorrect — A 127.0.0.1 mapping would not have a DNAT rule matching remote traffic to the host address.
02A tcpdump on the bridge shows web's SYN to 10.89.21.2:8000 followed immediately by a packet with flags R. from 10.89.21.2. Which conclusion is justified?
Incorrect — A drop produces silence and a client timeout, not a reset from the target.
Incorrect — Then web could not even send the SYN to that address over the bridge; here it arrived.
Correct — The target's kernel answered with a reset, so routing and rules worked; check the listening socket.
Incorrect — The reset came from 10.89.21.2 itself, so the address belongs to a live host on that network.
03A distroless API container has no shell, ss or netstat. You want to see which address it listens on without installing anything into it. Which approach works?
Correct — The sidecar shares the target's network namespace, so its netstat or ss sees the same sockets.
Incorrect — Distroless images have no package manager, and changing a running container is what you wanted to avoid.
Incorrect — docker port lists host-side mappings, not the application's listening address.
Incorrect — ExposedPorts comes from EXPOSE metadata and says nothing about what actually listens.

Try this

Work through “4. The host's rules” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from network troubleshooting, keep “4. The host's rules”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related