Network troubleshooting: sockets, drops and packets
ss, nstat, ip -s, tcpdump, mtr and iperf3.
Network faults reach you as vague symptoms: "the API is down", "uploads are slow", "the service stops answering after a few days". This lesson builds a small test network out of network namespaces, breaks it in four ways that real servers break, and diagnoses each fault the same way: state the symptom, form a hypothesis, take the one measurement that tests it, read the output, and only then change something. The tools are ss, nstat, ip -s link, tcpdump, mtr and iperf3. The previous lesson showed where the kernel counts each drop; the essentials course covered ss -tlnp and the difference between a refused and a timed-out connection, and this lesson builds on both.
A test network you can break
A network namespace is a separate copy of the kernel's network stack with its own interfaces, addresses, routes, sockets and counters. A veth pair is a virtual cable: two interfaces, and whatever enters one leaves the other. With three namespaces and two cables you get a client, a router and a server, and every fault you inject stays inside them, away from the host's own interface and your SSH session. The addresses come from the ranges RFC 5737 reserves for documentation (192.0.2.0/24 and 198.51.100.0/24).
Save the setup as a script, and save the small server from the last section (/var/tmp/nd-leaky.py) now as well. The first command below makes both executable and creates the one-line page that the web server in the first fault serves; then run the setup with sudo.
#!/bin/bash# nd-net.sh: a client, a router and a server, each in its own network namespace.set -efor ns in nd-client nd-router nd-server; doip netns add $nsip -n $ns link set lo updone# two virtual cables (veth pairs): client <-> router and router <-> serverip link add nd-c0 netns nd-client type veth peer name nd-r0 netns nd-routerip link add nd-r1 netns nd-router type veth peer name nd-s0 netns nd-serverip -n nd-client addr add 192.0.2.10/24 dev nd-c0ip -n nd-router addr add 192.0.2.1/24 dev nd-r0ip -n nd-router addr add 198.51.100.1/24 dev nd-r1ip -n nd-server addr add 198.51.100.10/24 dev nd-s0ip -n nd-client link set nd-c0 upip -n nd-router link set nd-r0 upip -n nd-router link set nd-r1 upip -n nd-server link set nd-s0 upip -n nd-client route add default via 192.0.2.1ip -n nd-server route add default via 198.51.100.1ip netns exec nd-router sysctl -qw net.ipv4.ip_forward=1
ip netns exec NAME command runs a command inside a namespace, and it needs root, so every command below starts with sudo ip netns exec. The reply's ttl=63 shows the path: the server answered with a time-to-live of 64 and the router took one off. On Ubuntu Server, tcpdump and mtr-tiny come with the default install; iperf3 does not (sudo apt install iperf3). On RHEL, sudo dnf install tcpdump mtr iperf3. Scanning a host's open ports with nmap is an exposure check and lives in the hardening course's services lesson.
Refused or silent: read the first packets
The first fault is a web server started with the wrong bind address. systemd-run starts it as a transient service under your own user, and NetworkNamespacePath= puts it inside the server namespace. The symptom arrives from the client.
"Could not connect" after 0 ms (under a millisecond) is fast, and a fast failure means something answered. The hypothesis is that nothing listens on 198.51.100.10:8080, so check the listening sockets on the server.
The server listens on 127.0.0.1:8080, the loopback address, which only processes inside the same namespace can reach. To see what the client actually received, capture the packets. tcpdump -n prints addresses instead of names, -i picks the interface, the expression tcp port 8080 is a filter that the kernel applies before copying anything, and -c 2 stops after two packets. Here it runs in the background (&) while curl makes the request; in practice you run it in a second terminal.
The client's SYN (Flags [S]) arrived, and the server's kernel answered at once with a reset ([R.], RST plus ACK): the host is reachable and nothing listens on that address and port. Restart the service bound to the server's address and the request succeeds.
A firewall that drops packets produces the other symptom. Add a drop rule for port 8080 inside the server namespace (the optional hardening course teaches nftables; here it only simulates a firewall) and capture on the client's side. -ttt prints the time since the previous packet.
The same SYN, with the same sequence number, leaves four times about one second apart and nothing comes back, until curl gives up after its four-second limit. The one-second spacing is not a fixed rule: these kernels retransmit the first SYNs at a linear one-second interval (net.ipv4.tcp_syn_linear_timeouts=4) and only then double the wait. The reading is what matters. A reset means something actively refused the connection: most often a reachable host with nothing on that port, but a firewall rule that rejects with a TCP reset (nftables reject with tcp reset), a load balancer or another middlebox answers the same way, so check which address the reset came from. Silence means the packet was dropped on the way or on arrival: a firewall rule, a routing or ARP problem, or a host that is down. "Nothing is listening" never produces silence.
-c, write to a file with -w only when you need to keep it, and delete the file when you are done. On a busy interface the capture also costs CPU and disk: keep the filter narrow, use -s 128 when headers are enough, check the packets dropped by kernel line tcpdump prints at the end (non-zero means the capture missed packets), and bound files with -C and -W (a ring of files of fixed size) or -G (rotate by time) so a capture cannot fill /var.Loss and latency: mtr, iperf3 and the retransmit counters
The next symptom is "transfers to the server are slow". Before you blame anything, measure what the path can carry. iperf3 -s on the server listens on TCP port 5201; iperf3 -c on the client sends as fast as it can for five seconds and reports the throughput per second and in total, with Retr (segments TCP had to send again) and Cwnd (the congestion window: how much unacknowledged data TCP allows in flight).
-b and keep -t short, prefer a path you own end to end, run the server as iperf3 -s -1 so it exits after one test instead of leaving an unauthenticated listener on port 5201, and close any firewall opening afterwards.About 97 Gbit/s, because a veth pair is a copy in memory and the figure measures this VM's CPUs, not a network; a real 10 Gbit/s link tops out near 9.4. The figure changes from run to run with whatever else the CPUs are doing (other runs of this lab measured 68 and 81 Gbit/s). Retr counts TCP segments sent again: 0 in this run, while other runs of the same test counted a few thousand, out of tens of millions of segments, because queues overflow now and then at that speed. A retransmission count means little until you set it against the number of segments sent. Now make the path bad. tc qdisc add ... netem attaches the network emulator as the queueing discipline of the router's interface toward the server, and every packet leaving it is delayed by 40 ms and dropped with a 5% probability.
mtr sends probes with an increasing time-to-live, so each router on the way answers with an ICMP "time exceeded" message, and it keeps doing so to show loss and latency per hop (-r prints a report after -c cycles, -n skips name lookups). The first report, with five probes a second (-i 0.2, which only root may use), shows 70% loss at the router and 10% at the destination. The router is not losing most packets: if it were, the destination behind it could not answer 90% of the probes. The kernel limits how often it sends ICMP errors to the same address (net.ipv4.icmp_ratelimit, one message per second with a small burst allowance, and time-exceeded is one of the rate-limited types), and routers from every vendor do something similar. The 10% at the destination is the real fault, the emulator's 5% loss (with 50 probes the measured rate is rough: 5 lost here, 1 in another run). At the default one-second interval the router answers every probe, and the 40 ms of added latency from hop 2 on is plain; this run's 20 probes happened to lose none at the destination, which at 5% loss happens about one time in three. Loss that starts at a hop and continues to the destination is real, loss at one hop that disappears after it is that router's reply policy, and a small loss rate needs a hundred probes or more before a 0% means anything. Where ICMP is filtered, mtr -T -P 443 probes with TCP SYNs instead.
Throughput fell from 97 Gbit/s to about 2 Mbit/s (2.30 at the sender, 1.87 at the receiver), with 41 retransmissions in about 1,000 segments (1.38 MBytes). TCP treats loss as congestion and shrinks its congestion window (Cwnd) to a few segments, and with a 40 ms round trip a small window means little data per second. A few percent of loss does far more damage than the percentage suggests. ss -ti shows the same state from inside the kernel for a live connection, and -m adds socket memory.
The first connection is iperf3's control channel; the second carries the data. rtt:42.944/0.373 is the smoothed round-trip time and its variation in milliseconds (40 ms of it is netem), cwnd:6 and ssthresh:9 show the window cut down after losses, and retrans:2/35 means 2 resent segments are still waiting for their acknowledgement and 35 were resent over the connection's life (bytes_retrans:50680 is the same 35 segments of 1448 bytes). On an established connection Send-Q counts bytes the peer has not acknowledged yet or that still wait in the socket: here 12 segments are in flight (unacked:12) and 324,352 bytes have not been sent at all (notsent). The skmem field from -m shows the memory behind that queue (w353248 bytes queued for sending against a send buffer tb of 470,016), so the sender is limited by the network, not by the application. nstat gives the host-wide view; in a namespace the counters belong to that namespace alone.
TcpRetransSegs against TcpOutSegs is the retransmission rate (117 of 42.2 million: almost every segment was sent by the fast baseline run and almost every retransmission came from the lossy runs, which is why a rate over an interval is more useful than totals since boot: nstat without -a prints the change since its last run). TcpExtTCPTimeouts counts retransmission timeouts, which stall a connection far longer than a fast retransmit. The last two commands show where this loss was counted: the netem qdisc reports dropped 72, while ip -s link for the same interface shows 0 dropped and 0 missed, because a qdisc drop happens before the packet reaches the driver. On a real host, loss inside the network never appears in your own counters at all: the retransmissions are the evidence. Remove the emulator and stop iperf3 when you are done.
Connections that never close: CLOSE-WAIT
The last symptom builds up slowly: a service works after a restart, gets slower over days, and finally logs "Too many open files" (the error EMFILE: the process has used up its file-descriptor limit, 1024 by default for a systemd service). Every connection is a file descriptor, so the hypothesis is connections the program never closes. This small server has that bug.
#!/usr/bin/python3# nd-leaky: answer every request, but never close the connection (the bug).import socketsrv = socket.create_server(('0.0.0.0', 9090))kept = []while True:conn, peer = srv.accept()conn.recv(1024)conn.sendall(b'HTTP/1.0 200 OK\r\nContent-Length: 3\r\n\r\nok\n')kept.append(conn) # the bug: the reply is sent, but conn.close() is never called
Each client read its reply and closed its end, which sends a FIN. The server's kernel acknowledged the FIN and moved the socket to CLOSE-WAIT: the peer is done, and the socket waits for the application to call close(). The kernel never does that on the application's behalf, so a CLOSE-WAIT socket lives as long as the process holds it. Recv-Q is 1 on every line because the FIN occupies one sequence number and the application never read up to it: it stopped using the socket altogether.
Twenty sockets in CLOSE-WAIT, all owned by nd-leaky.py, and 24 open descriptors (the standard three, the listening socket and the twenty leaked connections). The clients' side of each connection sits in FIN-WAIT-2, waiting for a FIN that never comes; Linux gives up on those after net.ipv4.tcp_fin_timeout (60 s), so the client host recovers by itself and the server does not. A CLOSE-WAIT count that only grows is a bug in the program, and no kernel setting fixes it. Restarting releases the descriptors, which buys time until the fix (closing the socket on every path, including errors) is deployed. Do not confuse it with TIME-WAIT, which appears on the side that closed first, lasts 60 seconds and is normal.
Try this
Move the fault to the other link: add netem loss 20% to nd-r0 (the router's interface toward the client) with sudo ip netns exec nd-router tc qdisc add dev nd-r0 root netem loss 20%, and predict what sudo ip netns exec nd-client mtr -n -r -c 20 198.51.100.10 shows before you run it. This time the replies from both hops cross the lossy link on their way back, so both hops show loss (5% and 20% in the lab run, 15% and 15% in another; with 20 probes the percentages are rough), and tc -s qdisc show dev nd-r0 in the router shows the drops. Remove the qdisc with sudo ip netns exec nd-router tc qdisc del dev nd-r0 root. Then tear the network down: deleting a namespace removes its interfaces and qdiscs with it.
Takeaway
Let the first packets and the socket states tell you which layer is at fault before you change anything: a reset means something refused the connection (usually a host with no listener), silence means something dropped the packet, loss that continues to the destination is real, and a growing CLOSE-WAIT count is a bug in the application.