Network troubleshooting: sockets, drops and packets

ss, nstat, ip -s, tcpdump, mtr and iperf3.

Advanced16 min · lesson 16 of 21

Network faults reach you as vague symptoms: "the API is down", "uploads are slow", "the service stops answering after a few days". This lesson builds a small test network out of network namespaces, breaks it in four ways that real servers break, and diagnoses each fault the same way: state the symptom, form a hypothesis, take the one measurement that tests it, read the output, and only then change something. The tools are ss, nstat, ip -s link, tcpdump, mtr and iperf3. The previous lesson showed where the kernel counts each drop; the essentials course covered ss -tlnp and the difference between a refused and a timed-out connection, and this lesson builds on both.

A test network you can break

A network namespace is a separate copy of the kernel's network stack with its own interfaces, addresses, routes, sockets and counters. A veth pair is a virtual cable: two interfaces, and whatever enters one leaves the other. With three namespaces and two cables you get a client, a router and a server, and every fault you inject stays inside them, away from the host's own interface and your SSH session. The addresses come from the ranges RFC 5737 reserves for documentation (192.0.2.0/24 and 198.51.100.0/24).

The lab network
1nd-client
192.0.2.10 on nd-c0
2nd-router
192.0.2.1 on nd-r0, 198.51.100.1 on nd-r1
3nd-server
198.51.100.10 on nd-s0
Faults are added on the router (tc netem) and on the server (the listener, a firewall rule).

Save the setup as a script, and save the small server from the last section (/var/tmp/nd-leaky.py) now as well. The first command below makes both executable and creates the one-line page that the web server in the first fault serves; then run the setup with sudo.

/var/tmp/nd-net.sh
#!/bin/bash
# nd-net.sh: a client, a router and a server, each in its own network namespace.
set -e
for ns in nd-client nd-router nd-server; do
ip netns add $ns
ip -n $ns link set lo up
done
# two virtual cables (veth pairs): client <-> router and router <-> server
ip link add nd-c0 netns nd-client type veth peer name nd-r0 netns nd-router
ip link add nd-r1 netns nd-router type veth peer name nd-s0 netns nd-server
ip -n nd-client addr add 192.0.2.10/24 dev nd-c0
ip -n nd-router addr add 192.0.2.1/24 dev nd-r0
ip -n nd-router addr add 198.51.100.1/24 dev nd-r1
ip -n nd-server addr add 198.51.100.10/24 dev nd-s0
ip -n nd-client link set nd-c0 up
ip -n nd-router link set nd-r0 up
ip -n nd-router link set nd-r1 up
ip -n nd-server link set nd-s0 up
ip -n nd-client route add default via 192.0.2.1
ip -n nd-server route add default via 198.51.100.1
ip netns exec nd-router sysctl -qw net.ipv4.ip_forward=1
deploy@web01 · Ubuntu 26.04 LTS
$ chmod 755 /var/tmp/nd-net.sh /var/tmp/nd-leaky.py mkdir -p /var/tmp/nd-www echo 'hello from nd-server' > /var/tmp/nd-www/index.html
$ sudo /var/tmp/nd-net.sh ip netns list
nd-server nd-router nd-client
$ sudo ip netns exec nd-client ping -c 3 198.51.100.10
PING 198.51.100.10 (198.51.100.10) 56(84) bytes of data. 64 bytes from 198.51.100.10: icmp_seq=1 ttl=63 time=2.08 ms 64 bytes from 198.51.100.10: icmp_seq=2 ttl=63 time=0.178 ms 64 bytes from 198.51.100.10: icmp_seq=3 ttl=63 time=0.338 ms --- 198.51.100.10 ping statistics --- 3 packets transmitted, 3 received, 0% packet loss, time 2031ms rtt min/avg/max/mdev = 0.178/0.865/2.081/0.861 ms

ip netns exec NAME command runs a command inside a namespace, and it needs root, so every command below starts with sudo ip netns exec. The reply's ttl=63 shows the path: the server answered with a time-to-live of 64 and the router took one off. On Ubuntu Server, tcpdump and mtr-tiny come with the default install; iperf3 does not (sudo apt install iperf3). On RHEL, sudo dnf install tcpdump mtr iperf3. Scanning a host's open ports with nmap is an exposure check and lives in the hardening course's services lesson.

Refused or silent: read the first packets

The first fault is a web server started with the wrong bind address. systemd-run starts it as a transient service under your own user, and NetworkNamespacePath= puts it inside the server namespace. The symptom arrives from the client.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=nd-web --uid=$USER -p NetworkNamespacePath=/run/netns/nd-server python3 -m http.server -b 127.0.0.1 -d /var/tmp/nd-www 8080
Running as unit: nd-web.service; invocation ID: 66f39e3c694d48b2a0609ef333954fdb
$ sudo ip netns exec nd-client curl -sS http://198.51.100.10:8080/
curl: (7) Failed to connect to 198.51.100.10 port 8080 after 0 ms: Could not connect to server

"Could not connect" after 0 ms (under a millisecond) is fast, and a fast failure means something answered. The hypothesis is that nothing listens on 198.51.100.10:8080, so check the listening sockets on the server.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-server ss -ltnp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess LISTEN 0 5 127.0.0.1:8080 0.0.0.0:* users:(("python3",pid=416210,fd=3))

The server listens on 127.0.0.1:8080, the loopback address, which only processes inside the same namespace can reach. To see what the client actually received, capture the packets. tcpdump -n prints addresses instead of names, -i picks the interface, the expression tcp port 8080 is a filter that the kernel applies before copying anything, and -c 2 stops after two packets. Here it runs in the background (&) while curl makes the request; in practice you run it in a second terminal.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-server tcpdump -ni nd-s0 -c 2 'tcp port 8080' & sleep 1 sudo ip netns exec nd-client curl -sS http://198.51.100.10:8080/ wait
tcpdump: verbose output suppressed, use -v[v]... for full protocol decode listening on nd-s0, link-type EN10MB (Ethernet), snapshot length 262144 bytes curl: (7) Failed to connect to 198.51.100.10 port 8080 after 0 ms: Could not connect to server 13:10:09.885085 IP 192.0.2.10.38704 > 198.51.100.10.8080: Flags [S], seq 2854800208, win 64240, options [mss 1460,sackOK,TS val 881206751 ecr 0,nop,wscale 9], length 0 13:10:09.885109 IP 198.51.100.10.8080 > 192.0.2.10.38704: Flags [R.], seq 0, ack 2854800209, win 0, length 0 2 packets captured 2 packets received by filter 0 packets dropped by kernel

The client's SYN (Flags [S]) arrived, and the server's kernel answered at once with a reset ([R.], RST plus ACK): the host is reachable and nothing listens on that address and port. Restart the service bound to the server's address and the request succeeds.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop nd-web sudo systemd-run --unit=nd-web --uid=$USER -p NetworkNamespacePath=/run/netns/nd-server python3 -m http.server -b 198.51.100.10 -d /var/tmp/nd-www 8080
Running as unit: nd-web.service; invocation ID: fee23ea8fb8347e7b544484836dfbc68
$ sudo ip netns exec nd-client curl -sS http://198.51.100.10:8080/
hello from nd-server

A firewall that drops packets produces the other symptom. Add a drop rule for port 8080 inside the server namespace (the optional hardening course teaches nftables; here it only simulates a firewall) and capture on the client's side. -ttt prints the time since the previous packet.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-server nft add table inet nd sudo ip netns exec nd-server nft add chain inet nd input '{ type filter hook input priority 0; }' sudo ip netns exec nd-server nft add rule inet nd input tcp dport 8080 drop
$ sudo ip netns exec nd-client tcpdump -ni nd-c0 -ttt -c 4 'tcp port 8080' & sleep 1 sudo ip netns exec nd-client curl -sS --max-time 4 http://198.51.100.10:8080/ wait
tcpdump: verbose output suppressed, use -v[v]... for full protocol decode listening on nd-c0, link-type EN10MB (Ethernet), snapshot length 262144 bytes 00:00:00.000000 IP 192.0.2.10.37738 > 198.51.100.10.8080: Flags [S], seq 4083699330, win 64240, options [mss 1460,sackOK,TS val 403285672 ecr 0,nop,wscale 9], length 0 00:00:01.039178 IP 192.0.2.10.37738 > 198.51.100.10.8080: Flags [S], seq 4083699330, win 64240, options [mss 1460,sackOK,TS val 403286711 ecr 0,nop,wscale 9], length 0 00:00:01.023838 IP 192.0.2.10.37738 > 198.51.100.10.8080: Flags [S], seq 4083699330, win 64240, options [mss 1460,sackOK,TS val 403287735 ecr 0,nop,wscale 9], length 0 00:00:01.023748 IP 192.0.2.10.37738 > 198.51.100.10.8080: Flags [S], seq 4083699330, win 64240, options [mss 1460,sackOK,TS val 403288759 ecr 0,nop,wscale 9], length 0 4 packets captured 4 packets received by filter 0 packets dropped by kernel curl: (28) Connection timed out after 4002 milliseconds
$ sudo ip netns exec nd-server nft delete table inet nd

The same SYN, with the same sequence number, leaves four times about one second apart and nothing comes back, until curl gives up after its four-second limit. The one-second spacing is not a fixed rule: these kernels retransmit the first SYNs at a linear one-second interval (net.ipv4.tcp_syn_linear_timeouts=4) and only then double the wait. The reading is what matters. A reset means something actively refused the connection: most often a reachable host with nothing on that port, but a firewall rule that rejects with a TCP reset (nftables reject with tcp reset), a load balancer or another middlebox answers the same way, so check which address the reset came from. Silence means the packet was dropped on the way or on arrival: a firewall rule, a routing or ARP problem, or a host that is down. "Nothing is listening" never produces silence.

Captures contain other people's data
tcpdump copies packets with their payload, and anything not encrypted (tokens, cookies, form data) is in the capture. It needs root or the CAP_NET_RAW capability. Use the narrowest filter that answers your question, limit it with -c, write to a file with -w only when you need to keep it, and delete the file when you are done. On a busy interface the capture also costs CPU and disk: keep the filter narrow, use -s 128 when headers are enough, check the packets dropped by kernel line tcpdump prints at the end (non-zero means the capture missed packets), and bound files with -C and -W (a ring of files of fixed size) or -G (rotate by time) so a capture cannot fill /var.

Loss and latency: mtr, iperf3 and the retransmit counters

The next symptom is "transfers to the server are slow". Before you blame anything, measure what the path can carry. iperf3 -s on the server listens on TCP port 5201; iperf3 -c on the client sends as fast as it can for five seconds and reports the throughput per second and in total, with Retr (segments TCP had to send again) and Cwnd (the congestion window: how much unacknowledged data TCP allows in flight).

iperf3 on a real network
Here iperf3 runs over veth pairs inside one VM. On a real path, an unthrottled test fills every shared link on the way (uplinks, VPN tunnels, WAN and cloud NAT) for as long as it runs, which is an outage you caused for everyone else on that path, and in a cloud it can cost egress fees. Agree a window, cap the rate with -b and keep -t short, prefer a path you own end to end, run the server as iperf3 -s -1 so it exits after one test instead of leaving an unauthenticated listener on port 5201, and close any firewall opening afterwards.
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=nd-iperf --uid=$USER -p NetworkNamespacePath=/run/netns/nd-server iperf3 -s
Running as unit: nd-iperf.service; invocation ID: 17f1d2f18c554dea937c0344b2968d92
$ sudo ip netns exec nd-client iperf3 -c 198.51.100.10 -t 5
Connecting to host 198.51.100.10, port 5201 [ 5] local 192.0.2.10 port 49510 connected to 198.51.100.10 port 5201 [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.00 sec 11.3 GBytes 96.5 Gbits/sec 0 1.83 MBytes [ 5] 1.00-2.00 sec 11.4 GBytes 97.7 Gbits/sec 0 2.25 MBytes [ 5] 2.00-3.00 sec 11.4 GBytes 97.9 Gbits/sec 0 2.25 MBytes [ 5] 3.00-4.00 sec 11.3 GBytes 97.5 Gbits/sec 0 2.63 MBytes [ 5] 4.00-5.00 sec 11.2 GBytes 95.8 Gbits/sec 0 2.63 MBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-5.00 sec 56.6 GBytes 97.2 Gbits/sec 0 sender [ 5] 0.00-5.00 sec 56.6 GBytes 97.2 Gbits/sec receiver …

About 97 Gbit/s, because a veth pair is a copy in memory and the figure measures this VM's CPUs, not a network; a real 10 Gbit/s link tops out near 9.4. The figure changes from run to run with whatever else the CPUs are doing (other runs of this lab measured 68 and 81 Gbit/s). Retr counts TCP segments sent again: 0 in this run, while other runs of the same test counted a few thousand, out of tens of millions of segments, because queues overflow now and then at that speed. A retransmission count means little until you set it against the number of segments sent. Now make the path bad. tc qdisc add ... netem attaches the network emulator as the queueing discipline of the router's interface toward the server, and every packet leaving it is delayed by 40 ms and dropped with a 5% probability.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-router tc qdisc add dev nd-r1 root netem delay 40ms loss 5%
$ sudo ip netns exec nd-client mtr -n -r -c 50 -i 0.2 198.51.100.10
Start: 2026-09-27T13:10:21+0000 HOST: web01 Loss% Snt Last Avg Best Wrst StDev 1.|-- 192.0.2.1 70.0% 50 0.1 0.2 0.1 0.6 0.2 2.|-- 198.51.100.10 10.0% 50 45.1 43.1 40.5 45.7 1.7
$ sudo ip netns exec nd-client mtr -n -r -c 20 198.51.100.10
Start: 2026-09-27T13:10:36+0000 HOST: web01 Loss% Snt Last Avg Best Wrst StDev 1.|-- 192.0.2.1 0.0% 20 0.1 0.2 0.1 0.6 0.1 2.|-- 198.51.100.10 0.0% 20 44.5 43.8 40.6 45.7 1.7

mtr sends probes with an increasing time-to-live, so each router on the way answers with an ICMP "time exceeded" message, and it keeps doing so to show loss and latency per hop (-r prints a report after -c cycles, -n skips name lookups). The first report, with five probes a second (-i 0.2, which only root may use), shows 70% loss at the router and 10% at the destination. The router is not losing most packets: if it were, the destination behind it could not answer 90% of the probes. The kernel limits how often it sends ICMP errors to the same address (net.ipv4.icmp_ratelimit, one message per second with a small burst allowance, and time-exceeded is one of the rate-limited types), and routers from every vendor do something similar. The 10% at the destination is the real fault, the emulator's 5% loss (with 50 probes the measured rate is rough: 5 lost here, 1 in another run). At the default one-second interval the router answers every probe, and the 40 ms of added latency from hop 2 on is plain; this run's 20 probes happened to lose none at the destination, which at 5% loss happens about one time in three. Loss that starts at a hop and continues to the destination is real, loss at one hop that disappears after it is that router's reply policy, and a small loss rate needs a hundred probes or more before a 0% means anything. Where ICMP is filtered, mtr -T -P 443 probes with TCP SYNs instead.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-client iperf3 -c 198.51.100.10 -t 5
Connecting to host 198.51.100.10, port 5201 [ 5] local 192.0.2.10 port 38684 connected to 198.51.100.10 port 5201 [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-1.01 sec 640 KBytes 5.22 Mbits/sec 12 9.90 KBytes [ 5] 1.01-2.00 sec 256 KBytes 2.10 Mbits/sec 5 11.3 KBytes [ 5] 2.00-3.00 sec 256 KBytes 2.10 Mbits/sec 10 8.48 KBytes [ 5] 3.00-4.00 sec 128 KBytes 1.05 Mbits/sec 6 7.07 KBytes [ 5] 4.00-5.00 sec 128 KBytes 1.05 Mbits/sec 8 7.07 KBytes - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-5.00 sec 1.38 MBytes 2.30 Mbits/sec 41 sender [ 5] 0.00-5.05 sec 1.12 MBytes 1.87 Mbits/sec receiver …

Throughput fell from 97 Gbit/s to about 2 Mbit/s (2.30 at the sender, 1.87 at the receiver), with 41 retransmissions in about 1,000 segments (1.38 MBytes). TCP treats loss as congestion and shrinks its congestion window (Cwnd) to a few segments, and with a 40 ms round trip a small window means little data per second. A few percent of loss does far more damage than the percentage suggests. ss -ti shows the same state from inside the kernel for a live connection, and -m adds socket memory.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-client iperf3 -c 198.51.100.10 -t 8 > /dev/null & sleep 4 sudo ip netns exec nd-client ss -tinm dst 198.51.100.10 wait
State Recv-Q Send-Q Local Address:Port Peer Address:Port … ESTAB 0 341728 192.0.2.10:38700 198.51.100.10:5201 skmem:(r0,rb131072,t0,tb470016,f3104,w353248,o0,bl0,d0) cubic wscale:9,9 rto:243 rtt:42.944/0.373 mss:1448 pmtu:1500 rcvmss:536 advmss:1448 cwnd:6 ssthresh:9 bytes_sent:1268485 bytes_retrans:50680 bytes_acked:1200430 segs_out:879 segs_in:220 data_segs_out:877 send 1618480bps lastsnd:11 lastrcv:3842 lastack:11 pacing_rate 3884328bps delivery_rate 2703256bps delivered:835 busy:3841ms unacked:12 retrans:2/35 lost:4 sacked:4 rcv_space:14480 rcv_ssthresh:64088 notsent:324352 minrtt:40 snd_wnd:444928 rcv_wnd:64512

The first connection is iperf3's control channel; the second carries the data. rtt:42.944/0.373 is the smoothed round-trip time and its variation in milliseconds (40 ms of it is netem), cwnd:6 and ssthresh:9 show the window cut down after losses, and retrans:2/35 means 2 resent segments are still waiting for their acknowledgement and 35 were resent over the connection's life (bytes_retrans:50680 is the same 35 segments of 1448 bytes). On an established connection Send-Q counts bytes the peer has not acknowledged yet or that still wait in the socket: here 12 segments are in flight (unacked:12) and 324,352 bytes have not been sent at all (notsent). The skmem field from -m shows the memory behind that queue (w353248 bytes queued for sending against a send buffer tb of 470,016), so the sender is limited by the network, not by the application. nstat gives the host-wide view; in a namespace the counters belong to that namespace alone.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-client nstat -asz TcpOutSegs TcpRetransSegs TcpExtTCPLostRetransmit TcpExtTCPSackRecovery TcpExtTCPTimeouts
#kernel TcpOutSegs 42225606 0.0 TcpRetransSegs 117 0.0 TcpExtTCPSackRecovery 50 0.0 TcpExtTCPLostRetransmit 4 0.0 TcpExtTCPTimeouts 3 0.0
$ sudo ip netns exec nd-router tc -s qdisc show dev nd-r1
qdisc netem 8007: root refcnt 3 limit 1000 delay 40ms loss 5% seed 2424876971619731161 Sent 3441441 bytes 2378 pkt (dropped 72, overlimits 0 requeues 0) backlog 67b 1p requeues 0
$ sudo ip netns exec nd-router ip -s link show nd-r1
3: nd-r1@if2: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc netem state UP mode DEFAULT group default qlen 1000 … RX: bytes packets errors dropped missed mcast 32884639 498160 0 0 0 0 TX: bytes packets errors dropped carrier collsns 60916081748 1389669 0 0 0 0

TcpRetransSegs against TcpOutSegs is the retransmission rate (117 of 42.2 million: almost every segment was sent by the fast baseline run and almost every retransmission came from the lossy runs, which is why a rate over an interval is more useful than totals since boot: nstat without -a prints the change since its last run). TcpExtTCPTimeouts counts retransmission timeouts, which stall a connection far longer than a fast retransmit. The last two commands show where this loss was counted: the netem qdisc reports dropped 72, while ip -s link for the same interface shows 0 dropped and 0 missed, because a qdisc drop happens before the packet reaches the driver. On a real host, loss inside the network never appears in your own counters at all: the retransmissions are the evidence. Remove the emulator and stop iperf3 when you are done.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-router tc qdisc del dev nd-r1 root sudo systemctl stop nd-iperf

Connections that never close: CLOSE-WAIT

The last symptom builds up slowly: a service works after a restart, gets slower over days, and finally logs "Too many open files" (the error EMFILE: the process has used up its file-descriptor limit, 1024 by default for a systemd service). Every connection is a file descriptor, so the hypothesis is connections the program never closes. This small server has that bug.

/var/tmp/nd-leaky.py
#!/usr/bin/python3
# nd-leaky: answer every request, but never close the connection (the bug).
import socket
srv = socket.create_server(('0.0.0.0', 9090))
kept = []
while True:
conn, peer = srv.accept()
conn.recv(1024)
conn.sendall(b'HTTP/1.0 200 OK\r\nContent-Length: 3\r\n\r\nok\n')
kept.append(conn) # the bug: the reply is sent, but conn.close() is never called
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --unit=nd-leaky --uid=$USER -p NetworkNamespacePath=/run/netns/nd-server /var/tmp/nd-leaky.py
Running as unit: nd-leaky.service; invocation ID: 0f71762af5eb4ff58e400dda9f91ba4b
$ sudo ip netns exec nd-client sh -c 'for i in $(seq 20); do curl -s http://198.51.100.10:9090/ > /dev/null; done'
$ sudo ip netns exec nd-server ss -tn state close-wait
Recv-Q Send-Q Local Address:Port Peer Address:Port 1 0 198.51.100.10:9090 192.0.2.10:50042 1 0 198.51.100.10:9090 192.0.2.10:50052 1 0 198.51.100.10:9090 192.0.2.10:49940 …

Each client read its reply and closed its end, which sends a FIN. The server's kernel acknowledged the FIN and moved the socket to CLOSE-WAIT: the peer is done, and the socket waits for the application to call close(). The kernel never does that on the application's behalf, so a CLOSE-WAIT socket lives as long as the process holds it. Recv-Q is 1 on every line because the FIN occupies one sequence number and the application never read up to it: it stopped using the socket altogether.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec nd-server ss -Htn state close-wait | wc -l sudo ip netns exec nd-server ss -Htnp state close-wait | head -1 ls /proc/$(systemctl show -P MainPID nd-leaky)/fd | wc -l
20 1 0 198.51.100.10:9090 192.0.2.10:50042 users:(("nd-leaky.py",pid=416882,fd=17)) 24
$ sudo ip netns exec nd-client ss -Htn state fin-wait-2 | wc -l
20

Twenty sockets in CLOSE-WAIT, all owned by nd-leaky.py, and 24 open descriptors (the standard three, the listening socket and the twenty leaked connections). The clients' side of each connection sits in FIN-WAIT-2, waiting for a FIN that never comes; Linux gives up on those after net.ipv4.tcp_fin_timeout (60 s), so the client host recovers by itself and the server does not. A CLOSE-WAIT count that only grows is a bug in the program, and no kernel setting fixes it. Restarting releases the descriptors, which buys time until the fix (closing the socket on every path, including errors) is deployed. Do not confuse it with TIME-WAIT, which appears on the side that closed first, lasts 60 seconds and is normal.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop nd-leaky sudo ip netns exec nd-server ss -Htn state close-wait | wc -l
0

Try this

Move the fault to the other link: add netem loss 20% to nd-r0 (the router's interface toward the client) with sudo ip netns exec nd-router tc qdisc add dev nd-r0 root netem loss 20%, and predict what sudo ip netns exec nd-client mtr -n -r -c 20 198.51.100.10 shows before you run it. This time the replies from both hops cross the lossy link on their way back, so both hops show loss (5% and 20% in the lab run, 15% and 15% in another; with 20 probes the percentages are rough), and tc -s qdisc show dev nd-r0 in the router shows the drops. Remove the qdisc with sudo ip netns exec nd-router tc qdisc del dev nd-r0 root. Then tear the network down: deleting a namespace removes its interfaces and qdiscs with it.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl stop nd-web for ns in nd-client nd-router nd-server; do sudo ip netns del $ns; done ip netns list

Takeaway

Let the first packets and the socket states tell you which layer is at fault before you change anything: a reset means something refused the connection (usually a host with no listener), silence means something dropped the packet, loss that continues to the destination is real, and a growing CLOSE-WAIT count is a bug in the application.

Quick check
01A client's connections to a partner API on port 443 time out after 30 seconds. tcpdump on the client shows the same SYN with an unchanged sequence number leaving several times, and nothing ever comes back. What does that rule out?
Incorrect — This is one of the explanations that fit: a drop produces exactly this silence, so it is not ruled out.
Incorrect — An unreachable host also gives silence (or at most an ICMP error), so this stays a candidate.
Correct — A reachable host with no listener answers the SYN with a reset at once, so silence rules this out.
Incorrect — A rule that discards packets is a drop, and drops look exactly like this capture.
02mtr -i 0.2 -c 100 to a database shows 65% loss at hop 3 and 0% loss at hops 4 to 6, including the database itself. Queries are slow. What do you conclude about hop 3?
Incorrect — If hop 3 dropped traffic, every hop after it would show at least that loss; hops 4 to 6 show none.
Correct — Fast probing makes a router's ICMP rate limit visible as loss at that hop only; the destination answers every probe.
Incorrect — Replies can take other paths, but the probes to hops 4 to 6 must cross hop 3, and they arrive.
Incorrect — 100 cycles is plenty; the pattern of loss at one hop only is the answer, and the slowness is elsewhere.
03An API service gets slower over several days and then logs "Too many open files". ss -tn state close-wait on the server shows 1,000 sockets, all owned by the service. What is the right next step?
Incorrect — A higher limit only delays the failure; the leaked sockets keep piling up until the new limit is reached.
Incorrect — tcp_fin_timeout applies to FIN-WAIT-2 on the side that closed first; the kernel never closes a CLOSE-WAIT socket itself.
Correct — CLOSE-WAIT means the peer closed and the application never called close(); only the application can release it.
Incorrect — The clients did close: that FIN is what put the server's sockets into CLOSE-WAIT.

Related