The network stack, sockets and name resolution
Packet path, queues, drops and the DNS path.
A packet that arrives at a server passes through the network card, the driver, the kernel's IP and TCP code and a socket queue before an application sees a single byte, and it can be dropped at every one of those steps. Each step counts its drops in a different place. This lesson follows the receive and transmit paths, reads the counter for each stage, produces an accept-queue overflow on purpose, and then follows the other path every connection starts with: turning a name into an address. By the end you can say where a lost packet or a failed lookup was lost, which is what the troubleshooting tools of the next lesson (tcpdump, mtr, iperf3) build on. The essentials course covered ss and DNS basics; this lesson assumes them.
The receive path, and where each drop is counted
A received packet passes six stages, each with its own queue and its own drop counter. The diagram names them in order; the paragraphs after it explain each one.
The receive ring is a circular list of memory buffers that the driver hands to the network card. The card copies each incoming frame straight into the next free buffer (DMA, direct memory access, so the CPU does no copying) and raises an interrupt.
The driver does not process packets inside that interrupt. It schedules a poll (NAPI, the kernel's polling interface for network drivers) that runs soon afterwards in a software interrupt, or softirq: kernel work deferred from the interrupt and run on the same CPU. Each poll round takes packets from the ring in a batch, limited by a packet budget and a time budget.
Some packets wait in a per-CPU backlog queue instead: packets that the kernel steers to another CPU to spread the load, and packets from devices without a poll loop of their own, such as loopback. After that, the IP layer runs the netfilter hooks (nftables rules and connection tracking), makes a routing decision, and hands packets for this host to TCP or UDP. They put each packet on a socket's receive queue, or a new connection on a listener's accept queue, until the application reads or accepts it.
ip -s link shows the interface counters the driver reports. On the RX line, errors counts damaged frames (bad checksums, framing), usually a cable, optic or duplex problem. missed counts packets the device dropped because it had no free buffer: the ring was full because the host did not empty it fast enough. dropped is often misread as the same thing; the kernel defines it as packets received but not processed, for example an unsupported protocol or frames filtered out, and a full ring is explicitly excluded. So a climbing missed points at ring size or CPU time for the softirq, and a climbing dropped points at what the traffic contains. All are zero here. The lab VM's interface is eth0; real servers use predictable names such as enp0s1 or ens3.
ethtool -g shows the ring sizes: 256 entries for receive and transmit, which is also this virtio device's maximum (a property of the lab's virtual NIC; physical NICs report their own limits). Physical NICs often allow larger rings (ethtool -G), which absorbs bursts at the cost of memory and latency; the change lasts until reboot unless NetworkManager or a systemd .link file (RxBufferSize=) sets it. On most drivers a ring-size change tears down and re-creates the NIC's queues, which drops traffic for a moment and on some drivers bounces the link, taking your SSH session or a bond member with it: make the change in a maintenance window, from the console or a second path, one host at a time. The driver's own counters, ethtool -S eth0, name drops in the driver's terms.
The softirq takes at most 300 packets or 2000 microseconds per round, and each CPU's backlog holds up to 1000 packets. /proc/net/softnet_stat has one line per CPU, in hexadecimal. The first column is packets processed, the second packets dropped because that CPU's backlog queue was full, and the third time_squeeze: how often the softirq stopped with work left because it had used its budget. CPU 0 processed 0x35a27 (219687) packets and ran out of budget 0xe (14) times, with no drops. A few squeezes are normal; a squeeze count that climbs with traffic, or any drops, means the CPUs handling packets cannot keep up.
Connection tracking, the netfilter part that remembers every flow for stateful firewall rules and NAT, has a table of fixed size. When it is full, new flows are dropped and the kernel log says nf_conntrack: table full, dropping packet; none of the counters above moves. Compare the two sysctls, and on NAT gateways, container hosts and busy proxies watch them over time (conntrack -S, from the conntrack package, adds per-CPU drop and insert_failed counts).
The module is loaded on this VM because the lab's VM tool installs a NAT table for its DNS forwarding; a server without NAT or stateful firewall rules may not load it at all, and then these sysctls do not exist.
Above the IP layer the counters live in the kernel's SNMP-style statistics, which nstat prints (-a absolute values since boot, -s without updating its history file, -z including zeros). The third column is an average rate that nstat fills in only when a background nstat -d keeps collecting; for a one-off reading it stays 0.0. TcpExtTCPBacklogDrop counts packets dropped while the application held the socket's lock: during that time packets wait in a per-socket backlog, limited to about twice the receive buffer, and anything beyond it is dropped. TcpExtTCPRcvQDrop counts packets dropped because the socket's receive buffer could not take them. The listen counters are the subject of the next section.
On the transmit side, data goes from the socket's send buffer through TCP, IP and the netfilter output hooks to the queueing discipline (qdisc) and then the driver's transmit ring. tc -s qdisc shows the qdisc: its dropped counts packets discarded because its queue was full, and requeues counts packets handed back because the driver was busy. This interface has pfifo_fast, the kernel's built-in default, although Ubuntu sets net.core.default_qdisc = fq_codel (in /usr/lib/sysctl.d/): a qdisc is chosen when an interface comes up, and this cloud image's initramfs brings the interface up with systemd-networkd before the root filesystem's sysctl files are applied, so it keeps the built-in default for as long as it stays up. The Rocky VM's interface has fq_codel, the value its sysctl files set too. Check your own hosts rather than assume either. The difference matters when you read the counters: fq_codel drops packets on purpose to keep queues short (active queue management), so its dropped includes deliberate drops, not only a full queue, and its ecn_mark counts packets it marked instead of dropping.
Listening sockets: the SYN queue and the accept queue
A listening TCP socket has two queues. A client's SYN creates an entry in the SYN queue, a half-open connection waiting for the client's final ACK, bounded by net.ipv4.tcp_max_syn_backlog; when it overflows, SYN cookies (tcp_syncookies=1 by default) let the kernel answer without storing state, which is the defence against SYN floods. When the handshake completes, the connection moves to the accept queue and waits there until the application calls accept(). The accept queue's size is the backlog the application passed to listen(), capped by net.core.somaxconn (4096 since kernel 5.4). When it is full, the kernel drops new SYNs without replying and counts them.
To watch that happen without touching real services, the lab uses a network namespace, a separate copy of the network stack with its own interfaces, sockets and counters, and two small programs. The listener asks for a backlog of 2 and never accepts; the clients open connections at once and keep them.
#!/usr/bin/python3# k-net-listener: listen on 127.0.0.1:8080 with a small backlog and never accept.import socket, sys, timebacklog = int(sys.argv[1])s = socket.socket()s.bind(('127.0.0.1', 8080))s.listen(backlog) # finished handshakes wait in the accept queue, up to this limittime.sleep(3600) # a stuck application: it never calls accept()
#!/usr/bin/python3# k-net-clients: open N connections to 127.0.0.1:8080 at once and keep them open.import socket, sys, timeconns = []for _ in range(int(sys.argv[1])):c = socket.socket()c.setblocking(False)c.connect_ex(('127.0.0.1', 8080)) # sends the SYN and returns without waitingconns.append(c)time.sleep(3600)
Make both executable (chmod 755). ip netns add creates the namespace, and systemd's NetworkNamespacePath= runs each program inside it as a transient service under your own user.
For a listening socket, ss reuses its two queue columns: Recv-Q is the number of connections waiting in the accept queue and Send-Q is the queue's limit, the backlog of 2. Now six clients connect at once.
The accept queue holds 3, one more than the backlog, which is how Linux counts it. The three queued connections appear as ESTAB pairs: the kernel completed their handshakes although the application never accepted them. The other three clients are stuck in SYN-SENT. They were not refused, which would have produced an error at once; their SYNs were dropped, so they retransmit and the application behind them sees a connect call that hangs.
Every dropped SYN raised both ListenOverflows and ListenDrops (the second also counts other listen-time drops). The 18 are six SYNs from each of the three waiting clients: the first one and five retransmissions. On these kernels the first five SYN retransmissions are one second apart before the interval starts doubling (net.ipv4.tcp_syn_linear_timeouts=4 gives four linear timeouts, and the first exponential step is still one second), and TCPSynRetrans counts the clients' 15 retransmissions, since the namespace holds both ends. somaxconn is 4096 in the namespace too, so the limit of 2 came from the application.
This is the pattern to recognise on a real server: clients time out connecting, the host looks idle, ss -ltn shows Recv-Q at or above Send-Q on the port, and ListenOverflows keeps rising. The network is fine; the application is not calling accept() fast enough, because its workers are all busy or its event loop is blocked. Find out why before you raise the backlog: a larger queue only absorbs bursts, and lowering somaxconn below its default makes things worse. On an established connection the same two columns mean something else: Recv-Q is bytes received that the application has not read yet, and Send-Q bytes sent that the peer has not acknowledged, which the troubleshooting lesson uses.
Name resolution: the path an application takes
The essentials lesson "DNS and name resolution" introduced the sources and the two tools; this section follows the path in detail. Before most connections there is a name. An application calls getaddrinfo() in the C library, and glibc's Name Service Switch (NSS) reads the hosts: line of /etc/nsswitch.conf and asks each source in turn: files is /etc/hosts, dns is glibc's resolver, which sends queries to the nameserver in /etc/resolv.conf.
On Ubuntu that nameserver is 127.0.0.53, the local stub of systemd-resolved, and /etc/resolv.conf is a symbolic link to the file resolved writes. resolved forwards queries to the servers configured per interface, here 192.168.5.2, which the VM tool's network provides; on a real server they come from DHCP or netplan. dig is different from an application: it builds DNS queries itself and never consults nsswitch.conf or /etc/hosts. getent hosts goes through NSS exactly as an application does. Add a name to /etc/hosts and ask both.
getent finds the name through files. dig finds it too, which surprises people: resolved answers queries on its stub from /etc/hosts as well, and resolvectl query says so with Data from: synthetic. Asking the upstream server directly (@192.168.5.2) returns NXDOMAIN, because no DNS server knows the name. On Ubuntu, then, a plain dig already includes /etc/hosts; only a query to a specific server shows what DNS itself says.
RHEL takes a shorter path. As the essentials lesson showed, NetworkManager writes the upstream server straight into /etc/resolv.conf and systemd-resolved is not installed, so there is no local stub and a plain dig never sees /etc/hosts: there, getent and dig disagree about a name that only /etc/hosts has. What the hosts: line adds is worth a look.
The line ends in myhostname, an NSS module that answers for the machine's own name and for _gateway, the current default gateway, which no DNS server could answer. Because it comes last, a name in /etc/hosts or DNS still wins.
A disagreement between the two paths is the usual reason dig succeeds while an application fails, or the reverse: a stale /etc/hosts entry sends the application to an old address while DNS already has the new one, a hosts: line lists a different order or source, or on Ubuntu resolved uses a per-interface server that a manual dig @server bypasses. Reproduce what the application sees with getent ahosts NAME, and use dig @server to ask what a particular DNS server says. Finally remove the test entry.
getent exits with status 2 when the key is not found.
Try this
Recreate the namespace (sudo ip netns add k-net-lab and sudo ip -n k-net-lab link set lo up), start the listener with a backlog of 16 and the six clients again, wait five seconds, and predict sudo ip netns exec k-net-lab ss -ltn and the three nstat counters before you look. Expect Recv-Q 6 and Send-Q 16, every client established, and ListenOverflows, ListenDrops and TCPSynRetrans all at 0: a new namespace starts with fresh counters, and the queue had room. Stop both services with sudo systemctl stop k-net-clients k-net-listener and delete the namespace with sudo ip netns del k-net-lab.
Takeaway
When clients cannot connect to a host that looks idle, compare Recv-Q with Send-Q on the listener and watch ListenOverflows before you suspect the network; when a name resolves differently for dig and for the application, believe getent, because it takes the application's path.