Troubleshooting a server, step by step
A method, applied to three real failures.
When something breaks on a server, what gets you to the cause fastest is rarely a clever command. It is a routine that makes you look at the right layer in the right order and stop guessing. This lesson gives you that routine and applies it to three failures broken on purpose on the lab server: a service that restarts until systemd gives up, a disk that stays full after you delete the file that filled it, and a connection that fails in three different ways. Almost every command comes from an earlier lesson (the one new tool, namei, is explained where it appears); what is new is the order.
The method: symptom, scope, change, then layer by layer
Start with the symptom, stated exactly: the command you ran and the message you got, copied rather than paraphrased. "The site is down" is a report; curl's (28) Connection timed out from one named client is a symptom you can work with.
Then the scope. Does it fail for one service or all of them, for one user or everyone, on this server only or on its neighbours too, and since when? A failure on every server points away from any single one of them. Then ask what changed, because most failures follow a change: a deployment, a package update (/var/log/apt/history.log), an edited configuration file, a restore from backup, a reboot, or slow growth such as a disk filling up. journalctl --since "2 hours ago" and the people who have access to the machine are both worth asking.
Then go through the layers, starting from the thing that fails and working outwards. At each layer one or two commands answer a yes-or-no question. Is the unit running (systemctl status)? What did it log (journalctl -u)? Are disk, inodes and memory available (df -h, df -i, free -h)? Does anything listen, can the client reach it, and does the name resolve (ss, curl, getent)? May the service's user reach every directory along the path (namei -l)? Does the configuration pass the program's own test (sshd -t, visudo -c)? Change one thing at a time, and check the symptom again after each change.
A service that keeps restarting
ts-inventory.service runs a small program as the system user ts-inventory. The program records each start in its data directory, then keeps running; set -e makes it stop at the first command that fails. To follow along, write the program and its unit with sudoedit (both files are root's):
#!/bin/sh# Stand-in for an application: record each start in its data directory, then keep running.set -edate -u +%FT%TZ >> /var/lib/ts-inventory/starts.logexec sleep infinity
[Unit]Description=Inventory sync[Service]Type=execUser=ts-inventoryExecStart=/usr/local/bin/ts-inventoryRestart=on-failure[Install]WantedBy=multi-user.target
Then create the system account (no home directory, no login shell, as in "Users, groups and the account files"), make the program executable, create its data directory owned by that account, and start the service:
It ran for days. This morning the data directory was restored from a backup as root and the service was restarted; these commands reproduce what happened:
Since then the inventory has not updated. The first question is which units have failed, and then what systemd knows about the one that did (-n 4 limits the log excerpt at the end of status to its last four lines):
status=2 is the program's own exit code. The Process: line says the program exited on its own (code=exited), and the journal lines tell the rest: Restart=on-failure scheduled restarts (the counter reached 5), and after too many starts in a short time (five in ten seconds by default) systemd stopped trying with Start request repeated too quickly and left the unit failed ("Services with systemd" explains the start limit). So the program fails as soon as it starts, and the next layer is its log:
The second line is the one that matters: line 4 of the script cannot create starts.log, Permission denied. In the third, systemd adds the conventional name for code 2, INVALIDARGUMENT, which tells you nothing about this program. The second command shows a trap. Filtering with -p err keeps only messages marked as errors (here the newest one), and systemd records everything a program prints at the ordinary info level, so the line with the cause disappears. Read a unit's log unfiltered first.
Permission denied means some directory along the path does not let the service's user do what it tried. namei -l lists every component of a path with its owner and mode, and running the same action as the service's user with sudo -u confirms the diagnosis before you change anything. It works for an account without a login shell, because sudo runs the command directly.
ts-inventory is owned by root with mode rwxr-xr-x, so the service's user may enter it but not create files in it. That is the change: the restore re-created the directory as root. (namei exits with status 1 because starts.log does not exist yet.) Give the directory back, restart, and check:
If you restart within ten seconds of the failures, systemd may still refuse with "start of the service was attempted too often"; sudo systemctl reset-failed ts-inventory clears that. The durable fix is to let systemd own the directory, which is the exercise at the end of this lesson.
A disk that stays full
ts-applog writes its log to /srv/ts-data/app.log, on a filesystem of its own. In the lab that filesystem is a 64 MiB file mounted through a loop device, so filling it never touches the real disk. Its unit says StandardOutput=append:/srv/ts-data/app.log: systemd opens the file once, and the program writes through that open file for as long as it runs. The lab filled the log to stand in for weeks of verbose logging. The symptom is the application complaining that it cannot write:
df shows the filesystem at 100% with nothing available, and the application's error, [Errno 28] No space left on device, is the kernel's answer to a write on a full filesystem. When df -h shows free space and writes still fail with this error, check df -i: the filesystem can also run out of inodes, the records that describe each file. The next question is which files use the space:
du adds up the files that have a name, and almost all of it is app.log. A tempting fix is to delete it:
Nothing came back. As "Disks, filesystems and mounts" explained, the kernel frees a file's blocks only when its last name is gone and no process has it open, and ts-applog still does. lsof +L1 lists open files with fewer than one link, meaning deleted but still open. lsof combines its selections with "or" unless you add -a ("and"), so +L1 with a path but without -a lists every deleted open file on the machine (in the lab, also an in-memory file that systemd keeps). With -a it lists only those on this filesystem:
That is the culprit: process ts-applog, file descriptor 1 (its standard output) open for writing (1w), about 57 MB, NLINK 0 and (deleted). Restarting the service closes the file:
The space is back, and the service has started a new, small app.log. What changed was a log deleted by hand instead of rotated: rotate logs with logrotate, covered in "Logs: the journal and /var/log", or log to the journal, which caps its own size.
A connection that fails
A client on this server must reach an API at http://app01.internal:8080/. In the lab, app01 is a network namespace called ts-app: a separate network stack on the same machine, with its own interface, address (203.0.113.218) and firewall, linked to this server's 203.0.113.217. Commands prefixed with sudo ip netns exec ts-app run inside it; on a real network you would log in to app01 and run the same command. The lab broke three things at once, so each fix uncovers the next message. curl -sS hides the progress meter but still prints errors, and each curl error has a number that points to a layer.
Error 6, Could not resolve host: curl never sent a packet to app01, because the name did not turn into an address. getent hosts asks the same resolver that applications use (/etc/hosts, then DNS, as "DNS and name resolution" showed); it prints nothing and exits with status 2. On a real network the fix is a DNS record. In the lab, an /etc/hosts line does the job:
sudo tee -a appends to a root-owned file, the pattern from "Pipes, redirection and exit status". The name resolves now, and the error changes to 28: packets went out and nothing came back within the five seconds --connect-timeout allowed. ping gets answers, so the host is up and the route works; only port 8080 goes unanswered, which is the signature of a firewall dropping packets. (Many cloud networks block ping, so a failed ping alone does not prove a host is down.) Look at the firewall on app01:
The rule tcp dport 8080 drop discarded every connection attempt without a reply, which is why curl waited. Writing firewall rules is a subject of the hardening course; this lab table exists only for the demonstration, so it is simply deleted. On a real server, find out why the rule is there first, then change that rule alone, with the firewall's own tool so the change survives a reload. The error changes to 7, Could not connect to server, and at once: app01 answered immediately and refused, which is what a host does when nothing listens on the port. The next layer is the listener and the service behind it:
ss -tln shows no listening sockets inside app01, and the unit is inactive. Once it starts, the request succeeds. On a real server, ask why it was stopped before you start it: its journal may say, and systemctl is-enabled shows whether it will come back after a reboot (sudo systemctl enable --now if it should).
A checklist to keep
1. Symptom the exact command and the exact message2. Scope one service or all? one client or all? this server only? since when?3. Change apt history.log, journalctl --since, recent edits, restores, reboots4. Unit systemctl --failed; systemctl status NAME5. Logs journalctl -u NAME -b, unfiltered first6. Resources df -h; df -i; free -h; uptime7. Network ss -tln on the server; curl -sS and ping from the client; getent hosts NAME8. Access namei -l PATH; the same action with sudo -u SERVICEUSER9. Config the program's own test (sshd -t, visudo -c) before any reload10. Fix ask why it was in that state; fix one thing, verify the symptom,write down what you changed
Try this
Make the first fix permanent. Break the service again with sudo chown root:root /var/lib/ts-inventory, then add a drop-in with sudo systemctl edit ts-inventory containing two lines, [Service] and StateDirectory=ts-inventory. Restart the service and run ls -ld /var/lib/ts-inventory: the directory belongs to ts-inventory again, because with StateDirectory= systemd creates /var/lib/NAME and gives it to the unit's user every time the service starts. Explain in one sentence why this prevents the failure from this lesson's first scenario coming back after the next restore. To remove the practice service afterwards, stop it, delete /etc/systemd/system/ts-inventory.service and its .d directory, /usr/local/bin/ts-inventory and /var/lib/ts-inventory with sudo rm -r, run sudo systemctl daemon-reload, and delete the account with sudo userdel ts-inventory.
Takeaway
Copy the exact error, ask what changed, and check one layer at a time, from the unit outwards, before you change anything. After each change, run the command that showed the symptom again.