Troubleshooting a server, step by step

A method, applied to three real failures.

Beginner18 min · lesson 28 of 29

When something breaks on a server, what gets you to the cause fastest is rarely a clever command. It is a routine that makes you look at the right layer in the right order and stop guessing. This lesson gives you that routine and applies it to three failures broken on purpose on the lab server: a service that restarts until systemd gives up, a disk that stays full after you delete the file that filled it, and a connection that fails in three different ways. Almost every command comes from an earlier lesson (the one new tool, namei, is explained where it appears); what is new is the order.

The method: symptom, scope, change, then layer by layer

Start with the symptom, stated exactly: the command you ran and the message you got, copied rather than paraphrased. "The site is down" is a report; curl's (28) Connection timed out from one named client is a symptom you can work with.

Then the scope. Does it fail for one service or all of them, for one user or everyone, on this server only or on its neighbours too, and since when? A failure on every server points away from any single one of them. Then ask what changed, because most failures follow a change: a deployment, a package update (/var/log/apt/history.log), an edited configuration file, a restore from backup, a reboot, or slow growth such as a disk filling up. journalctl --since "2 hours ago" and the people who have access to the machine are both worth asking.

Then go through the layers, starting from the thing that fails and working outwards. At each layer one or two commands answer a yes-or-no question. Is the unit running (systemctl status)? What did it log (journalctl -u)? Are disk, inodes and memory available (df -h, df -i, free -h)? Does anything listen, can the client reach it, and does the name resolve (ss, curl, getent)? May the service's user reach every directory along the path (namei -l)? Does the configuration pass the program's own test (sshd -t, visudo -c)? Change one thing at a time, and check the symptom again after each change.

A service that keeps restarting

ts-inventory.service runs a small program as the system user ts-inventory. The program records each start in its data directory, then keeps running; set -e makes it stop at the first command that fails. To follow along, write the program and its unit with sudoedit (both files are root's):

/usr/local/bin/ts-inventory
#!/bin/sh
# Stand-in for an application: record each start in its data directory, then keep running.
set -e
date -u +%FT%TZ >> /var/lib/ts-inventory/starts.log
exec sleep infinity
/etc/systemd/system/ts-inventory.service
[Unit]
Description=Inventory sync
[Service]
Type=exec
User=ts-inventory
ExecStart=/usr/local/bin/ts-inventory
Restart=on-failure
[Install]
WantedBy=multi-user.target

Then create the system account (no home directory, no login shell, as in "Users, groups and the account files"), make the program executable, create its data directory owned by that account, and start the service:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo useradd --system --no-create-home --shell /usr/sbin/nologin ts-inventory sudo chmod 755 /usr/local/bin/ts-inventory sudo install -d -o ts-inventory -g ts-inventory /var/lib/ts-inventory sudo systemctl daemon-reload sudo systemctl start ts-inventory systemctl is-active ts-inventory
active

It ran for days. This morning the data directory was restored from a backup as root and the service was restarted; these commands reproduce what happened:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo rm -r /var/lib/ts-inventory sudo mkdir /var/lib/ts-inventory sudo systemctl restart ts-inventory

Since then the inventory has not updated. The first question is which units have failed, and then what systemd knows about the one that did (-n 4 limits the log excerpt at the end of status to its last four lines):

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl --failed
UNIT LOAD ACTIVE SUB DESCRIPTION ● ts-inventory.service loaded failed failed Inventory sync … 1 loaded units listed.
$ systemctl status ts-inventory -n 4
× ts-inventory.service - Inventory sync Loaded: loaded (/etc/systemd/system/ts-inventory.service; disabled; preset: enabled) Active: failed (Result: exit-code) since Sun 2026-09-27 12:31:14 UTC; 160ms ago Duration: 1ms Invocation: c1ceb6ea648c45f1834d7b3677187e98 Process: 398113 ExecStart=/usr/local/bin/ts-inventory (code=exited, status=2) Main PID: 398113 (code=exited, status=2) Mem peak: 1.7M CPU: 5ms Sep 27 12:31:14 web01 systemd[1]: ts-inventory.service: Scheduled restart job, restart counter is at 5. Sep 27 12:31:14 web01 systemd[1]: ts-inventory.service: Start request repeated too quickly. Sep 27 12:31:14 web01 systemd[1]: ts-inventory.service: Failed with result 'exit-code'. Sep 27 12:31:14 web01 systemd[1]: Failed to start ts-inventory.service - Inventory sync.

status=2 is the program's own exit code. The Process: line says the program exited on its own (code=exited), and the journal lines tell the rest: Restart=on-failure scheduled restarts (the counter reached 5), and after too many starts in a short time (five in ten seconds by default) systemd stopped trying with Start request repeated too quickly and left the unit failed ("Services with systemd" explains the start limit). So the program fails as soon as it starts, and the next layer is its log:

deploy@web01 · Ubuntu 26.04 LTS
$ journalctl -u ts-inventory -n 8 --no-hostname
Sep 27 12:31:14 systemd[1]: Started ts-inventory.service - Inventory sync. Sep 27 12:31:14 ts-inventory[398113]: /usr/local/bin/ts-inventory: 4: cannot create /var/lib/ts-inventory/starts.log: Permission denied Sep 27 12:31:14 systemd[1]: ts-inventory.service: Main process exited, code=exited, status=2/INVALIDARGUMENT Sep 27 12:31:14 systemd[1]: ts-inventory.service: Failed with result 'exit-code'. Sep 27 12:31:14 systemd[1]: ts-inventory.service: Scheduled restart job, restart counter is at 5. Sep 27 12:31:14 systemd[1]: ts-inventory.service: Start request repeated too quickly. Sep 27 12:31:14 systemd[1]: ts-inventory.service: Failed with result 'exit-code'. Sep 27 12:31:14 systemd[1]: Failed to start ts-inventory.service - Inventory sync.
$ journalctl -u ts-inventory -p err -n 1 --no-hostname
Sep 27 12:31:14 systemd[1]: Failed to start ts-inventory.service - Inventory sync.

The second line is the one that matters: line 4 of the script cannot create starts.log, Permission denied. In the third, systemd adds the conventional name for code 2, INVALIDARGUMENT, which tells you nothing about this program. The second command shows a trap. Filtering with -p err keeps only messages marked as errors (here the newest one), and systemd records everything a program prints at the ordinary info level, so the line with the cause disappears. Read a unit's log unfiltered first.

Permission denied means some directory along the path does not let the service's user do what it tried. namei -l lists every component of a path with its owner and mode, and running the same action as the service's user with sudo -u confirms the diagnosis before you change anything. It works for an account without a login shell, because sudo runs the command directly.

deploy@web01 · Ubuntu 26.04 LTS
$ namei -l /var/lib/ts-inventory/starts.log
f: /var/lib/ts-inventory/starts.log drwxr-xr-x root root / drwxr-xr-x root root var drwxr-xr-x root root lib drwxr-xr-x root root ts-inventory starts.log - No such file or directory
$ sudo -u ts-inventory touch /var/lib/ts-inventory/starts.log
touch: cannot touch '/var/lib/ts-inventory/starts.log': Permission denied

ts-inventory is owned by root with mode rwxr-xr-x, so the service's user may enter it but not create files in it. That is the change: the restore re-created the directory as root. (namei exits with status 1 because starts.log does not exist yet.) Give the directory back, restart, and check:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo chown ts-inventory:ts-inventory /var/lib/ts-inventory sudo systemctl restart ts-inventory systemctl is-active ts-inventory
active
$ ls -l /var/lib/ts-inventory
total 4 -rw-r--r-- 1 ts-inventory ts-inventory 21 Sep 27 12:31 starts.log

If you restart within ten seconds of the failures, systemd may still refuse with "start of the service was attempted too often"; sudo systemctl reset-failed ts-inventory clears that. The durable fix is to let systemd own the directory, which is the exercise at the end of this lesson.

A disk that stays full

ts-applog writes its log to /srv/ts-data/app.log, on a filesystem of its own. In the lab that filesystem is a 64 MiB file mounted through a loop device, so filling it never touches the real disk. Its unit says StandardOutput=append:/srv/ts-data/app.log: systemd opens the file once, and the program writes through that open file for as long as it runs. The lab filled the log to stand in for weeks of verbose logging. The symptom is the application complaining that it cannot write:

deploy@web01 · Ubuntu 26.04 LTS
$ df -h /srv/ts-data
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 55M 0 100% /srv/ts-data
$ journalctl -u ts-applog -n 2 --no-hostname
Sep 27 12:31:27 systemd[1]: Started ts-applog.service - Application that logs to /srv/ts-data. Sep 27 12:31:29 ts-applog[398415]: ts-applog: cannot write log: [Errno 28] No space left on device

df shows the filesystem at 100% with nothing available, and the application's error, [Errno 28] No space left on device, is the kernel's answer to a write on a full filesystem. When df -h shows free space and writes still fail with this error, check df -i: the filesystem can also run out of inodes, the records that describe each file. The next question is which files use the space:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo du -sh /srv/ts-data/*
55M /srv/ts-data/app.log 16K /srv/ts-data/lost+found

du adds up the files that have a name, and almost all of it is app.log. A tempting fix is to delete it:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo rm /srv/ts-data/app.log df -h /srv/ts-data
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 55M 0 100% /srv/ts-data

Nothing came back. As "Disks, filesystems and mounts" explained, the kernel frees a file's blocks only when its last name is gone and no process has it open, and ts-applog still does. lsof +L1 lists open files with fewer than one link, meaning deleted but still open. lsof combines its selections with "or" unless you add -a ("and"), so +L1 with a path but without -a lists every deleted open file on the machine (in the lab, also an in-memory file that systemd keeps). With -a it lists only those on this filesystem:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo lsof -a +L1 /srv/ts-data
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME ts-applog 398415 ts-applog 1w REG 7,0 57180160 0 13 /srv/ts-data/app.log (deleted)

That is the culprit: process ts-applog, file descriptor 1 (its standard output) open for writing (1w), about 57 MB, NLINK 0 and (deleted). Restarting the service closes the file:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl restart ts-applog df -h /srv/ts-data
Filesystem Size Used Avail Use% Mounted on /dev/loop0 56M 152K 55M 1% /srv/ts-data

The space is back, and the service has started a new, small app.log. What changed was a log deleted by hand instead of rotated: rotate logs with logrotate, covered in "Logs: the journal and /var/log", or log to the journal, which caps its own size.

A connection that fails

A client on this server must reach an API at http://app01.internal:8080/. In the lab, app01 is a network namespace called ts-app: a separate network stack on the same machine, with its own interface, address (203.0.113.218) and firewall, linked to this server's 203.0.113.217. Commands prefixed with sudo ip netns exec ts-app run inside it; on a real network you would log in to app01 and run the same command. The lab broke three things at once, so each fix uncovers the next message. curl -sS hides the progress meter but still prints errors, and each curl error has a number that points to a layer.

deploy@web01 · Ubuntu 26.04 LTS
$ curl -sS http://app01.internal:8080/
curl: (6) Could not resolve host: app01.internal
$ getent hosts app01.internal

Error 6, Could not resolve host: curl never sent a packet to app01, because the name did not turn into an address. getent hosts asks the same resolver that applications use (/etc/hosts, then DNS, as "DNS and name resolution" showed); it prints nothing and exits with status 2. On a real network the fix is a DNS record. In the lab, an /etc/hosts line does the job:

deploy@web01 · Ubuntu 26.04 LTS
$ echo '203.0.113.218 app01.internal' | sudo tee -a /etc/hosts
203.0.113.218 app01.internal
$ getent hosts app01.internal
203.0.113.218 app01.internal
$ curl -sS --connect-timeout 5 http://app01.internal:8080/
curl: (28) Connection timed out after 5010 milliseconds
$ ping -c 2 app01.internal
PING app01.internal (203.0.113.218) 56(84) bytes of data. 64 bytes from app01.internal (203.0.113.218): icmp_seq=1 ttl=64 time=0.656 ms 64 bytes from app01.internal (203.0.113.218): icmp_seq=2 ttl=64 time=0.194 ms --- app01.internal ping statistics --- 2 packets transmitted, 2 received, 0% packet loss, time 1004ms rtt min/avg/max/mdev = 0.194/0.425/0.656/0.231 ms

sudo tee -a appends to a root-owned file, the pattern from "Pipes, redirection and exit status". The name resolves now, and the error changes to 28: packets went out and nothing came back within the five seconds --connect-timeout allowed. ping gets answers, so the host is up and the route works; only port 8080 goes unanswered, which is the signature of a firewall dropping packets. (Many cloud networks block ping, so a failed ping alone does not prove a host is down.) Look at the firewall on app01:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec ts-app nft list ruleset
table inet ts-filter { chain input { type filter hook input priority filter; policy accept; tcp dport 8080 drop } }
$ sudo ip netns exec ts-app nft delete table inet ts-filter
$ curl -sS http://app01.internal:8080/
curl: (7) Failed to connect to app01.internal port 8080 after 0 ms: Could not connect to server

The rule tcp dport 8080 drop discarded every connection attempt without a reply, which is why curl waited. Writing firewall rules is a subject of the hardening course; this lab table exists only for the demonstration, so it is simply deleted. On a real server, find out why the rule is there first, then change that rule alone, with the firewall's own tool so the change survives a reload. The error changes to 7, Could not connect to server, and at once: app01 answered immediately and refused, which is what a host does when nothing listens on the port. The next layer is the listener and the service behind it:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo ip netns exec ts-app ss -tln
State Recv-Q Send-Q Local Address:Port Peer Address:Port
$ systemctl is-active ts-web
inactive
$ sudo systemctl start ts-web
$ curl -sS http://app01.internal:8080/
inventory API: ok

ss -tln shows no listening sockets inside app01, and the unit is inactive. Once it starts, the request succeeds. On a real server, ask why it was stopped before you start it: its journal may say, and systemctl is-enabled shows whether it will come back after a reboot (sudo systemctl enable --now if it should).

What curl's error tells you
curl http://host:port/ fails
read the error number
(6)
Name does not resolve
getent hosts; fix DNS or /etc/hosts
(28)
Packets dropped on the way
ping, then the firewall rules
(7), fast
Nothing listens on the port
ss -tln, then systemctl status
HTTP error
The application answered
read the application's own log

A checklist to keep

troubleshooting checklist
1. Symptom the exact command and the exact message
2. Scope one service or all? one client or all? this server only? since when?
3. Change apt history.log, journalctl --since, recent edits, restores, reboots
4. Unit systemctl --failed; systemctl status NAME
5. Logs journalctl -u NAME -b, unfiltered first
6. Resources df -h; df -i; free -h; uptime
7. Network ss -tln on the server; curl -sS and ping from the client; getent hosts NAME
8. Access namei -l PATH; the same action with sudo -u SERVICEUSER
9. Config the program's own test (sshd -t, visudo -c) before any reload
10. Fix ask why it was in that state; fix one thing, verify the symptom,
write down what you changed

Try this

Make the first fix permanent. Break the service again with sudo chown root:root /var/lib/ts-inventory, then add a drop-in with sudo systemctl edit ts-inventory containing two lines, [Service] and StateDirectory=ts-inventory. Restart the service and run ls -ld /var/lib/ts-inventory: the directory belongs to ts-inventory again, because with StateDirectory= systemd creates /var/lib/NAME and gives it to the unit's user every time the service starts. Explain in one sentence why this prevents the failure from this lesson's first scenario coming back after the next restore. To remove the practice service afterwards, stop it, delete /etc/systemd/system/ts-inventory.service and its .d directory, /usr/local/bin/ts-inventory and /var/lib/ts-inventory with sudo rm -r, run sudo systemctl daemon-reload, and delete the account with sudo userdel ts-inventory.

Takeaway

Copy the exact error, ask what changed, and check one layer at a time, from the unit outwards, before you change anything. After each change, run the command that showed the symptom again.

Quick check
01The journal of the failing ts-inventory.service shows "Main process exited, code=exited, status=2/INVALIDARGUMENT". What does INVALIDARGUMENT tell you about the cause?
Incorrect — A broken unit file stops the unit from loading or starting. Here the program started and ended with an exit code of its own.
Incorrect — The name is only systemd's conventional label for exit code 2. Many programs, such as the shell here, use 2 for any kind of failure.
Incorrect — Codes from 200 upwards, such as 203/EXEC, are systemd's own setup failures. Code 2 is below that range, so the program ran and chose it.
Correct — Only the program's own messages say why; here the shell's line reported Permission denied.
02Five application servers all started timing out against the same payment API at 10:02, and nothing was changed on any of them. Where do you look first?
Incorrect — All five fail the same way at the same moment, so the cause is unlikely to sit on any one of them.
Correct — A failure that appears everywhere at once points away from the individual servers and towards a dependency they have in common.
Incorrect — The journals will show the same timeout on all five. They describe the symptom, not a cause they share.
Incorrect — Each server has its own rules, and the same new rule on all five at 10:02 would itself be a change, which the question rules out.
03Three things are wrong with a connection at once. Under pressure you add the DNS record, delete the firewall rule and start the service together, and it works. What has that cost you?
Incorrect — The symptom is gone, but you cannot say which change was needed, or whether one of them opened something that should have stayed closed.
Incorrect — It was faster this time. What it lost is information, not time.
Correct — One change at a time, with the symptom checked after each, shows what every change did and which ones to keep or undo.
Incorrect — Nothing in systemd restores deleted nftables rules. The cost is in not knowing which change mattered.

Related