Services with systemd
Read state, write a unit, debug a failure.
A service is a program that runs in the background without anyone logged in to start it: the SSH server that lets you log in, the time synchronisation daemon, a web server, your own application. On Ubuntu 26.04 and RHEL 10, as on nearly every current distribution, services are started and supervised by systemd. The kernel starts systemd first, as process ID 1, and systemd then starts services at boot in the right order, restarts them when they crash, collects their output in the journal, and stops them cleanly at shutdown. You give it instructions with one command, systemctl.
PID 1 belongs to systemd and runs as root. Every other userspace process descends from it (kernel threads hang from PID 2, kthreadd, as the first lesson showed). systemd also puts each service's processes into its own control group, or cgroup, a kernel grouping of processes. That is how it can find and stop everything a service started, even a process that detached from its parent.
Reading a service's state
Start with a service that is certainly running: the SSH server you are probably logged in through. On Ubuntu it is called ssh.
Read it from the top. Loaded names the unit file systemd read, here /usr/lib/systemd/system/ssh.service, the directory where packages install their units. The same bracket says whether the unit is enabled to start at boot and what the distribution's preset would choose. Active is the state right now and how long it has been in it. Process shows a command that ran before the main one: sshd -t checks the configuration, so a broken sshd_config stops the start instead of taking the server down. Main PID is the process to look for in ps, and CGroup lists every process systemd counts as part of the service. The last lines come from the journal, systemd's log, so you see why the service did what it did without opening a log file.
One line needs explaining. The unit says disabled, yet it is running and TriggeredBy names ssh.socket. Since Ubuntu 22.10 the SSH server is socket-activated: ssh.socket is the unit that is enabled and listens on port 22, and systemd starts ssh.service when the first connection arrives. The two units answer the two questions differently.
To see every service running right now, list the service units in the running state. The list below leaves some lines out, among them the agent of the virtual machine tool the lab runs in.
Each line is one unit: its name, whether its file loaded, and its high-level and low-level state. This is the quickest way to learn what a server you have just been handed actually runs.
sshd.service, it is enabled directly, and sshd.socket exists but is disabled, so the daemon starts at boot and listens itself. Scripts that restart SSH on both families need the right name: systemctl restart ssh on Ubuntu, systemctl restart sshd on RHEL.Running now versus starting at boot
systemd keeps two questions apart. Is the service running this second? start, stop and restart change that, reload asks a running service to re-read its configuration, and systemctl is-active reports it. Will it start at the next boot? enable and disable change that, and systemctl is-enabled reports it. Neither set of commands touches the other question. sudo systemctl enable --now NAME does both in one step, which is what you want when you deploy something.
Writing a unit for your own program
Unit files live in two main places. /usr/lib/systemd/system belongs to packages: a package upgrade can replace any file there, so you never edit them. /etc/systemd/system belongs to you, and a unit there with the same name as a packaged one takes its place. Between the two sits /run/systemd/system, for units created at runtime that disappear at reboot, so the order of precedence is /etc, then /run, then /usr/lib. Your own services go in /etc/systemd/system.
As a stand-in for an application, the example runs Python's built-in web server, serving one page from /srv/myapp as a system account called myapp that cannot log in. Create the account and the page first:
Then write the unit file as root, for example with sudoedit /etc/systemd/system/myapp.service:
[Unit]Description=Demo web app (Python http.server)[Service]Type=execUser=myappExecStart=/usr/bin/python3 -m http.server 8080 --bind 127.0.0.1 --directory /srv/myappRestart=on-failure[Install]WantedBy=multi-user.target
[Unit] describes the unit; ordering lines such as After= would go here too, and this app needs none because it only listens on the local address. [Service] says how to run it. ExecStart is the command, with an absolute path. User=myapp runs it without root privileges. Restart=on-failure restarts it if it exits with an error. Type=exec makes systemctl start wait until the program has actually been executed and report an error if it could not be; the default, Type=simple, reports success as soon as systemd has forked, even when the program does not exist. [Install] is read only by enable: WantedBy=multi-user.target attaches the service to the normal boot.
The service answers, and the status line reads disabled: it is running, but nothing will start it after a reboot. That combination is the usual cause of "it worked until we patched and rebooted". systemd found the new file without being told because it loads a unit it has never seen the first time you name it. Changing a unit it has already loaded is different, as the next section shows.
After=network.target. That only orders the service after the network management service has started; it does not wait for an address or a route. A service that genuinely needs a working network at startup uses Wants=network-online.target and After=network-online.target. Most services, including anything that only listens, need neither.When a service fails
Now break it the way people really do: a typo in the program path. After editing the file, a plain restart looks like it worked.
systemd warns and restarts the old definition it still holds in memory: the service is running the command from before the edit. After you change a unit file by hand, run sudo systemctl daemon-reload so systemd reads it again. This time the restart fails.
Follow the advice in the message and start with systemctl status.
The cross and failed say the service is down. The Process line shows the exact command systemd tried and status=203/EXEC. Codes from 200 upwards are systemd's own and mean it failed while setting the process up, before your program ran: 203 means the program could not be executed at all (systemd.exec(5) lists them). The journal lines show that Restart=on-failure tried again, then gave up: systemd refuses to start a unit more than five times in ten seconds by default, so a broken service cannot loop forever.
journalctl -u myapp shows only this unit's messages, and the first line spells out the reason: Failed at step EXEC spawning /usr/bin/pyhton3: No such file or directory, so the executable does not exist. The next line is the 203/EXEC status again, and after the restart attempts (left out) comes Start request repeated too quickly, the rate limit giving up. When you do not know which service is in trouble, systemctl --failed lists every unit in the failed state. It is a good first command on a machine that misbehaves after a reboot.
The lines left out are per-user service managers (user@UID.service) of throwaway accounts that other lab exercises deleted while those accounts were still logged in; on your own machine the list should hold only what is really broken.
Fix the path, reload, restart, and check the result instead of assuming it.
If you retry within ten seconds of the failures you may instead see "start of the service was attempted too often". sudo systemctl reset-failed myapp clears the counter.
Enabling it for boot
Enabling is nothing more than that symbolic link: multi-user.target pulls in every unit linked from its .wants directory when the machine boots. disable removes the link and leaves the running service alone.
A disabled unit can still be started by hand or pulled in by another unit. When a service must not run at all, mask it: systemd puts a link to /dev/null in /etc/systemd/system under the unit's name, which hides the packaged file, and every attempt to start it fails until you unmask it. Here it is on the rsync daemon, a packaged unit that is installed but not running.
Because masking works by occupying the unit's name in /etc/systemd/system, it cannot mask a unit whose own file already lives there: sudo systemctl mask myapp answers Failed to mask unit: File '/etc/systemd/system/myapp.service' already exists. For your own units, disable them or remove the file instead.
Changing a unit without editing it
To change a packaged unit, or one a colleague owns, add a drop-in instead of editing the file. sudo systemctl edit myapp opens an editor on /etc/systemd/system/myapp.service.d/override.conf and reloads systemd when you save; in a script, --stdin reads the new content from a pipe. Here the drop-in moves the app to port 9090. The empty ExecStart= line first clears the command inherited from the main file; without it systemd would refuse the unit for having two commands.
systemctl cat shows every file that makes up the unit, in the order they apply, which is the quickest way to find out why a service is not behaving the way its main file says it should.
Service settings go much further than this: the hardening course uses the same drop-ins to take privileges and filesystem access away from a service. The journal that journalctl -u reads is the subject of the next lesson.
Try this
Build myapp from this lesson on a lab machine, enable it, and confirm it answers with curl. Then add a second drop-in, separate from override.conf: sudo systemctl edit --drop-in=zz-user myapp opens /etc/systemd/system/myapp.service.d/zz-user.conf; put [Service] and User=nosuchuser in it, save, and restart. Find the reason in systemctl status myapp and journalctl -u myapp: the status line reads status=217/USER, systemd's code for a user it could not set. Remove that file with sudo rm, run sudo systemctl daemon-reload and sudo systemctl reset-failed myapp, and start it again. Finish with sudo systemctl disable --now myapp, so is-active and is-enabled both report it off, then remove everything: sudo rm -r /etc/systemd/system/myapp.service /etc/systemd/system/myapp.service.d /srv/myapp, sudo systemctl daemon-reload and sudo userdel myapp.
Takeaway
systemctl status NAME and journalctl -u NAME answer most questions about a service. Keep your own units in /etc/systemd/system, change packaged ones with drop-ins, and run daemon-reload after any edit you make by hand.