systemd units and the dependency graph
Ordering, requirements, jobs and generators.
An application that starts before the database it needs is ready, a restart that takes a second service with it, a mount unit nobody wrote: all three come from how systemd turns unit files into a dependency graph and the graph into jobs. In this lesson you build three small units and use them to see the difference between requiring a unit and being ordered after it, what "started" means for each service type, how one start request becomes a transaction of jobs, which dependencies every service gets without asking, how stops and crashes travel along the graph, and where generated and socket-activated units come from. The Linux essentials lesson on systemd (unit files, drop-ins, daemon-reload, socket-activated SSH) is assumed.
Unit types and where units come from
Everything systemd manages is a unit, and the suffix of a unit's name gives its type. Counting the units this server has loaded, by type:
Services are the processes systemd supervises. Devices are kernel devices reported by udev, such as disks, which other units can wait for. Targets are named synchronisation points such as multi-user.target that group other units and run nothing themselves. Mounts, automounts and swap describe filesystems and swap space, sockets and timers start services on a connection or at a time, and paths start them when a file appears. Slices are the nodes of the cgroup tree, the subject of the resource-control lesson, and scopes hold processes systemd tracks but did not start itself, such as login sessions. --all includes inactive units and names that other units merely refer to, which is why there are over two hundred services.
systemd reads unit files from these directories, and a file higher in the list replaces a file of the same name further down. /etc/systemd/system (yours) beats /run/systemd/system (runtime, gone at reboot), which beats /usr/lib/systemd/system (packages). system.control holds settings written by systemctl set-property and transient holds units created by systemd-run, both used in the resource-control lesson. The three generator directories hold units that no person wrote.
A generator is a small program that systemd runs early in boot and again at every daemon-reload, before it loads any unit. Each one translates some other configuration into unit files. The best known is systemd-fstab-generator, which turns /etc/fstab into mount units:
systemctl cat names the file it read: /run/systemd/generator/boot.mount, written from the fstab line, with SourcePath= pointing back to it. The generator added dependencies nobody typed: the filesystem check of that device must finish first (Requires= and After= on systemd-fsck@...), and the mount must be done before local-fs.target. Because the unit is generated, an edit to fstab changes nothing in systemd until sudo systemctl daemon-reload runs the generators again; mount prints a hint when you forget.
Other generators read the kernel command line, /etc/crypttab, netplan's network YAML (on Ubuntu), leftover SysV init scripts (systemd-sysv-generator) and, on Ubuntu, the SSH server's port settings (sshd-socket-generator, in the last section). systemd-ssh-generator adds SSH sockets on a virtual machine's vsock device, which is why the lab VM has an sshd-vsock.socket that a physical server does not.
Requiring is not ordering
Unit files express two independent relations. Requirement directives say which units a start pulls in: Wants= pulls one in and ignores its failure, Requires= pulls one in and takes its failure seriously. Ordering directives, After= and Before=, say which of two units that are both starting waits for the other. Neither implies the other; without ordering, systemd.unit(5) says, both units start at the same time. Two lab units make that visible. Create each file below as root, for example with sudoedit, and run sudo systemctl daemon-reload after writing them. The first stands in for a database that needs five seconds before it can take connections.
[Unit]Description=Lab database (ready after five seconds)[Service]Type=notifyNotifyAccess=allRuntimeDirectory=sd-units-dbExecStart=/bin/sh -c 'sleep 5; touch /run/sd-units-db/ready; systemd-notify --ready; exec sleep infinity'
RuntimeDirectory= creates /run/sd-units-db when the service starts and removes it when it stops, so the ready file exists only while the database is up and ready. The application refuses to run without that file. It requires the database and has no ordering line.
[Unit]Description=Lab application that needs the databaseRequires=sd-units-db.service[Service]Type=execExecStart=/bin/sh -c 'test -e /run/sd-units-db/ready || { echo "database not ready"; exit 1; }; echo "connected to the database"; exec sleep infinity'
The start reported success, and a moment later the application had failed while the database was still activating. Both units were started at the same moment and the application lost the race: database not ready. (systemctl is-active exits 0 when any of the listed units is active, 3 when none is.) static in the Loaded line means the unit has no [Install] section, so it runs only when started by hand or pulled in by another unit. Add the ordering in a drop-in, stop the database so that both units start from nothing (the application has already failed), and try again:
This time the start took five seconds. systemd started the database, waited until it reported that it was ready, and only then started the application, which connected. With After= in place, Requires= is also strict about failure: if the database fails to start, the application is not started at all and its job fails with "Dependency failed". Almost every real dependency needs both lines, and both go in the unit that waits.
What "started" means: Type= and readiness
After= waits until the other unit has finished starting, and that unit's Type= decides when that is. The database uses Type=notify: it counts as started when it sends READY=1 to systemd, which systemd-notify --ready does after the five seconds. (NotifyAccess=all lets a child process of the service send the message; the default accepts it only from the main process.) The other types, from systemd.service(5): simple, the default when a unit has ExecStart= and no Type=, counts as started as soon as systemd has forked the process, before the program has even been executed; exec once the program has been executed, so a missing binary fails the start; forking when the process systemd launched exits and leaves its child running, the traditional Unix daemon; oneshot when the process has finished; dbus when the program has taken its bus name. Switch the database to simple and see what After= is worth then:
The database was active at once, so the application started at once and failed just as it did without After=. systemctl revert deleted the drop-in and returned the unit to what its own file says. The database was still running, so the last command stops it and clears the application's failed state, and both units are inactive again. Only notify, dbus and a well-written forking daemon tell systemd that a service is ready to serve. If your program supports sd_notify, use Type=notify; if not, use Type=exec and make its clients retry, or let systemd own the socket, as the last section shows.
Jobs and transactions
systemctl start does not start anything itself. It asks the manager for a job, and the manager builds a transaction: the requested job plus a job for every unit the requirement dependencies pull in, checked for contradictions and ordering loops before any of it runs. Each job then waits for the jobs it is ordered after. With both units stopped, start the application without waiting for the result:
One request, two jobs. The database's start job is running and the application's is waiting for it, because of After=. A request that would reverse a queued job is refused if you ask for that:
--job-mode=fail makes systemd refuse any transaction that would cancel a job already queued; this stop would have cancelled the queued start, so the transaction is "destructive". The default mode, replace, would have replaced the start with the stop. Once the jobs finish, look at the dependencies the manager actually holds for the application:
The file names one dependency; the rest are default dependencies, which every service gets unless it sets DefaultDependencies=no (systemd.service(5)). It requires and is ordered after sysinit.target, so early boot (local filesystems, udev, the journal) is finished; it is ordered after basic.target; and it conflicts with, and is ordered before, shutdown.target. That conflict is how shutdown works: starting shutdown.target stops every unit that conflicts with it, in the reverse of the start order. system.slice is the cgroup the service runs in, and systemd-journald.socket is there because the service's output goes to the journal. Only units that run very early or very late, such as filesystem checks, turn default dependencies off.
systemctl list-dependencies draws the requirement tree without ordering, expanding target units recursively (--all expands every unit): the database, the slice, and under sysinit.target everything early boot involves. --reverse answers the opposite question, which units pull this one in. systemd-analyze dot prints both kinds of edge in the format of the Graphviz dot program, which can draw them as a picture, with a colour per dependency type:
Stops, restarts and crashes
Requirements also carry stops and restarts. An explicit stop or restart of a unit is passed to every unit that Requires= it: stopping the database stops the application. PartOf= does only that and pulls nothing in on start. It is the directive most often written the wrong way round: PartOf=X in unit Y means that when X is stopped or restarted, Y is too, and nothing that happens to Y affects X. A worker that belongs to the application:
[Unit]Description=Lab worker that belongs to the applicationPartOf=sd-units-app.serviceAfter=sd-units-app.service[Service]Type=execExecStart=/usr/bin/sleep infinity
systemd records the relation on the other side as well: the application lists the worker under ConsistsOf=, a property you cannot set directly. Restarting the application restarted the worker, which has a new main PID; stopping the worker left the application alone. What Requires= does not pass on is a unit that stops by itself. Kill the database's processes, as a crash would:
The database is failed and the application runs on without it. systemd.unit(5) states this: a unit that deactivates on its own, such as a service whose process exits, is not propagated to units that require it. BindsTo= is the stronger form that also stops the dependent unit when its dependency dies, and it is right for a unit that must never run without the other. An application that reconnects by itself is better left running, with Restart= on the database bringing it back.
Socket activation
The essentials lesson showed that Ubuntu starts SSH through ssh.socket. The mechanism also explains why socket-activated services need little ordering. A socket unit makes systemd create the listening socket itself:
Each listening socket is held by two processes: systemd (PID 1), which created it, and sshd, which received it as an open file descriptor when ssh.service started. Accept=no in the socket unit means one service instance receives the listening socket and accepts all connections itself. For a listening socket, Send-Q shows the backlog limit, 4096: a client that connects while the service is still starting waits in that queue instead of being refused.
Socket units are started by sockets.target, basic.target is ordered after sockets.target, and every ordinary service is ordered after basic.target by its default dependencies. So every socket-activated server is already listening before any ordinary service starts, and its clients need no After= on it. On Ubuntu, sshd-socket-generator copies a non-default Port or ListenAddress from the SSH server configuration into a drop-in for ssh.socket, so a port change needs sudo systemctl daemon-reload and sudo systemctl restart ssh.socket (Ubuntu's README.Debian for openssh-server says so). RHEL 10 runs sshd.service as an ordinary service that opens its own socket.
Try this
Make the application stop when the database dies. Add a second drop-in without touching the first: printf '[Unit]\nBindsTo=sd-units-db.service\n' | sudo systemctl edit --drop-in=bindsto --stdin sd-units-app, then sudo systemctl start sd-units-app, which also starts the database again. Kill the database with sudo systemctl kill --signal=KILL sd-units-db and predict systemctl is-active sd-units-db sd-units-app before you run it: expect failed and inactive. Clean up with sudo systemctl stop sd-units-worker sd-units-app and sudo systemctl reset-failed sd-units-db, delete the three unit files and the sd-units-app.service.d directory from /etc/systemd/system, and run sudo systemctl daemon-reload.
Takeaway
Write every dependency as a requirement plus an ordering in the unit that waits, and make sure the unit it waits for has a Type= that means ready. When a start, stop or crash travels the graph in a way you did not expect, systemctl show -p Requires -p After -p ConsistsOf shows the dependencies systemd is actually using.