Sandboxing services with systemd
Users, capabilities, seccomp and read-only views.
A service that is compromised runs with whatever the service was given. If that is root and the whole filesystem, a bug in one daemon becomes control of the machine. systemd can take almost all of that away from a service without changing its code: run it as a throwaway user, make the filesystem read-only, hand it one capability instead of root, and allow only the system calls it needs. In this lesson you build a small service, measure how exposed it is, then harden it one drop-in at a time, proving at each step that the restriction actually bites, watching the exposure score fall, and rolling the whole thing back. The Linux essentials course covered writing and changing a unit; this lesson assumes that and adds the security directives.
Measure before you change anything
systemd-analyze security scores a unit from 0 (locked down) to 10 (wide open), listing every setting it checks. Start with a stand-in application: a small Python HTTP server that answers with its effective user ID and, on /write?path=..., tries to create a file where the request says, so the sandbox is visible from a browser. Save it as /usr/local/lib/hard-sandbox/app.py:
#!/usr/bin/env python3# A stand-in for an application: it answers on a port and, on /write?path=..., tries to# create a file where the query says. That makes the sandbox visible from a browser.import osfrom http.server import BaseHTTPRequestHandler, HTTPServerfrom urllib.parse import urlparse, parse_qsclass H(BaseHTTPRequestHandler):def do_GET(self):u = urlparse(self.path)if u.path == "/write":target = parse_qs(u.query).get("path", ["/var/lib/hard-sandbox/marker"])[0]try:with open(target, "w") as f:f.write("ok\n")body = "wrote " + target + "\n"except OSError as e:body = "FAILED " + target + ": " + e.strerror + "\n"else:body = "hello, euid=%d\n" % os.geteuid()data = body.encode()self.send_response(200)self.send_header("Content-Type", "text/plain")self.send_header("Content-Length", str(len(data)))self.end_headers()self.wfile.write(data)def log_message(self, *a):returnHTTPServer(("127.0.0.1", int(os.environ.get("PORT", "80"))), H).serve_forever()
The unit runs it with no restrictions, as most packaged units still do. (The lab VM's tool lets any user bind low ports, net.ipv4.ip_unprivileged_port_start=0; the lab sets the kernel default, 1024, as on a real Ubuntu or RHEL server.)
[Unit]Description=Demo API (hard-sandbox)[Service]Environment=PORT=80ExecStart=/usr/bin/python3 /usr/local/lib/hard-sandbox/app.pyRestart=on-failure[Install]WantedBy=multi-user.target
The service answers as euid=0, writes wherever it likes, and scores 9.6 UNSAFE. A flaw in this app is a root flaw with full write access to the disk. Everything below drives that number down.
One rule about the files first
Unit files take comments only on their own line. A # after a value becomes part of the value, and the setting is dropped with only a warning in the journal, while the unit still starts. It is worth seeing once, because it is the most common way a "hardened" unit turns out not to be. systemctl edit --stdin writes a drop-in from standard input (without --stdin it opens an editor) and reloads systemd itself. Add a drop-in with a trailing comment and ask systemd to check it:
systemd-analyze verify reports the parse failure, and systemctl show confirms the result: NoNewPrivileges=no, the setting ignored entirely. The service you thought you hardened runs exactly as before. Every snippet in this lesson keeps its comments on their own lines.
An identity and a smaller world
The first real change gives the service its own identity and takes away the filesystem. DynamicUser=yes runs it as a system user that systemd creates on start and removes on stop, so there is no standing account to target and nothing it owns to abuse. ProtectSystem=strict mounts the entire filesystem read-only for the service, StateDirectory= carves out one writable directory under /var/lib, and ProtectHome=, PrivateTmp=, PrivateDevices= and the ProtectKernel* settings hide the users' homes, give it a private /tmp, and cut it off from device nodes and kernel tunables. ProtectProc=invisible gives the service its own view of /proc in which other users' processes do not exist, so a compromised service cannot read their command lines or look for targets; systemd.exec(5) recommends it for most services, and it needs a non-root user, which DynamicUser= provides.
[Service]# Run under a throwaway system user that systemd allocates for each start.DynamicUser=yes# The only writable place: /var/lib/hard-sandbox, created and owned for us.StateDirectory=hard-sandbox# Comments go on their own line: systemd ignores a whole-line #, not a trailing one.NoNewPrivileges=yesProtectSystem=strictProtectHome=yes# Hide other users' processes from the service's view of /proc.ProtectProc=invisiblePrivateTmp=yesPrivateDevices=yesProtectKernelTunables=yesProtectKernelModules=yesProtectKernelLogs=yesProtectControlGroups=yes
Restart with that drop-in and the service does not come back. It now runs as a non-root user, and the app binds port 80, which is a privileged port: only root, or a process holding the right capability, may bind a port below 1024. The journal names the exact failure.
A capability instead of root
The reflex is to put the service back to root. The better fix is to grant the one privilege it needs and nothing else. Linux splits root's power into capabilities, and binding a low port is CAP_NET_BIND_SERVICE. AmbientCapabilities= gives the running process that one capability even though it is not root, and CapabilityBoundingSet= caps the set it could ever hold at just that one, so even a successful exploit cannot pick up another.
[Service]# Let the service bind a port below 1024 without being root, and allow nothing else.AmbientCapabilities=CAP_NET_BIND_SERVICECapabilityBoundingSet=CAP_NET_BIND_SERVICE
Now it serves on port 80 as a dynamic, non-root user. User=hard-sandbox is the transient account systemd made; it does not exist before the service starts or after it stops, so there is no login, no home and no standing target. This is capabilities doing the job people reach for root to do. With the service running confined, check that the filesystem restrictions bite. Writing to its state directory works; writing to /etc fails because the filesystem is read-only under it; writing into a home directory fails because ProtectHome hid it.
Three requests, three different outcomes: the carve-out is writable, /etc reports "Read-only file system" (ProtectSystem=strict), and /home reports "Permission denied" (ProtectHome). The exposure score has already dropped from 9.6 UNSAFE to 4.6 OK, the middle of the scale.
ProtectProc=invisible is easiest to see with two throwaway services started by systemd-run, one without it and one with it, each counting the process directories it can see:
Without it the throwaway service sees 131 process directories, every process on the host. With it, 3: its own shell, ls and wc; ps -e then lists only itself and the shell. Note what it does not do: it hides other processes from the service, not the service's own command line from other users; that needs hidepid= on the host's /proc (the secrets lesson).
Seccomp and the last knobs
The final layer restricts the system calls themselves. SystemCallFilter=@system-service allows only the calls a normal service makes and blocks the rest with a seccomp filter (a kernel feature that screens system calls). SystemCallArchitectures=native stops the process from smuggling calls in through a foreign CPU ABI. RestrictAddressFamilies= limits it to IP and Unix sockets, RestrictNamespaces= stops it creating namespaces, and MemoryDenyWriteExecute= refuses memory that is writable and executable at once. That last one breaks every just-in-time compiler (systemd.exec(5) says so): leave it out for Java, Node.js, .NET or a PCRE JIT, and keep it for plain interpreted or compiled code like this Python app. Resource caps such as TasksMax= and LimitNOFILE= are capacity settings, not hardening; size them from the service's observed peak, or a busy service fails with "Too many open files".
[Service]# Allow the system-service syscall set only, on this CPU's native ABI, and make a# blocked call return EPERM instead of killing the process. Sockets: IP and Unix only.SystemCallFilter=@system-serviceSystemCallErrorNumber=EPERMSystemCallArchitectures=nativeRestrictAddressFamilies=AF_INET AF_INET6 AF_UNIXRestrictNamespaces=yesLockPersonality=yesMemoryDenyWriteExecute=yes
The service still serves, because its normal work uses only allowed calls, and the score falls again, to 1.8 OK.
Before you enforce a filter on a real service, find out what it calls. SystemCallLog= logs the listed calls through the audit subsystem without blocking them; run the service for a representative period with SystemCallLog=~@system-service (every call outside the set) and read the log. Here a throwaway unit logs @chown calls:
The call was allowed (the file now belongs to deploy), and the audit log names it: syscall=fchownat from the uutils chown, with arch=aarch64 (an x86_64 server logs its own architecture and call names may differ). With auditd running, as here, the record lands in audit.log; without auditd the kernel writes it to the kernel log. It is worth seeing how the filter behaves when a call is blocked. By default systemd kills the process with SIGSYS; with SystemCallErrorNumber=EPERM the call instead returns an error and the program carries on. A transient systemd-run unit that denies the @chown group and then tries to change a file's owner shows both.
The first is killed (status=31/SYS is SIGSYS); the second is refused with "Operation not permitted" and keeps running. SystemCallErrorNumber=EPERM is usually kinder to a real service, which can handle an error but not a sudden death mid-request. The price is diagnosis: an EPERM return leaves no record naming the blocked call, so the service reports an odd error far from the cause. That is one more reason to run the SystemCallLog= discovery first. Blocking too much is the common self-inflicted outage here, which is why @system-service (an allow-list systemd maintains) is the sane default rather than a hand-picked list.
ReadWritePaths or a LogsDirectory, one that runs helper programs cannot take a strict syscall filter, and one that talks to a Unix socket needs AF_UNIX left in RestrictAddressFamilies. Read the failure in journalctl -u NAME and loosen the one directive that caused it, rather than reaching for User=root again. And never sandbox a process that starts other people's work: cron, atd, sshd, getty, systemd-logind or a CI runner pass their restrictions on to every job and login they start. A read-only /etc inherited by every root cron job, or ProtectHome= hiding every authorized_keys from sshd, is an outage. Sandbox the job's own unit instead, as the scheduled jobs lesson does with its timer.Verify, then roll it back
Confirm the live settings with systemctl show, which reports what is actually in effect, not what the files say. Then remember the escape hatch: every drop-in is removable. systemctl revert deletes them all and returns the unit to its packaged definition, which is the rollback for any of these changes.
Reverted, the service is back to euid=0 and 9.6 UNSAFE, proving the drop-ins were the only thing holding it down. On a packaged daemon you do the same through systemctl edit, measuring before and after and restarting to test. The example is nginx (sudo apt install nginx), a leaf daemon: it starts only its own worker processes. The drop-in lets it write only its logs, its cache and its pid file:
nginx restarted with the drop-in, answered HTTP 200, and moved from 9.6 UNSAFE to 7.4 MEDIUM; the revert removed the drop-in and the score went back to 9.6. More directives (a system call filter, a bounding set of the few capabilities its master process needs) would lower it further, each tested the same way. Not every daemon can take every restriction, which is why you measure, apply what fits, restart, and test that the service still does its job before you keep a drop-in.
Try this
Practise on the lesson's demo service, on nginx, or on another leaf service you can afford to break on a lab machine; never on ssh, networking, journald, dbus or logind, which can cut you off. Run systemd-analyze security NAME to read its score, and add a drop-in with sudo systemctl edit NAME containing ProtectSystem=strict, ProtectHome=yes, NoNewPrivileges=yes and PrivateTmp=yes (comments on their own lines). Restart it, confirm it still works, and read the new score. Then break something on purpose: add SystemCallFilter=@system-service and SystemCallArchitectures=native, restart, and if the service fails, read journalctl -u NAME for the killed call. Finish with sudo systemctl revert NAME and confirm systemctl show -p DropInPaths NAME is empty and the service is back to normal.
Takeaway
Measure a service with systemd-analyze security, then take away what it does not need: a dynamic user instead of root, a read-only filesystem with one writable StateDirectory, a single capability such as CAP_NET_BIND_SERVICE in place of root, and the @system-service syscall allow-list. Verify with systemctl show, keep comments on their own lines, and roll back with systemctl revert.