Sandboxing services with systemd

Users, capabilities, seccomp and read-only views.

Intermediate18 min · lesson 13 of 24

A service that is compromised runs with whatever the service was given. If that is root and the whole filesystem, a bug in one daemon becomes control of the machine. systemd can take almost all of that away from a service without changing its code: run it as a throwaway user, make the filesystem read-only, hand it one capability instead of root, and allow only the system calls it needs. In this lesson you build a small service, measure how exposed it is, then harden it one drop-in at a time, proving at each step that the restriction actually bites, watching the exposure score fall, and rolling the whole thing back. The Linux essentials course covered writing and changing a unit; this lesson assumes that and adds the security directives.

Measure before you change anything

systemd-analyze security scores a unit from 0 (locked down) to 10 (wide open), listing every setting it checks. Start with a stand-in application: a small Python HTTP server that answers with its effective user ID and, on /write?path=..., tries to create a file where the request says, so the sandbox is visible from a browser. Save it as /usr/local/lib/hard-sandbox/app.py:

/usr/local/lib/hard-sandbox/app.py
#!/usr/bin/env python3
# A stand-in for an application: it answers on a port and, on /write?path=..., tries to
# create a file where the query says. That makes the sandbox visible from a browser.
import os
from http.server import BaseHTTPRequestHandler, HTTPServer
from urllib.parse import urlparse, parse_qs
class H(BaseHTTPRequestHandler):
def do_GET(self):
u = urlparse(self.path)
if u.path == "/write":
target = parse_qs(u.query).get("path", ["/var/lib/hard-sandbox/marker"])[0]
try:
with open(target, "w") as f:
f.write("ok\n")
body = "wrote " + target + "\n"
except OSError as e:
body = "FAILED " + target + ": " + e.strerror + "\n"
else:
body = "hello, euid=%d\n" % os.geteuid()
data = body.encode()
self.send_response(200)
self.send_header("Content-Type", "text/plain")
self.send_header("Content-Length", str(len(data)))
self.end_headers()
self.wfile.write(data)
def log_message(self, *a):
return
HTTPServer(("127.0.0.1", int(os.environ.get("PORT", "80"))), H).serve_forever()

The unit runs it with no restrictions, as most packaged units still do. (The lab VM's tool lets any user bind low ports, net.ipv4.ip_unprivileged_port_start=0; the lab sets the kernel default, 1024, as on a real Ubuntu or RHEL server.)

/etc/systemd/system/hard-sandbox.service
[Unit]
Description=Demo API (hard-sandbox)
[Service]
Environment=PORT=80
ExecStart=/usr/bin/python3 /usr/local/lib/hard-sandbox/app.py
Restart=on-failure
[Install]
WantedBy=multi-user.target
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl start hard-sandbox
$ curl -s http://127.0.0.1/
hello, euid=0
$ curl -s "http://127.0.0.1/write?path=/var/tmp/hard-sandbox-probe" ls -l /var/tmp/hard-sandbox-probe
wrote /var/tmp/hard-sandbox-probe …
$ systemd-analyze security hard-sandbox.service | tail -n 1
→ Overall exposure level for hard-sandbox.service: 9.6 UNSAFE :-{

The service answers as euid=0, writes wherever it likes, and scores 9.6 UNSAFE. A flaw in this app is a root flaw with full write access to the disk. Everything below drives that number down.

One rule about the files first

Unit files take comments only on their own line. A # after a value becomes part of the value, and the setting is dropped with only a warning in the journal, while the unit still starts. It is worth seeing once, because it is the most common way a "hardened" unit turns out not to be. systemctl edit --stdin writes a drop-in from standard input (without --stdin it opens an editor) and reloads systemd itself. Add a drop-in with a trailing comment and ask systemd to check it:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl edit --stdin --drop-in=05-badcomment hard-sandbox <<'EOF' [Service] NoNewPrivileges=true # belt and braces EOF
Successfully installed edited file '/etc/systemd/system/hard-sandbox.service.d/05-badcomment.conf'.
$ sudo systemd-analyze verify hard-sandbox.service
/etc/systemd/system/hard-sandbox.service.d/05-badcomment.conf:2: Failed to parse NoNewPrivileges=true # belt and braces, ignoring: Invalid argument …
$ sudo systemctl daemon-reload systemctl show -p NoNewPrivileges hard-sandbox.service
NoNewPrivileges=no

systemd-analyze verify reports the parse failure, and systemctl show confirms the result: NoNewPrivileges=no, the setting ignored entirely. The service you thought you hardened runs exactly as before. Every snippet in this lesson keeps its comments on their own lines.

An identity and a smaller world

The first real change gives the service its own identity and takes away the filesystem. DynamicUser=yes runs it as a system user that systemd creates on start and removes on stop, so there is no standing account to target and nothing it owns to abuse. ProtectSystem=strict mounts the entire filesystem read-only for the service, StateDirectory= carves out one writable directory under /var/lib, and ProtectHome=, PrivateTmp=, PrivateDevices= and the ProtectKernel* settings hide the users' homes, give it a private /tmp, and cut it off from device nodes and kernel tunables. ProtectProc=invisible gives the service its own view of /proc in which other users' processes do not exist, so a compromised service cannot read their command lines or look for targets; systemd.exec(5) recommends it for most services, and it needs a non-root user, which DynamicUser= provides.

/etc/systemd/system/hard-sandbox.service.d/10-confine.conf
[Service]
# Run under a throwaway system user that systemd allocates for each start.
DynamicUser=yes
# The only writable place: /var/lib/hard-sandbox, created and owned for us.
StateDirectory=hard-sandbox
# Comments go on their own line: systemd ignores a whole-line #, not a trailing one.
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
# Hide other users' processes from the service's view of /proc.
ProtectProc=invisible
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes

Restart with that drop-in and the service does not come back. It now runs as a non-root user, and the app binds port 80, which is a privileged port: only root, or a process holding the right capability, may bind a port below 1024. The journal names the exact failure.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl restart hard-sandbox
$ systemctl status hard-sandbox --lines 0
× hard-sandbox.service - Demo API (hard-sandbox) … Active: failed (Result: exit-code) since Sun 2026-09-27 09:28:57 UTC; 2s ago …
$ journalctl -u hard-sandbox --no-hostname | grep -iE "permission denied|errno" | tail -n 1
Sep 27 09:28:57 python3[243121]: PermissionError: [Errno 13] Permission denied

A capability instead of root

The reflex is to put the service back to root. The better fix is to grant the one privilege it needs and nothing else. Linux splits root's power into capabilities, and binding a low port is CAP_NET_BIND_SERVICE. AmbientCapabilities= gives the running process that one capability even though it is not root, and CapabilityBoundingSet= caps the set it could ever hold at just that one, so even a successful exploit cannot pick up another.

/etc/systemd/system/hard-sandbox.service.d/20-caps.conf
[Service]
# Let the service bind a port below 1024 without being root, and allow nothing else.
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl reset-failed hard-sandbox sudo systemctl restart hard-sandbox
$ curl -s http://127.0.0.1/
hello, euid=61325
$ systemctl show -p User -p DynamicUser hard-sandbox.service
User=hard-sandbox DynamicUser=yes

Now it serves on port 80 as a dynamic, non-root user. User=hard-sandbox is the transient account systemd made; it does not exist before the service starts or after it stops, so there is no login, no home and no standing target. This is capabilities doing the job people reach for root to do. With the service running confined, check that the filesystem restrictions bite. Writing to its state directory works; writing to /etc fails because the filesystem is read-only under it; writing into a home directory fails because ProtectHome hid it.

deploy@web01 · Ubuntu 26.04 LTS
$ curl -s "http://127.0.0.1/write?path=/var/lib/hard-sandbox/marker"
wrote /var/lib/hard-sandbox/marker
$ curl -s "http://127.0.0.1/write?path=/etc/hard-sandbox-probe"
FAILED /etc/hard-sandbox-probe: Read-only file system
$ curl -s "http://127.0.0.1/write?path=/home/deploy/hard-sandbox-probe"
FAILED /home/deploy/hard-sandbox-probe: Permission denied

Three requests, three different outcomes: the carve-out is writable, /etc reports "Read-only file system" (ProtectSystem=strict), and /home reports "Permission denied" (ProtectHome). The exposure score has already dropped from 9.6 UNSAFE to 4.6 OK, the middle of the scale.

deploy@web01 · Ubuntu 26.04 LTS
$ systemd-analyze security hard-sandbox.service | tail -n 1
→ Overall exposure level for hard-sandbox.service: 4.6 OK :-)

ProtectProc=invisible is easiest to see with two throwaway services started by systemd-run, one without it and one with it, each counting the process directories it can see:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --pipe --wait --quiet -p DynamicUser=yes sh -c 'ls -d /proc/[0-9]* | wc -l'
131
$ sudo systemd-run --pipe --wait --quiet -p DynamicUser=yes -p ProtectProc=invisible sh -c 'ls -d /proc/[0-9]* | wc -l; ps -e'
3 PID TTY TIME CMD 243496 ? 00:00:00 sh 243499 ? 00:00:00 ps

Without it the throwaway service sees 131 process directories, every process on the host. With it, 3: its own shell, ls and wc; ps -e then lists only itself and the shell. Note what it does not do: it hides other processes from the service, not the service's own command line from other users; that needs hidepid= on the host's /proc (the secrets lesson).

What each directive takes away
Identity
DynamicUser=
no standing account
NoNewPrivileges=
no setuid escalation
Capability*Set=
one privilege, not root
Filesystem
ProtectSystem=strict
read-only disk
ProtectHome=
homes hidden
PrivateTmp=
private /tmp
Kernel and calls
ProtectKernel*=
no tunables/modules
SystemCallFilter=
a seccomp allow-list
RestrictAddressFamilies=
only the sockets it needs
Each line removes one avenue an attacker would use. systemd-analyze security scores what is left.

Seccomp and the last knobs

The final layer restricts the system calls themselves. SystemCallFilter=@system-service allows only the calls a normal service makes and blocks the rest with a seccomp filter (a kernel feature that screens system calls). SystemCallArchitectures=native stops the process from smuggling calls in through a foreign CPU ABI. RestrictAddressFamilies= limits it to IP and Unix sockets, RestrictNamespaces= stops it creating namespaces, and MemoryDenyWriteExecute= refuses memory that is writable and executable at once. That last one breaks every just-in-time compiler (systemd.exec(5) says so): leave it out for Java, Node.js, .NET or a PCRE JIT, and keep it for plain interpreted or compiled code like this Python app. Resource caps such as TasksMax= and LimitNOFILE= are capacity settings, not hardening; size them from the service's observed peak, or a busy service fails with "Too many open files".

/etc/systemd/system/hard-sandbox.service.d/30-syscall.conf
[Service]
# Allow the system-service syscall set only, on this CPU's native ABI, and make a
# blocked call return EPERM instead of killing the process. Sockets: IP and Unix only.
SystemCallFilter=@system-service
SystemCallErrorNumber=EPERM
SystemCallArchitectures=native
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
RestrictNamespaces=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes

The service still serves, because its normal work uses only allowed calls, and the score falls again, to 1.8 OK.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl restart hard-sandbox
$ curl -s http://127.0.0.1/
hello, euid=61325
$ systemd-analyze security hard-sandbox.service | tail -n 1
→ Overall exposure level for hard-sandbox.service: 1.8 OK :-)

Before you enforce a filter on a real service, find out what it calls. SystemCallLog= logs the listed calls through the audit subsystem without blocking them; run the service for a representative period with SystemCallLog=~@system-service (every call outside the set) and read the log. Here a throwaway unit logs @chown calls:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --pipe --wait -p 'SystemCallLog=@chown' chown deploy /var/tmp/hard-sandbox-probe ls -l /var/tmp/hard-sandbox-probe
… Finished with result: success Main processes terminated with: code=exited, status=0/SUCCESS … -rw-r--r-- 1 deploy root 4 Sep 27 09:29 /var/tmp/hard-sandbox-probe
$ sudo ausearch --input-logs -m SECCOMP -ts recent -i | grep -o 'comm=chown.*syscall=[a-z0-9_]*' | tail -n 1
comm=chown exe=/usr/lib/cargo/bin/coreutils/chown sig=SIG0 arch=aarch64 syscall=fchownat

The call was allowed (the file now belongs to deploy), and the audit log names it: syscall=fchownat from the uutils chown, with arch=aarch64 (an x86_64 server logs its own architecture and call names may differ). With auditd running, as here, the record lands in audit.log; without auditd the kernel writes it to the kernel log. It is worth seeing how the filter behaves when a call is blocked. By default systemd kills the process with SIGSYS; with SystemCallErrorNumber=EPERM the call instead returns an error and the program carries on. A transient systemd-run unit that denies the @chown group and then tries to change a file's owner shows both.

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemd-run --pipe --wait -p 'SystemCallFilter=~@chown' chown deploy /var/tmp/hard-sandbox-probe
Running as unit: run-p243782-i240709.service; invocation ID: 734aa73966d64b3b9074f5a5ba7e2b95 Finished with result: core-dump Main processes terminated with: code=dumped, status=31/SYS Service runtime: 584ms CPU time consumed: 22ms Memory peak: 1.8M (swap: 0B)
$ sudo systemd-run --pipe --wait -p 'SystemCallFilter=~@chown' -p SystemCallErrorNumber=EPERM chown deploy /var/tmp/hard-sandbox-probe
Running as unit: run-p243822-i240732.service chown: changing ownership of '/var/tmp/hard-sandbox-probe': Operation not permitted (os error 1) Finished with result: exit-code Main processes terminated with: code=exited, status=1/FAILURE Service runtime: 10ms CPU time consumed: 6ms Memory peak: 1.7M (swap: 0B)

The first is killed (status=31/SYS is SIGSYS); the second is refused with "Operation not permitted" and keeps running. SystemCallErrorNumber=EPERM is usually kinder to a real service, which can handle an error but not a sudden death mid-request. The price is diagnosis: an EPERM return leaves no record naming the blocked call, so the service reports an odd error far from the cause. That is one more reason to run the SystemCallLog= discovery first. Blocking too much is the common self-inflicted outage here, which is why @system-service (an allow-list systemd maintains) is the sane default rather than a hand-picked list.

Add what the service needs, not every directive
A tighter score is not the goal in itself; a service that no longer works is worse than one that scores 5. Add directives that match what the service actually does and test after each: a daemon that writes logs needs a ReadWritePaths or a LogsDirectory, one that runs helper programs cannot take a strict syscall filter, and one that talks to a Unix socket needs AF_UNIX left in RestrictAddressFamilies. Read the failure in journalctl -u NAME and loosen the one directive that caused it, rather than reaching for User=root again. And never sandbox a process that starts other people's work: cron, atd, sshd, getty, systemd-logind or a CI runner pass their restrictions on to every job and login they start. A read-only /etc inherited by every root cron job, or ProtectHome= hiding every authorized_keys from sshd, is an outage. Sandbox the job's own unit instead, as the scheduled jobs lesson does with its timer.

Verify, then roll it back

Confirm the live settings with systemctl show, which reports what is actually in effect, not what the files say. Then remember the escape hatch: every drop-in is removable. systemctl revert deletes them all and returns the unit to its packaged definition, which is the rollback for any of these changes.

deploy@web01 · Ubuntu 26.04 LTS
$ systemctl show -p User,NoNewPrivileges,ProtectSystem,CapabilityBoundingSet,SystemCallArchitectures hard-sandbox.service
CapabilityBoundingSet=cap_net_bind_service User=hard-sandbox ProtectSystem=strict NoNewPrivileges=yes SystemCallArchitectures=native
$ sudo systemctl revert hard-sandbox sudo systemctl restart hard-sandbox
Removed '/etc/systemd/system/hard-sandbox.service.d/30-syscall.conf'. Removed '/etc/systemd/system/hard-sandbox.service.d/10-confine.conf'. Removed '/etc/systemd/system/hard-sandbox.service.d/20-caps.conf'. Removed '/etc/systemd/system/hard-sandbox.service.d'.
$ curl -s http://127.0.0.1/ systemd-analyze security hard-sandbox.service | tail -n 1
hello, euid=0 → Overall exposure level for hard-sandbox.service: 9.6 UNSAFE :-{

Reverted, the service is back to euid=0 and 9.6 UNSAFE, proving the drop-ins were the only thing holding it down. On a packaged daemon you do the same through systemctl edit, measuring before and after and restarting to test. The example is nginx (sudo apt install nginx), a leaf daemon: it starts only its own worker processes. The drop-in lets it write only its logs, its cache and its pid file:

deploy@web01 · Ubuntu 26.04 LTS
$ sudo systemctl start nginx systemd-analyze security nginx.service | tail -n 1
→ Overall exposure level for nginx.service: 9.6 UNSAFE :-{
$ sudo systemctl edit --stdin --drop-in=50-secopslog nginx.service <<'EOF' [Service] # nginx starts only its own worker processes, so a sandbox fits it. ProtectSystem=strict # Where nginx writes: its logs, its cache and its pid file in /run. ReadWritePaths=/var/log/nginx /var/lib/nginx /run ProtectHome=yes PrivateTmp=yes PrivateDevices=yes NoNewPrivileges=yes ProtectKernelTunables=yes ProtectKernelModules=yes ProtectControlGroups=yes EOF
Successfully installed edited file '/etc/systemd/system/nginx.service.d/50-secopslog.conf'.
$ sudo systemctl restart nginx systemctl is-active nginx curl -s -o /dev/null -w 'HTTP %{http_code}\n' http://127.0.0.1/ systemd-analyze security nginx.service | tail -n 1
active HTTP 200 → Overall exposure level for nginx.service: 7.4 MEDIUM :-|
$ sudo systemctl revert nginx.service sudo systemctl restart nginx systemctl show -p DropInPaths nginx.service systemd-analyze security nginx.service | tail -n 1
Removed '/etc/systemd/system/nginx.service.d/50-secopslog.conf'. Removed '/etc/systemd/system/nginx.service.d'. DropInPaths= → Overall exposure level for nginx.service: 9.6 UNSAFE :-{

nginx restarted with the drop-in, answered HTTP 200, and moved from 9.6 UNSAFE to 7.4 MEDIUM; the revert removed the drop-in and the score went back to 9.6. More directives (a system call filter, a bounding set of the few capabilities its master process needs) would lower it further, each tested the same way. Not every daemon can take every restriction, which is why you measure, apply what fits, restart, and test that the service still does its job before you keep a drop-in.

Try this

Practise on the lesson's demo service, on nginx, or on another leaf service you can afford to break on a lab machine; never on ssh, networking, journald, dbus or logind, which can cut you off. Run systemd-analyze security NAME to read its score, and add a drop-in with sudo systemctl edit NAME containing ProtectSystem=strict, ProtectHome=yes, NoNewPrivileges=yes and PrivateTmp=yes (comments on their own lines). Restart it, confirm it still works, and read the new score. Then break something on purpose: add SystemCallFilter=@system-service and SystemCallArchitectures=native, restart, and if the service fails, read journalctl -u NAME for the killed call. Finish with sudo systemctl revert NAME and confirm systemctl show -p DropInPaths NAME is empty and the service is back to normal.

Takeaway

Measure a service with systemd-analyze security, then take away what it does not need: a dynamic user instead of root, a read-only filesystem with one writable StateDirectory, a single capability such as CAP_NET_BIND_SERVICE in place of root, and the @system-service syscall allow-list. Verify with systemctl show, keep comments on their own lines, and roll back with systemctl revert.

Quick check
01A service needs to listen on port 443. A teammate sets User=root so it can bind the port. What is the smaller change that meets the same need?
Incorrect — System accounts have no special right to low ports; only root or a process holding CAP_NET_BIND_SERVICE may bind them.
Correct — That single capability lets a non-root process bind a low port, so the service never runs as root; cap the bounding set to it as well.
Incorrect — NoNewPrivileges blocks gaining privileges; it does not grant the port-binding right, so the bind would still fail.
Incorrect — That changes the requirement rather than meeting it, and the question is how to bind 443 safely, which a capability does.
02You add ProtectSystem=strict, ProtectHome=yes and NoNewPrivileges=true to a unit drop-in, but the last one is written as "NoNewPrivileges=true # hardening". After daemon-reload, what is the effect of that line?
Incorrect — systemd does not strip trailing comments in unit files; only whole comment lines are ignored.
Incorrect — The other two lines parse fine; only the malformed line is dropped, which is what makes the mistake easy to miss.
Correct — The comment becomes part of the value, the boolean fails, and the setting defaults off with only a journal warning while the unit looks hardened.
Incorrect — The service still starts; the setting is just ignored, which is more dangerous than a hard failure would be.
03With SystemCallFilter=@system-service set, a service occasionally makes a call outside that set. By default, what happens, and what does SystemCallErrorNumber=EPERM change?
Correct — The default action is a SIGSYS kill; EPERM turns it into a survivable error, which is usually better for a service handling live requests.
Incorrect — Reversed: the default is not an errno at all but a SIGSYS kill; SystemCallErrorNumber replaces that kill with an errno.
Incorrect — A blocked call does not disable the unit; it affects only the calling process, and the setting changes how that call fails.
Incorrect — A filtered call never succeeds; the choice is only between killing the process and returning an error.

Related