Sandboxing Linux services with systemd security directives
Turn a plain unit file into a locked-down sandbox with ProtectSystem, NoNewPrivileges, and capability limits.
systemd-analyze security api.service | tail -3→ Overall exposure level for api.service: 9.6 UNSAFE 😨systemd-analyze security api.service | grep -E "✗" | head -5✗ PrivateTmp= Service has access to other software's temporary files 0.1✗ ProtectSystem= Service has full access to the OS file hierarchy 0.2✗ NoNewPrivileges= Service processes may acquire new privileges 0.2✗ CapabilityBoundingSet= Service runs with full capability set 0.2✗ SystemCallFilter= Service does not filter system calls 0.2the number is exposure: 10 is everything open, 0 is nothing left to take awayA service that runs as root with the whole filesystem writable, every capability and every syscall available is the default unit file, and it is what most packaged daemons ship with. systemd can take each of those away in the unit itself, with no container and no change to the program: a read-only view of the OS, a private /tmp, an empty capability set, a syscall allowlist, an ephemeral uid. The security verb of systemd-analyze scores what is left exposed and names the next directive to add, which turns hardening into a loop rather than a leap.
Directives by what they take away
The layers, in the order to add them
| Layer | Directives | What the process loses | What usually breaks |
|---|---|---|---|
| identity | DynamicUser=yes (or User=), NoNewPrivileges=yes | root; the ability to gain privileges through setuid or file capabilities | daemons that expect a fixed uid in config files or on existing data directories |
| filesystem | ProtectSystem=strict, ProtectHome=yes, PrivateTmp=yes, PrivateDevices=yes, ReadWritePaths= | writes anywhere except the listed paths; /home; a shared /tmp; physical devices | logs written under /var/log/<app> without a ReadWritePaths entry; plugins loaded from a home directory |
| kernel surface | ProtectKernelTunables=yes, ProtectKernelModules=yes, ProtectKernelLogs=yes, ProtectControlGroups=yes, ProtectProc=invisible | sysctl writes, module loading, dmesg, cgroup edits, other processes in /proc | monitoring agents that read other processes; tuning scripts that write sysctls |
| capabilities | CapabilityBoundingSet= (empty) or a short list, AmbientCapabilities= for the one it needs | everything root could do that is not in the list | binding a port below 1024 (add CAP_NET_BIND_SERVICE), chown of dropped files |
| syscalls and misc | SystemCallFilter=@system-service, SystemCallErrorNumber=EPERM, RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX, RestrictNamespaces=yes, LockPersonality=yes, MemoryDenyWriteExecute=yes, RestrictRealtime=yes | unusual syscalls, raw and packet sockets, namespace creation, W+X memory | JIT runtimes (Java, Node, .NET) with MemoryDenyWriteExecute; anything that shells out to a tool the filter blocks |
[Service]# identityDynamicUser=yesNoNewPrivileges=yes# filesystemProtectSystem=strictProtectHome=yesPrivateTmp=yesPrivateDevices=yesReadWritePaths=/var/lib/api /var/log/api# kernel surfaceProtectKernelTunables=yesProtectKernelModules=yesProtectKernelLogs=yesProtectControlGroups=yesProtectProc=invisible# capabilities: none, except the one needed to bind :443CapabilityBoundingSet=CAP_NET_BIND_SERVICEAmbientCapabilities=CAP_NET_BIND_SERVICE# syscalls and miscSystemCallFilter=@system-serviceSystemCallErrorNumber=EPERMRestrictAddressFamilies=AF_INET AF_INET6 AF_UNIXRestrictNamespaces=yesLockPersonality=yesRestrictRealtime=yesUMask=0077
A drop-in under api.service.d/ keeps the hardening separate from the vendor's unit, so a package upgrade replaces the vendor file and leaves yours; systemctl daemon-reload then restart applies it. DynamicUser=yes does more than pick a uid: it implies PrivateTmp, RemoveIPC and a few protections, and it allocates the uid at start, so the /var/lib/api state directory is better declared as StateDirectory=api and owned by systemd for whatever uid is chosen. MemoryDenyWriteExecute is deliberately absent above because it breaks every JIT; add it for a Go or C daemon.
One layer, restart, exercise, read the journal
Add one layer, restart, run the service through its real work (a request, a scheduled job, a log rotation), and read the journal for EPERM, EACCES or Read-only file system. SystemCallErrorNumber=EPERM is what makes a blocked syscall show up as an error the program reports rather than as a silent kill, and strace -f -p <pid> on a stuck process names the exact call. When a layer breaks something, the fix is almost always narrower than removing the layer: one more entry in ReadWritePaths, one capability in the bounding set, one syscall group added to the filter (SystemCallFilter=@system-service @mount, for a daemon that must mount).
systemctl daemon-reload && systemctl restart api && sleep 2 && curl -sS -o /dev/null -w "%{http_code}\n" https://localhost/health200journalctl -u api --since "-1m" -p warning --no-pagerapi[8123]: open /var/log/api/access.log: read-only file systemProtectSystem=strict did its job; /var/log/api was missing from ReadWritePaths (added above)systemd-analyze security api.service | tail -1→ Overall exposure level for api.service: 1.8 OK 🙂from 9.6 to 1.8; the remaining lines are the trade-offs documented in the drop-inTwo directives make the sandbox easier to live with rather than harder. StateDirectory=api and LogsDirectory=api have systemd create /var/lib/api and /var/log/api with the right owner at start and add them to the writable set, so a DynamicUser service does not need ReadWritePaths at all for its own data. And systemd-analyze security --threshold= turns the score into a gate: with --threshold=3 the command exits non-zero for a unit that exposes more than that, which is the check to run in the pipeline that ships unit files, so a drop-in that someone loosened to make a test pass does not reach the fleet unnoticed.
The capabilities layer is the same mechanism Linux capabilities describes for containers, with AmbientCapabilities= playing the role of cap_add; the syscall filter is the seccomp profile a container runtime applies by default, applied here to a plain process. A service hardened this way and a container with restricted Pod Security have removed roughly the same things, which is a useful way to explain to a team why the container is not the only path to a small attack surface.