Sandboxing Linux services with systemd security directives

Turn a plain unit file into a locked-down sandbox with ProtectSystem, NoNewPrivileges, and capability limits.

Mar 10, 2026·Updated ·5 min readAdvanced·By SecOpsLog · documentation-verified
bash — what a default unit exposes
systemd-analyze security api.service | tail -3
→ Overall exposure level for api.service: 9.6 UNSAFE 😨
systemd-analyze security api.service | grep -E "✗" | head -5
✗ PrivateTmp= Service has access to other software's temporary files 0.1
✗ ProtectSystem= Service has full access to the OS file hierarchy 0.2
✗ NoNewPrivileges= Service processes may acquire new privileges 0.2
✗ CapabilityBoundingSet= Service runs with full capability set 0.2
✗ SystemCallFilter= Service does not filter system calls 0.2
the number is exposure: 10 is everything open, 0 is nothing left to take away

A service that runs as root with the whole filesystem writable, every capability and every syscall available is the default unit file, and it is what most packaged daemons ship with. systemd can take each of those away in the unit itself, with no container and no change to the program: a read-only view of the OS, a private /tmp, an empty capability set, a syscall allowlist, an ephemeral uid. The security verb of systemd-analyze scores what is left exposed and names the next directive to add, which turns hardening into a loop rather than a leap.

Directives by what they take away

The layers, in the order to add them

LayerDirectivesWhat the process losesWhat usually breaks
identityDynamicUser=yes (or User=), NoNewPrivileges=yesroot; the ability to gain privileges through setuid or file capabilitiesdaemons that expect a fixed uid in config files or on existing data directories
filesystemProtectSystem=strict, ProtectHome=yes, PrivateTmp=yes, PrivateDevices=yes, ReadWritePaths=writes anywhere except the listed paths; /home; a shared /tmp; physical deviceslogs written under /var/log/<app> without a ReadWritePaths entry; plugins loaded from a home directory
kernel surfaceProtectKernelTunables=yes, ProtectKernelModules=yes, ProtectKernelLogs=yes, ProtectControlGroups=yes, ProtectProc=invisiblesysctl writes, module loading, dmesg, cgroup edits, other processes in /procmonitoring agents that read other processes; tuning scripts that write sysctls
capabilitiesCapabilityBoundingSet= (empty) or a short list, AmbientCapabilities= for the one it needseverything root could do that is not in the listbinding a port below 1024 (add CAP_NET_BIND_SERVICE), chown of dropped files
syscalls and miscSystemCallFilter=@system-service, SystemCallErrorNumber=EPERM, RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX, RestrictNamespaces=yes, LockPersonality=yes, MemoryDenyWriteExecute=yes, RestrictRealtime=yesunusual syscalls, raw and packet sockets, namespace creation, W+X memoryJIT runtimes (Java, Node, .NET) with MemoryDenyWriteExecute; anything that shells out to a tool the filter blocks
/etc/systemd/system/api.service.d/hardening.conf
[Service]
# identity
DynamicUser=yes
NoNewPrivileges=yes
# filesystem
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ReadWritePaths=/var/lib/api /var/log/api
# kernel surface
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectProc=invisible
# capabilities: none, except the one needed to bind :443
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
AmbientCapabilities=CAP_NET_BIND_SERVICE
# syscalls and misc
SystemCallFilter=@system-service
SystemCallErrorNumber=EPERM
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
RestrictNamespaces=yes
LockPersonality=yes
RestrictRealtime=yes
UMask=0077

A drop-in under api.service.d/ keeps the hardening separate from the vendor's unit, so a package upgrade replaces the vendor file and leaves yours; systemctl daemon-reload then restart applies it. DynamicUser=yes does more than pick a uid: it implies PrivateTmp, RemoveIPC and a few protections, and it allocates the uid at start, so the /var/lib/api state directory is better declared as StateDirectory=api and owned by systemd for whatever uid is chosen. MemoryDenyWriteExecute is deliberately absent above because it breaks every JIT; add it for a Go or C daemon.

One layer, restart, exercise, read the journal

Add one layer, restart, run the service through its real work (a request, a scheduled job, a log rotation), and read the journal for EPERM, EACCES or Read-only file system. SystemCallErrorNumber=EPERM is what makes a blocked syscall show up as an error the program reports rather than as a silent kill, and strace -f -p <pid> on a stuck process names the exact call. When a layer breaks something, the fix is almost always narrower than removing the layer: one more entry in ReadWritePaths, one capability in the bounding set, one syscall group added to the filter (SystemCallFilter=@system-service @mount, for a daemon that must mount).

bash — the loop, one iteration
systemctl daemon-reload && systemctl restart api && sleep 2 && curl -sS -o /dev/null -w "%{http_code}\n" https://localhost/health
200
journalctl -u api --since "-1m" -p warning --no-pager
api[8123]: open /var/log/api/access.log: read-only file system
ProtectSystem=strict did its job; /var/log/api was missing from ReadWritePaths (added above)
systemd-analyze security api.service | tail -1
→ Overall exposure level for api.service: 1.8 OK 🙂
from 9.6 to 1.8; the remaining lines are the trade-offs documented in the drop-in

Two directives make the sandbox easier to live with rather than harder. StateDirectory=api and LogsDirectory=api have systemd create /var/lib/api and /var/log/api with the right owner at start and add them to the writable set, so a DynamicUser service does not need ReadWritePaths at all for its own data. And systemd-analyze security --threshold= turns the score into a gate: with --threshold=3 the command exits non-zero for a unit that exposes more than that, which is the check to run in the pipeline that ships unit files, so a drop-in that someone loosened to make a test pass does not reach the fleet unnoticed.

The score measures exposure, and a low number is not the goal by itself
A database that needs a fixed uid and wide ReadWritePaths will not reach 1.0 and should not be forced to. Record why each directive is absent as a comment in the drop-in, re-run the analysis after package upgrades (a vendor unit can add ExecStartPre lines that need paths yours forbids), and keep the drop-in in configuration management so a manual edit on one host cannot drift.
Sandboxes well vs needs care
Sandboxes well
HTTP APIs and workers that write one directory
Exporters and agents that read the network
Static Go and Rust binaries
Anything already containerised
Needs care
Databases: fixed uid, large ReadWritePaths
JIT runtimes and MemoryDenyWriteExecute
Daemons that shell out to system tools
Monitoring agents that read /proc for all users
Go deeper in a courseLinux hardeningsystemd sandboxing next to SSH, the firewall, SELinux and auditd on one host.View course

The capabilities layer is the same mechanism Linux capabilities describes for containers, with AmbientCapabilities= playing the role of cap_add; the syscall filter is the seccomp profile a container runtime applies by default, applied here to a plain process. A service hardened this way and a container with restricted Pod Security have removed roughly the same things, which is a useful way to explain to a team why the container is not the only path to a small attack surface.

Related posts

Quick reference