seccomp: filter system calls
Docker's default profile, a stricter one built from it, and how to find the call that was refused.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/seccomp.tar.gz && tar -xzf seccomp.tar.gz, which creates ~/lab/seccomp/. SHA-256: 7d92be2603523fe0e22c44dc3020cbcc3ec323024bac123ec39d5ca26fd84c56Docker attaches a seccomp filter to every container it starts, and most teams never look at it until something fails with Operation not permitted. This lesson covers what the default profile allows and refuses, how to read it, how to build a stricter profile from it, how to find out which system call was refused, and why an allow-list traced from your application with strace stops the container before your application starts. Along the way it settles a common review comment: security_opt: [seccomp=unconfined] on a debugging sidecar, "because strace needs this". On Docker 29 strace runs under the default profile, and that line would remove the filter between the container and a few hundred kernel entry points it never uses.
Use the main lab VM, in ~/lab/seccomp. The lesson files hold a Dockerfile for a small test image and a deliberately broken profile, min.json, used at the end.
What a seccomp filter does
Every request a process makes to the kernel, opening a file, creating a socket, starting a process, is a system call with a number. seccomp ("secure computing") lets a process attach a small BPF program that the kernel runs on each system call before doing any work. The program sees the system call number, the CPU architecture the call came through and the arguments, and returns an action: allow it, fail it with an error number, log it, or kill the process. Filters are inherited by child processes and can only be made stricter, never removed. runc attaches Docker's profile to every container it starts unless the container is --privileged or started with seccomp=unconfined:
docker info reports profile=builtin, the default profile compiled into the engine. Inside a container, Seccomp: 2 means filter mode and Seccomp_filters: 1 means one filter is attached. An unconfined container shows Seccomp: 0.
The default profile is an allow-list
The test image adds keyctl, strace and the kernel headers to Alpine:
FROM alpine:3.22# keyctl to call add_key, strace to trace system calls, and the kernel headers# that map a system call number to its nameRUN apk add --no-cache keyutils strace linux-headers
add_key stores a key in the kernel keyring, which is not namespaced, so containers would share it with the host. The default profile does not allow it, and the call fails with Operation not permitted (EPERM). The same error comes back when a capability is missing or AppArmor refuses something, so the message alone does not say which control said no. Running the same command once with --security-opt seccomp=unconfined does: here the key is created and its serial number printed, so the filter was the cause. Do that A/B test on a throwaway container while diagnosing, and never leave unconfined in a Compose file or a run script.
The last command answers the review comment from the opening. strace traced chmod under the default profile, as an ordinary root container with no added capability. ptrace is allowed on kernel 4.8 and later; on older kernels it could be used to get around seccomp and Docker blocked it. Attaching to a process that is not your descendant is a different matter. With Yama's ptrace_scope at 1, Ubuntu's default, it needs CAP_SYS_PTRACE (--cap-add SYS_PTRACE), and so does tracing a process of another UID on any host. On AppArmor hosts the profile must also allow it ("AppArmor and SELinux").
Reading the profile
The builtin profile is maintained in the moby/profiles repository, and the engine's go.mod says which version it was built with. The JSON form, default.json, is the starting point for any custom profile:
Engine 29.8.2 uses moby/profiles/seccomp v0.2.3; for another engine version, replace the tag in the first URL. The --retry options are there because a single dropped connection to GitHub is common enough to break a script. The profile's defaultAction is SCMP_ACT_ERRNO with defaultErrnoRet 1: any system call not matched by a rule fails with EPERM. That makes it an allow-list, and the rest of the file is the list, grouped into rules. Some rules apply only under conditions. The two that mention ptrace show both kinds: one allows ptrace and process_vm_readv/writev on kernels 4.8 and later, the other adds more tracing calls when the container has CAP_SYS_PTRACE. Adding a capability therefore also opens the system calls that go with it, which is how --cap-add SYS_ADMIN makes mount reachable. add_key appears in no rule at all (null), so the default action refuses it.
archMap lists the CPU architectures the filter understands, each with sub-architectures: 32-bit x86 and x32 for x86-64, 32-bit ARM for aarch64. System call numbers differ between these: read is 0 on x86-64 and 3 on 32-bit x86. Rules are written as names and translated into numbers for each architecture in the list, so 32-bit programs keep working. A call that arrives through an architecture the filter does not list at all gets libseccomp's "bad architecture" action, which kills the process by default. Old advice warns that a 32-bit call can slip past a profile. That applies to filters that allow everything by default and deny a few calls, or that never check the architecture. With an allow-list like Docker's, an unmatched call is refused, so forgetting a name breaks the application instead of opening a hole.
A stricter profile, built from the default
Suppose a service has no reason to change file modes. Start from default.json, remove names, and test. Removing only chmod changes nothing on this VM:
arm64 has no chmod system call. Like newer architectures, it only has the *at forms, so busybox's chmod uses fchmodat, and a name that does not exist on an architecture is skipped there. x86-64 still has chmod, and which call a binary makes depends on its C library, so test a profile on every architecture you deploy. The second profile removes the whole family, including fchmodat2 (added in kernel 6.6), and the container still writes and reads files while chmod fails with Operation not permitted.
A profile forked from default.json is frozen at the version you copied. Engine upgrades bring new rules for new system calls (clone3 and fchmodat2 are recent examples), and when a new C library starts using one, containers on the fork fail with EPERM while containers on the builtin profile work. At every engine upgrade, diff your fork against the new default.json and carry the additions over, then recreate the containers: the profile is copied into the container when it is created, as the inspect output below shows.
--security-opt seccomp=PATH reads the file on the client and sends its content to the daemon (Compose security_opt takes the same string, and the older seccomp:PATH spelling is still accepted). That has a consequence for audits:
A container on the builtin profile has no security options (null). An unconfined one shows seccomp=unconfined. A custom one shows the whole JSON document inline, not a file name, so a fleet check cannot look for a path; compare or hash the JSON against the profile you approved. To change the profile for every container, the daemon setting seccomp-profile in daemon.json takes a path ("Configuring the daemon safely" in Docker in depth covers the safe change procedure).
Finding out which call was refused
Denials from the default profile are silent. The process gets EPERM and nothing is logged. To see them without weakening the filter, add the SECCOMP_FILTER_FLAG_LOG flag to a copy of the profile, run the failing command once with each, and read the kernel log:
Both runs were refused, and only the flagged one left a record. The fields that matter: type=1326 marks a seccomp record, comm and exe name the program, arch=c00000b7 is aarch64 (an amd64 host shows c000003e), syscall=217 is the number in that architecture's table, and code=0x50000 is the action taken, here ERRNO. The number becomes a name with the kernel headers of the same architecture; the test image has them, and on amd64 the same grep in the same image reads the x86-64 table, where add_key is 248. Hosts running auditd have ausyscall for the same lookup. In a test environment, "defaultAction": "SCMP_ACT_LOG" goes one step further. Everything the profile would refuse is allowed and logged, which is how you collect what a stricter profile is missing.
Why an allow-list from strace fails
The tempting shortcut is to trace the application and allow exactly what it called:
{"defaultAction": "SCMP_ACT_ERRNO","archMap": [{ "architecture": "SCMP_ARCH_AARCH64", "subArchitectures": ["SCMP_ARCH_ARM"] }],"syscalls": [{"names": ["brk", "execve", "exit_group", "getuid", "mmap", "mprotect", "set_tid_address", "write"],"action": "SCMP_ACT_ALLOW"}]}
The container never starts, and the error comes from runc, not from echo. runc loads the filter while setting up the container and then still has work to do under it before it executes your program: closing file descriptors, changing to the working directory, opening files under /proc. strace inside the container never sees those calls. Switch the same list to SCMP_ACT_LOG to find them:
Every logged call came from runc:[2:INIT], the runc process that becomes the container, and the names (fcntl, chdir, openat, close, close_range, plus Go runtime calls such as epoll_ctl or futex, which vary from run to run) are runc's own. The same gap applies to the C library start-up and to code paths your test did not reach. The working method is the one used above. Start from default.json, remove what you can show is unused, and run the result with the log flag under real traffic before you enforce it. For most services the default profile is the right setting, and the stricter profiles are worth their upkeep on the few services that would hurt most if compromised. Clean up:
security_opt: [seccomp=unconfined] with the comment "strace needs this". It runs as root with Docker's default capabilities on a current kernel. What is the right review response?Operation not permitted and you suspect seccomp. You want to know which system call is refused without turning the filter off. What do you do?"defaultAction": "SCMP_ACT_ALLOW" and denies chmod, fchmod and fchmodat. After a kernel and base-image upgrade, an updated tool changes file modes inside containers again, although its calls were refused before. What happened?Try this
Work through “Why an allow-list from strace fails” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from seccomp: filter system calls, keep “Why an allow-list from strace fails”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.