seccomp: filter system calls

Docker's default profile, a stricter one built from it, and how to find the call that was refused.

Advanced15 min · lesson 12 of 24
Lesson files
The scripts, test data and local test servers this lesson uses, exactly as they ran on the lab machine (2 files, 1 KB): seccomp.tar.gz. The lab VM shares no folders with your computer, so fetch them inside the VM: cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/seccomp.tar.gz && tar -xzf seccomp.tar.gz, which creates ~/lab/seccomp/. SHA-256: 7d92be2603523fe0e22c44dc3020cbcc3ec323024bac123ec39d5ca26fd84c56

Docker attaches a seccomp filter to every container it starts, and most teams never look at it until something fails with Operation not permitted. This lesson covers what the default profile allows and refuses, how to read it, how to build a stricter profile from it, how to find out which system call was refused, and why an allow-list traced from your application with strace stops the container before your application starts. Along the way it settles a common review comment: security_opt: [seccomp=unconfined] on a debugging sidecar, "because strace needs this". On Docker 29 strace runs under the default profile, and that line would remove the filter between the container and a few hundred kernel entry points it never uses.

Use the main lab VM, in ~/lab/seccomp. The lesson files hold a Dockerfile for a small test image and a deliberately broken profile, min.json, used at the end.

What a seccomp filter does

Every request a process makes to the kernel, opening a file, creating a socket, starting a process, is a system call with a number. seccomp ("secure computing") lets a process attach a small BPF program that the kernel runs on each system call before doing any work. The program sees the system call number, the CPU architecture the call came through and the arguments, and returns an action: allow it, fail it with an error number, log it, or kill the process. Filters are inherited by child processes and can only be made stricter, never removed. runc attaches Docker's profile to every container it starts unless the container is --privileged or started with seccomp=unconfined:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker info --format '{{json .SecurityOptions}}'
["name=apparmor,profile=default","name=seccomp,profile=builtin","name=cgroupns"]
$ docker run --rm alpine:3.22 grep Seccomp /proc/self/status
Seccomp: 2 Seccomp_filters: 1

docker info reports profile=builtin, the default profile compiled into the engine. Inside a container, Seccomp: 2 means filter mode and Seccomp_filters: 1 means one filter is attached. An unconfined container shows Seccomp: 0.

The default profile is an allow-list

The test image adds keyctl, strace and the kernel headers to Alpine:

Dockerfile
FROM alpine:3.22
# keyctl to call add_key, strace to trace system calls, and the kernel headers
# that map a system call number to its name
RUN apk add --no-cache keyutils strace linux-headers
ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ docker build -q -t lab-seccomp:1 .
sha256:40acc1ad5eb653502ed56c0c6a206f9bab9de919f50f2f4ddade0c297d2cf449
ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm lab-seccomp:1 keyctl add user lab-demo lab-value @s
add_key: Operation not permitted
$ docker run --rm --security-opt seccomp=unconfined lab-seccomp:1 keyctl add user lab-demo lab-value @s
240258953
$ docker run --rm lab-seccomp:1 strace -e trace=fchmodat chmod 600 /etc/hostname
fchmodat(AT_FDCWD, "/etc/hostname", 0600) = 0 +++ exited with 0 +++

add_key stores a key in the kernel keyring, which is not namespaced, so containers would share it with the host. The default profile does not allow it, and the call fails with Operation not permitted (EPERM). The same error comes back when a capability is missing or AppArmor refuses something, so the message alone does not say which control said no. Running the same command once with --security-opt seccomp=unconfined does: here the key is created and its serial number printed, so the filter was the cause. Do that A/B test on a throwaway container while diagnosing, and never leave unconfined in a Compose file or a run script.

The last command answers the review comment from the opening. strace traced chmod under the default profile, as an ordinary root container with no added capability. ptrace is allowed on kernel 4.8 and later; on older kernels it could be used to get around seccomp and Docker blocked it. Attaching to a process that is not your descendant is a different matter. With Yama's ptrace_scope at 1, Ubuntu's default, it needs CAP_SYS_PTRACE (--cap-add SYS_PTRACE), and so does tracing a process of another UID on any host. On AppArmor hosts the profile must also allow it ("AppArmor and SELinux").

Reading the profile

The builtin profile is maintained in the moby/profiles repository, and the engine's go.mod says which version it was built with. The JSON form, default.json, is the starting point for any custom profile:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ curl -fsSL --retry 5 --retry-all-errors https://raw.githubusercontent.com/moby/moby/docker-v29.8.2/go.mod | grep profiles/seccomp
github.com/moby/profiles/seccomp v0.2.3
$ curl -fsSL --retry 5 --retry-all-errors -o default.json https://raw.githubusercontent.com/moby/profiles/seccomp/v0.2.3/seccomp/default.json sha256sum default.json jq -c "{defaultAction, defaultErrnoRet, arches: [.archMap[].architecture]}" default.json
536529b665dd0972c37bfb569f5d4ac8a53592e7b00752bc39ff063ca9864c74 default.json {"defaultAction":"SCMP_ACT_ERRNO","defaultErrnoRet":1,"arches":["SCMP_ARCH_X86_64","SCMP_ARCH_AARCH64","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_S390X","SCMP_ARCH_RISCV64","SCMP_ARCH_LOONGARCH64"]}
$ jq -c ".syscalls[] | select(.names | index(\"ptrace\")) | {names, includes}" default.json jq "[.syscalls[].names[]] | index(\"add_key\")" default.json
{"names":["process_vm_readv","process_vm_writev","ptrace"],"includes":{"minKernel":"4.8"}} {"names":["kcmp","pidfd_getfd","process_madvise","process_vm_readv","process_vm_writev","ptrace"],"includes":{"caps":["CAP_SYS_PTRACE"]}} null

Engine 29.8.2 uses moby/profiles/seccomp v0.2.3; for another engine version, replace the tag in the first URL. The --retry options are there because a single dropped connection to GitHub is common enough to break a script. The profile's defaultAction is SCMP_ACT_ERRNO with defaultErrnoRet 1: any system call not matched by a rule fails with EPERM. That makes it an allow-list, and the rest of the file is the list, grouped into rules. Some rules apply only under conditions. The two that mention ptrace show both kinds: one allows ptrace and process_vm_readv/writev on kernels 4.8 and later, the other adds more tracing calls when the container has CAP_SYS_PTRACE. Adding a capability therefore also opens the system calls that go with it, which is how --cap-add SYS_ADMIN makes mount reachable. add_key appears in no rule at all (null), so the default action refuses it.

archMap lists the CPU architectures the filter understands, each with sub-architectures: 32-bit x86 and x32 for x86-64, 32-bit ARM for aarch64. System call numbers differ between these: read is 0 on x86-64 and 3 on 32-bit x86. Rules are written as names and translated into numbers for each architecture in the list, so 32-bit programs keep working. A call that arrives through an architecture the filter does not list at all gets libseccomp's "bad architecture" action, which kills the process by default. Old advice warns that a 32-bit call can slip past a profile. That applies to filters that allow everything by default and deny a few calls, or that never check the architecture. With an allow-list like Docker's, an unmatched call is refused, so forgetting a name breaks the application instead of opening a hole.

One system call meets Docker's profile
1process makes a system call
number + architecture + arguments
2architecture in archMap?
no: bad-architecture action, process killed
3number matches an allow rule?
conditions: kernel version, capabilities, arguments
4yes: the kernel runs it
no: defaultAction, EPERM
Docker's default profile refuses anything it does not list. A profile that allows by default inverts this: everything it fails to name gets through.

A stricter profile, built from the default

Suppose a service has no reason to change file modes. Start from default.json, remove names, and test. Removing only chmod changes nothing on this VM:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ jq "(.syscalls[].names) -= [\"chmod\"]" default.json > only-chmod.json docker run --rm --security-opt seccomp=only-chmod.json alpine:3.22 chmod 600 /etc/hostname && echo "chmod still works"
chmod still works
$ jq "(.syscalls[].names) -= [\"chmod\", \"fchmod\", \"fchmodat\", \"fchmodat2\"]" default.json > no-chmod.json docker run --rm --security-opt seccomp=no-chmod.json alpine:3.22 sh -c "echo data > /tmp/f && cat /tmp/f && chmod 600 /tmp/f"
data chmod: /tmp/f: Operation not permitted

arm64 has no chmod system call. Like newer architectures, it only has the *at forms, so busybox's chmod uses fchmodat, and a name that does not exist on an architecture is skipped there. x86-64 still has chmod, and which call a binary makes depends on its C library, so test a profile on every architecture you deploy. The second profile removes the whole family, including fchmodat2 (added in kernel 6.6), and the container still writes and reads files while chmod fails with Operation not permitted.

A profile forked from default.json is frozen at the version you copied. Engine upgrades bring new rules for new system calls (clone3 and fchmodat2 are recent examples), and when a new C library starts using one, containers on the fork fail with EPERM while containers on the builtin profile work. At every engine upgrade, diff your fork against the new default.json and carry the additions over, then recreate the containers: the profile is copied into the container when it is created, as the inspect output below shows.

--security-opt seccomp=PATH reads the file on the client and sends its content to the daemon (Compose security_opt takes the same string, and the older seccomp:PATH spelling is still accepted). That has a consequence for audits:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ docker run -d --name lab-sc-default alpine:3.22 sleep 300 >/dev/null docker run -d --name lab-sc-custom --security-opt seccomp=no-chmod.json alpine:3.22 sleep 300 >/dev/null docker run -d --name lab-sc-off --security-opt seccomp=unconfined alpine:3.22 sleep 300 >/dev/null for c in lab-sc-default lab-sc-custom lab-sc-off; do printf '%s ' $c; docker inspect -f '{{json .HostConfig.SecurityOpt}}' $c | cut -c1-80 done
lab-sc-default null lab-sc-custom ["seccomp={\"defaultAction\":\"SCMP_ACT_ERRNO\",\"defaultErrnoRet\":1,\"archMap\ lab-sc-off ["seccomp=unconfined"]

A container on the builtin profile has no security options (null). An unconfined one shows seccomp=unconfined. A custom one shows the whole JSON document inline, not a file name, so a fleet check cannot look for a path; compare or hash the JSON against the profile you approved. To change the profile for every container, the daemon setting seccomp-profile in daemon.json takes a path ("Configuring the daemon safely" in Docker in depth covers the safe change procedure).

Finding out which call was refused

Denials from the default profile are silent. The process gets EPERM and nothing is logged. To see them without weakening the filter, add the SECCOMP_FILTER_FLAG_LOG flag to a copy of the profile, run the failing command once with each, and read the kernel log:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ jq ".flags = [\"SECCOMP_FILTER_FLAG_LOG\"]" default.json > default-log.json date "+%F %T" > since docker run --rm lab-seccomp:1 keyctl add user lab-demo lab-value @s docker run --rm --security-opt seccomp=default-log.json lab-seccomp:1 keyctl add user lab-demo lab-value @s
add_key: Operation not permitted add_key: Operation not permitted
$ sudo journalctl -k --since "$(cat since)" --no-pager -o cat --grep 'type=1326'
audit: type=1326 audit(1791403638.027:1993): auid=4294967295 uid=0 gid=0 ses=4294967295 subj=docker-default pid=534845 comm="keyctl" exe="/bin/keyctl" sig=0 arch=c00000b7 syscall=217 compat=0 ip=0xf672cf0eb438 code=0x50000
ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker run --rm lab-seccomp:1 grep -w 217 /usr/include/asm/unistd_64.h
#define __NR_add_key 217

Both runs were refused, and only the flagged one left a record. The fields that matter: type=1326 marks a seccomp record, comm and exe name the program, arch=c00000b7 is aarch64 (an amd64 host shows c000003e), syscall=217 is the number in that architecture's table, and code=0x50000 is the action taken, here ERRNO. The number becomes a name with the kernel headers of the same architecture; the test image has them, and on amd64 the same grep in the same image reads the x86-64 table, where add_key is 248. Hosts running auditd have ausyscall for the same lookup. In a test environment, "defaultAction": "SCMP_ACT_LOG" goes one step further. Everything the profile would refuse is allowed and logged, which is how you collect what a stricter profile is missing.

Why an allow-list from strace fails

The tempting shortcut is to trace the application and allow exactly what it called:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ docker run --rm lab-seccomp:1 sh -c 'strace -f -qq -o /tmp/t /bin/echo hello >/dev/null; sed -E "s/^[0-9]+ +//; s/\(.*//" /tmp/t | grep -v "^+++" | sort -u | xargs'
brk execve exit_group getuid mmap mprotect set_tid_address write
min.json
{
"defaultAction": "SCMP_ACT_ERRNO",
"archMap": [
{ "architecture": "SCMP_ARCH_AARCH64", "subArchitectures": ["SCMP_ARCH_ARM"] }
],
"syscalls": [
{
"names": ["brk", "execve", "exit_group", "getuid", "mmap", "mprotect", "set_tid_address", "write"],
"action": "SCMP_ACT_ALLOW"
}
]
}
ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ docker run --rm --security-opt seccomp=min.json alpine:3.22 /bin/echo hello
docker: Error response from daemon: failed to create task for container: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: error closing exec fds: get handle to /proc/thread-self/fd: fstatfs fsmount:fscontext:proc: operation not permitted Run 'docker run --help' for more information

The container never starts, and the error comes from runc, not from echo. runc loads the filter while setting up the container and then still has work to do under it before it executes your program: closing file descriptors, changing to the working directory, opening files under /proc. strace inside the container never sees those calls. Switch the same list to SCMP_ACT_LOG to find them:

ubuntu@secopslog-docker:~/lab/seccomp · Docker 29.8.2
$ jq ".defaultAction = \"SCMP_ACT_LOG\"" min.json > min-log.json S=$(date "+%F %T") docker run --rm --security-opt seccomp=min-log.json alpine:3.22 /bin/echo hello sleep 1 sudo journalctl -k --since "$S" --no-pager -o cat --grep "type=1326" > would-deny.log grep -o "comm=\"[^\"]*\"" would-deny.log | sort | uniq -c
hello 9 comm="runc:[2:INIT]"
$ N=$(grep -o "syscall=[0-9]*" would-deny.log | cut -d= -f2 | sort -nu | paste -sd "|") docker run --rm lab-seccomp:1 grep -wE "($N)" /usr/include/asm/unistd_64.h
#define __NR_epoll_ctl 21 #define __NR_fcntl 25 #define __NR_chdir 49 #define __NR_openat 56 #define __NR_close 57 #define __NR_close_range 436

Every logged call came from runc:[2:INIT], the runc process that becomes the container, and the names (fcntl, chdir, openat, close, close_range, plus Go runtime calls such as epoll_ctl or futex, which vary from run to run) are runc's own. The same gap applies to the C library start-up and to code paths your test did not reach. The working method is the one used above. Start from default.json, remove what you can show is unused, and run the result with the log flag under real traffic before you enforce it. For most services the default profile is the right setting, and the stricter profiles are worth their upkeep on the few services that would hurt most if compromised. Clean up:

ubuntu@secopslog-docker:~ · Docker 29.8.2
$ docker rm -f lab-sc-default lab-sc-custom lab-sc-off docker rmi lab-seccomp:1 rm -rf ~/lab/seccomp
lab-sc-default lab-sc-custom lab-sc-off Untagged: lab-seccomp:1 Deleted: sha256:40acc1ad5eb653502ed56c0c6a206f9bab9de919f50f2f4ddade0c297d2cf449
Quick check
01A debugging container has security_opt: [seccomp=unconfined] with the comment "strace needs this". It runs as root with Docker's default capabilities on a current kernel. What is the right review response?
Incorrect — The default profile has no deny list; it allows ptrace on kernel 4.8 and later, as the profile's rules show.
Incorrect — Privileged mode disables the seccomp filter as well, a much larger change than the one requested.
Incorrect — Adding capabilities does extend it (CAP_SYS_PTRACE adds more tracing calls), and plain ptrace needs no change at all.
Correct — The lab traces chmod under the builtin profile with no options.
02A service fails with Operation not permitted and you suspect seccomp. You want to know which system call is refused without turning the filter off. What do you do?
Incorrect — Docker records no such thing; inspect shows only the profile.
Correct — The lab's flagged run is still refused, and the record names syscall 217, which the headers map to add_key.
Incorrect — Default-profile refusals are silent; the lab's unflagged run left no record.
Incorrect — The A/B test tells you seccomp was the cause, not which call; nothing is printed.
03A team's profile uses "defaultAction": "SCMP_ACT_ALLOW" and denies chmod, fchmod and fchmodat. After a kernel and base-image upgrade, an updated tool changes file modes inside containers again, although its calls were refused before. What happened?
Incorrect — Filters are not tied to a kernel build; the old rules still match the old calls.
Incorrect — A library does not change ABI because of a kernel upgrade; nothing about the architecture changed.
Correct — A profile that allows by default lets any call it does not name through, including new ones.
Incorrect — A custom profile replaces the builtin one; nothing is merged.

Try this

Work through “Why an allow-list from strace fails” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.

Takeaway

If you keep one thing from seccomp: filter system calls, keep “Why an allow-list from strace fails”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.

Related