AppArmor and SELinux
Docker's AppArmor profile, a custom one, and how to tell a security module refused you.
cd ~/lab && curl -fsSLO https://secopslog.com/lab-files/docker-int/lsm.tar.gz && tar -xzf lsm.tar.gz, which creates ~/lab/lsm/. SHA-256: 8be695f0097279880f57bf23f38b21406c5b7e4fd96837359e909d49673ef15a/etc/apparmor.d. If the VM does not exist yet, create it on your workstation from the lab kit folder with ./setup/create-lab.sh --profile sec, then open a shell with multipass shell secopslog-docker-sec (limactl shell secopslog-docker-sec on Lima). Reset it afterwards with ./setup/create-lab.sh --profile sec --recreate. On the sec VM the ubuntu user is not in the docker group, so every Docker command below uses sudo.mount: mounting none on /mnt/x failed: Permission denied, from a container that was started with --cap-add SYS_ADMIN for exactly that mount. The capability is there, and Docker's seccomp profile allows mount once CAP_SYS_ADMIN is added ("seccomp: filter system calls"). The kernel log has nothing to say about it. The refusal comes from a third layer, a Linux Security Module, and this one does not log. This lesson covers the two modules Docker integrates with, AppArmor and SELinux: what Docker's AppArmor profile blocks, how to prove that it was the cause, how to write, load and debug a profile of your own, and what an SELinux host needs before its labels apply to containers.
Mandatory access control in one paragraph
Ordinary file permissions are discretionary: the owner of a file decides, and root passes every check. A Linux Security Module (LSM) adds hooks in the kernel that run after those checks and can refuse an operation even for root with every capability. That is mandatory access control, a policy loaded by the host administrator, which processes cannot change. AppArmor writes its policy as per-program profiles of paths and permissions, and ships enabled on Ubuntu, Debian and their derivatives. SELinux labels every process and file with a type and allows access by type, and ships enforcing on Fedora, RHEL and its rebuilds. A host runs one of the two as its main MAC module.
docker-default
On an AppArmor host, Docker generates a profile called docker-default, loads it into the kernel when the daemon starts, and applies it to every container that does not ask for something else:
docker info lists apparmor with profile=default. Inside the container, /proc/self/attr/apparmor/current reads docker-default (enforce). On kernels with several stacked LSMs that path is the reliable one; the older /proc/self/attr/current still works on this host. docker inspect shows the profile in AppArmorProfile while SecurityOpt stays empty, and aa-status on the host lists the profile and every confined process, here the container's busybox running sleep.
The profile is generated from a template in the moby/profiles repository. It allows ordinary file and network access and leaves capability decisions to Docker's capability set. What it adds is a set of denials for things containers should never do even when someone hands them a capability: mounting filesystems, writing to most of /proc (sysrq-trigger, kcore, sys/kernel and others) and to /sys, and tracing processes outside the profile. --privileged replaces the profile with unconfined.
A refusal that leaves no log line
The mount fails, and the kernel log of the last minute holds no AppArmor denial for docker-default. That is not a logging problem. docker-default refuses mounts with an explicit deny mount rule, and AppArmor does not log explicit deny rules unless the rule is written audit deny. The third command is the diagnosis. The same container with --security-opt apparmor=unconfined mounts the tmpfs. If removing the profile changes the result and nothing else changed, the profile was the cause.
Run that A/B test only on a disposable host or a throwaway container, and treat it as a measurement, never as a fix. With the capability still added, the unconfined container could mount anything its other settings allow, which is the escape path "Capability and kernel escapes" describes. The same reasoning sorts out any Permission denied or Operation not permitted from a container:
A profile of your own
docker-default is broad so that it fits every image. A profile for one workload can say what the process may do and refuse the rest, and those refusals are logged. This one lets a container read anything, run busybox, and write only under /data:
abi <abi/4.0>,include <tunables/global># lab-app: read the image, run busybox, write only under /data.# Anything not allowed here is denied, and AppArmor logs the denial.profile lab-app flags=(attach_disconnected,mediate_deleted) {include <abstractions/base>/** r,/bin/busybox mrix,/data/** w,signal (receive) peer=unconfined,}
abi <abi/4.0> pins the policy language version and tunables/global defines variables such as @{PROC}. attach_disconnected makes paths resolve inside the container's own mount namespace, which every container profile needs. abstractions/base covers what almost every program touches, such as shared libraries, /dev/null and locale files. /** r allows reading everything; /bin/busybox mrix allows mapping and executing busybox, and ix keeps the new program in the same profile. /data/** w is the only write rule. signal (receive) peer=unconfined lets the container runtime deliver docker stop. Everything else is refused, including every capability, because the profile has no capability rule, and all networking, because it has no network rule.
Docker does not load custom profiles. You load them with apparmor_parser, and keeping the file in /etc/apparmor.d lets the AppArmor service load it again at boot:
The container runs under lab-app (enforce). Writing /data/out works; running apk, a separate binary, fails, and so does creating /etc/lab-x. Unlike docker-default's explicit denies, these refusals come from rules the profile does not have, and AppArmor logs each one with the operation (exec, mknod), the profile and the path, which is the information you need to decide whether a rule is missing. Without auditd the kernel rate-limits these messages, so a large burst of denials can lose lines; hosts running auditd keep them in /var/log/audit/audit.log. The same command under docker-default runs apk and writes /etc/lab-x, confirming that the profile made the difference.
Writing a profile is an iterative job, and complain mode is the tool for it. A profile in complain mode logs what it would refuse and allows it. apparmor_parser -C loads a profile in complain mode:
In complain mode the write succeeds and the log shows apparmor="ALLOWED" for each step of it (mknod, open, file_perm). The working method is to load a new profile in complain mode, run the application through its real workload, turn the logged accesses that are legitimate into rules (aa-logprof from apparmor-utils can propose them), and then reload it in enforce mode, which the last command of that step did. In Compose the setting is security_opt: [apparmor=lab-app]. The failure you will meet first with custom profiles is a name the kernel does not know, for example after a host was rebuilt and the file was never copied to /etc/apparmor.d:
runc cannot switch the container to a profile that is not loaded, so the container is not created; the useful part of the message is unable to apply apparmor profile. Check aa-status on that host. Clean up, which unloads the profile with apparmor_parser -R and removes its file:
SELinux hosts
The lab VMs run Ubuntu, so SELinux is described here from the Docker and Fedora documentation, not demonstrated. With SELinux, container processes run as the type container_t and files meant for containers carry container_file_t; the policy allows container_t to touch little else. Each container also gets its own pair of MCS categories, such as s0:c12,c34, so two containers with the same UID still cannot read each other's files. ps -eZ shows the process labels and ls -Z the file labels.
Docker Engine from Docker's own packages does not enable SELinux support by default. Until the daemon runs with it, containers are not given container_t labels and the :z and :Z volume options relabel nothing:
{"selinux-enabled": true}
Distribution packages such as Fedora's moby-engine enable it already, so check docker info: an SELinux-enabled daemon lists selinux under Security Options. Bind mounts then need a label before a container can use them. :z gives the content a shared container label and :Z a label private to one container; "Volumes, bind mounts and tmpfs in practice" in Docker in depth covers which to use and the host paths never to relabel. A refusal appears as an AVC record, which ausearch -m avc -ts recent finds on hosts running auditd. For a single container, --security-opt label=disable is the equivalent of apparmor=unconfined and belongs in the same A/B test, never in production. setenforce 0 turns SELinux off for the whole host and every workload on it; it is not a fix for one container.
--cap-add SYS_ADMIN cannot mount a tmpfs. The host's kernel log shows no AppArmor denial. With --security-opt apparmor=unconfined added, the mount works. What is going on?--security-opt apparmor=web-strict fail to start with "unable to apply apparmor profile". Other containers run fine. What is the likely cause?:Z seem to make no difference, and ps -eZ does not show container processes as container_t. What is missing?Try this
Work through “SELinux hosts” yourself on a sandbox you can throw away, following the commands above in order. Then break one step deliberately and re-run, so you have seen the failure before it finds you.
Takeaway
If you keep one thing from apparmor and selinux, keep “SELinux hosts”. Decide now which check you will run when this shows up on a live system, and write it somewhere your team will find it.